Katalog Plus
Bibliothek der Frankfurt UAS
Bald neuer Katalog: sichern Sie sich schon vorab Ihre persönlichen Merklisten im Nutzerkonto: Anleitung.
Dieses Ergebnis aus BASE kann Gästen nicht angezeigt werden.  Login für vollen Zugriff.

Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin

Title: Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
Authors: Clérice, Thibault; Bawden, Rachel; Glaise, Anthony; Pinche, Ariane; Smith, David
Publication Year: 2026
Collection: ArXiv.org (Cornell University Library)
Subject Terms: Computation and Language
Description: Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and normalized digital editions. While ATR models trained on more palaeographically-oriented datasets such as CATMuS have shown greater generalizability, their raw outputs remain poorly compatible with most readers and downstream NLP tools, thus creating a usability gap. On the other hand, ATR models trained to produce normalized outputs have been shown to struggle to adapt to new domains and tend to over-normalize and hallucinate. We introduce the task of Pre-Editorial Normalization (PEN), which consists in normalizing graphemic ATR output according to editorial conventions, which has the advantage of keeping an intermediate step with palaeographic fidelity while providing a normalized version for practical usability. We present a new dataset derived from the CoMMA corpus and aligned with digitized Old French and Latin editions using passim. We also produce a manually corrected gold-standard evaluation set. We benchmark this resource using ByT5-based sequence-to-sequence models on normalization and pre-annotation tasks. Our contributions include the formal definition of PEN, a 4.66M-sample silver training corpus, a 1.8k-sample gold evaluation set, and a normalization model achieving a 6.7% CER, substantially outperforming previous models for this task.
Document Type: text
Language: unknown
Relation: http://arxiv.org/abs/2602.13905
Availability: http://arxiv.org/abs/2602.13905
Accession Number: edsbas.63AF2C07
Database: BASE