News
Automatic Restoration of Birchbark Manuscripts Using Masked Language Modeling
Abstract
This work addresses the automatic restoration of lacunae in Old Novgorodian birchbark manuscripts, a challenging task due to the small available corpus and the distinctive dialectal features of the target variety. We systematically compare three BERT-like encoders – mBERT, BERTislav, and ModernBERT – in character-level and token-level prediction modes, zero-shot and after fine-tuning, on a purpose-built corpus of Old Russian and Old Church Slavonic texts. After domain fine-tuning, ModernBERT achieves the strongest restoration, reaching token-level top-1 accuracy of 30.72% on real editorial lacunae and 91.40% on artificially masked text. At the character level, however, the fine-tuned encoders are matched or surpassed by a simple character n-gram baseline, revealing a structural mismatch between subword pre-training and single-character prediction. We further probe the frozen embeddings for document genre and date: they carry useful signal for dating, whereas genre proves largely a surface-orthographic property that a character TF-IDF baseline captures as well or better.
Keywords
Edition
Proceedings of the Institute for System Programming, vol. 38, issue 6, part 1, 2026, pp. 177-192
ISSN 2220-6426 (Online), ISSN 2079-8156 (Print).
DOI: 10.15514/ISPRAS-2026-38(6)-11
For citation
Full text of the paper in pdf
Back to the contents of the volume