BDRC/monlamai-transcriptions
Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.
Update README.md
Add files using upload-large-folder tool
Reduce columns to image/transcription/id/mw_id/technology/script/script_4; card fixes
MonlamAI transcriptions (3,072 · manuscript u-med, CC0) — private test upload
initial commit
