CoolFace
Datasetpublic

BDRC/monlamai-transcriptions

Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.

sourceHugging Facecc0-1.0updated 1mo agoView on Hugging Face
0likes55downloads
5 commits on main
ace7f5b1mo ago

Update README.md

Eroux
d092c041mo ago

Add files using upload-large-folder tool

Eroux
260ab281mo ago

Reduce columns to image/transcription/id/mw_id/technology/script/script_4; card fixes

Eroux
a258bea1mo ago

MonlamAI transcriptions (3,072 · manuscript u-med, CC0) — private test upload

Eroux
7b5ad0a1mo ago

initial commit

Eroux