tibetan
Datasets
All datasets matching “tibetan”tibetan-metadata-extracted
Tibetan Metadata Extracted Documents
Raw BDRC outliner exports used to build ganga4364/tibetan-metadata-detector.
Contents
3,794 approved documents with annotated title/author spans. Each row:
Field
Description
doc_id
Document UUID
filename
Source filename
text
Full document text (UTF-8)
annotations_json
JSON with segments and flat annotations (title/author spans)
Related repos
Window splits + model training data:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-metadata-extracted.tibetan_monolingual_A_merged_123_linesTibetan-0310tibetan_monolingual_Atibetan_monolingual_A_merged_135_linestibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.
