datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.tibetan-ocr-benchmark
BDRC Tibetan OCR Benchmark (open subset)
A hand-transcribed benchmark for evaluating Tibetan OCR across writing styles and
technologies. This is an open-access set of 472 page images with ground-truth
transcriptions and per-page metadata (script, technology, legibility).
Companion to the model BDRC/tibetan-ocr
and the leaderboard
(dozens of OCR systems scored on this benchmark).
The transcriptions were produced by Dharmaduta.
The images were selected through detailed research… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-ocr-benchmark.KhyentseWangpo
Dataset Card for KhyentseWangpo
A line-to-text dataset for Tibetan OCR.
Dataset Details
Dataset Description
This dataset consists of 13,527 rows with three columns:
id (string): Unique identifier for each line
image (image): Image file containing a line of Tibetan text
transcription (string): Tibetan text transcription in Unicode format
Curated by: Buddhist Digital Resource Center
Language: Tibetan
License: Open Data Commons Attribution License (ODC-By)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/KhyentseWangpo.tibetan-quotation-detection
Tibetan Quotation Detection Benchmark
254 classical Tibetan books with 23,095 human-reviewed quotation spans, split by book into train / validation / test.
Character offsets are inclusive on both ends: the span is text[start:end+1].
Source
Classical Buddhist commentaries digitized with support from the Tsadra Foundation and OpenPecha, annotated in two batches: an old batch (125 books, Quotation layer) and a new batch (195 books, Citation layer). Both layers mark… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-quotation-detection.8-class-tibetan-page-classification-dataset
8-Class Tibetan Page Classification
Page-level classification of BDRC manuscript / print images into 8 classes: the
six script categories of BDRC/6-class-tibetan-script-classification-dataset
plus two page-type classes — blank and nonplaintext — so a downstream
OCR pipeline can route pages (skip blanks, handle tables/illustrations/scores separately).
Trained classifier: BDRC/8-class-tibetan-page-classifier.
Classes
Class
Description
danyig_pedri
Danyig… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/8-class-tibetan-page-classification-dataset.danyig-pedri-binary-script-classifier
Danyig vs Pedri Binary Script Classification Dataset
Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed.
Images per class
Class
train
val
test
All
Danyig
480
60
60
600
Pedri
480
60
60
600
Total
960
120
120
1,200
Splits
Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.ALL-BDRC-alignments
Tibetan OCR — ALL-BDRC-alignments
79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with
hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come
from the Asian Classics Input Project (ACIP) Sungbum corpus via the
Asian Legacy Library (ALL), normalized to Unicode
and manually matched to BDRC scans. This is the largest clean uchen woodblock set in
the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments.tibetan-spelling-correction-dataset
Tibetan Spelling Correction
Sentence-level spelling correction pairs for Tibetan. Each row pairs an annotator's transcription of a manuscript page segment with the final reviewer's corrected version, from the double-annotation workflow of the BDRC Etext Corpus.
For training and evaluating post-correction models such as TiSpell.
Source batches: Ume 1-4, Uchen 1-4. 4,672 pages.
Contents
train
validation
test
All
error pairs
43,015
2,268
2,229
47,512… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-spelling-correction-dataset.monlamai-transcriptions
Tibetan OCR — MonlamAI transcriptions
3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode
transcriptions, contributed by MonlamAI over BDRC manuscript
scans and aligned page by page.
This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes:
filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.monlamai-handwritten
Tibetan OCR — MonlamAI handwritten (synthetic pages)
11,755 synthetic page images of handwritten Tibetan dbu-med (u-med) text with
line-accurate Unicode transcriptions (line breaks preserved). Each page is
reconstructed from single-line handwritten crops produced during the MonlamAI project (see below) composed in natural reading order into multi-line pages (~6 rows per
page) with margins and inter-line gaps. Because the layout is synthetic, this set is
intended for training… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-handwritten.palri-parkhang-transcriptions
Tibetan OCR — Palri Parkhang
11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly
dbu-med (u-med) manuscripts with a small uchen portion.
The transcriptions were
produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely
Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.tibetan-ocr-diagnostic-benchmark
Tibetan OCR Diagnostic Benchmark (OFAT)
A small, controlled diagnostic OCR benchmark of 300 synthetic Tibetan pecha-page images with exact, noise-free ground truth.
It is a scientific instrument for measuring how OCR character error rate (CER) responds to individual difficulty factors one at a time (OFAT) — not a coverage-maximizing training set.
Code & regeneration: https://github.com/buda-base/synthetic-ocr-benchmark-tools (diagnostic_benchmark/)
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-ocr-diagnostic-benchmark.berkeley-transcriptions
Tibetan OCR — Berkeley
8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly
dbu-med (u-med) manuscripts with a small woodblock (uchen) portion.
These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley.
This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.TiBLAD
TiBLAD — Tibetan Book Layout Analysis Dataset
YOLO-format object-detection dataset for TiBLA (Tibetan Book Layout Analysis).
It contains bounding-box annotations for four layout classes (header, text-area,
footer, footnote) on scanned modern Tibetan book pages, split into training,
validation, and test sets.
Models, code & paper
Primary model: BDRC/TiBLA-RTDETR (RT-DETR-l, AGPL-3.0)
Permissive alternatives: BDRC/TiBLA-PP-DocLayout-L (Apache-2.0)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/TiBLAD.PageToChar-Outline-benchmark
Page-overlap benchmark
This benchmark dataset for detecitng the text boundaries in page where text boundaires overlap while transferring BDRC page-level outlines into character-level outlines over OCR text.
Most outlined texts line up with page boundaries. For those cases the character span is the first character of the start page through the last character of the end page. Some texts do not: one work can end mid-page while another begins on the same page, or several short works… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/PageToChar-Outline-benchmark.tibetan-script-classification-benchmark
Tibetan Script Classification Benchmark
Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training.
All images are BDRC manuscript page scans, balanced by subclass.
Class
Images
Subclasses
Danyig
60
DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6
Druma
60
Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6
Gyuyig
60
Khyuyig: 31, Tsumachug: 15, Yigchung: 14
Pedri
60
Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.LineSegmentationLayoutSegmentation_DatasetKarmapa8ScriptClassificationModernBooksLayout_v1Bo_metadata_benchmark
Bo Metadata Benchmark
This benchmark evaluates how well large language models extract bibliographic metadata from Tibetan texts when given only a segment of the work (Text Head and Text Last), not the full text.
The model is expected to recover fields such as titles, authors, translators, revisors, scribes, revealers, publisher, date, and place from those windows.
The published file is benchmark_dataset.csv (699 rows).
Collections
Collection
Texts
Share… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/Bo_metadata_benchmark.NorbuketakaNumbers
