BDRC
Datasets
All datasets matching “BDRC”tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.tibetan-ocr-benchmark
BDRC Tibetan OCR Benchmark (open subset)
A hand-transcribed benchmark for evaluating Tibetan OCR across writing styles and
technologies. This is an open-access set of 472 page images with ground-truth
transcriptions and per-page metadata (script, technology, legibility).
Companion to the model BDRC/tibetan-ocr
and the leaderboard
(dozens of OCR systems scored on this benchmark).
The transcriptions were produced by Dharmaduta.
The images were selected through detailed research… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-ocr-benchmark.KhyentseWangpo
Dataset Card for KhyentseWangpo
A line-to-text dataset for Tibetan OCR.
Dataset Details
Dataset Description
This dataset consists of 13,527 rows with three columns:
id (string): Unique identifier for each line
image (image): Image file containing a line of Tibetan text
transcription (string): Tibetan text transcription in Unicode format
Curated by: Buddhist Digital Resource Center
Language: Tibetan
License: Open Data Commons Attribution License (ODC-By)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/KhyentseWangpo.tibetan-quotation-detection
Tibetan Quotation Detection Benchmark
254 classical Tibetan books with 23,095 human-reviewed quotation spans, split by book into train / validation / test.
Character offsets are inclusive on both ends: the span is text[start:end+1].
Source
Classical Buddhist commentaries digitized with support from the Tsadra Foundation and OpenPecha, annotated in two batches: an old batch (125 books, Quotation layer) and a new batch (195 books, Citation layer). Both layers mark… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-quotation-detection.8-class-tibetan-page-classification-dataset
8-Class Tibetan Page Classification
Page-level classification of BDRC manuscript / print images into 8 classes: the
six script categories of BDRC/6-class-tibetan-script-classification-dataset
plus two page-type classes — blank and nonplaintext — so a downstream
OCR pipeline can route pages (skip blanks, handle tables/illustrations/scores separately).
Trained classifier: BDRC/8-class-tibetan-page-classifier.
Classes
Class
Description
danyig_pedri
Danyig… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/8-class-tibetan-page-classification-dataset.danyig-pedri-binary-script-classifier
Danyig vs Pedri Binary Script Classification Dataset
Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed.
Images per class
Class
train
val
test
All
Danyig
480
60
60
600
Pedri
480
60
60
600
Total
960
120
120
1,200
Splits
Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.
