CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BDRC /tibetan-page-orientation-classifier-dataset Tibetan Page Orientation Dataset Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen. Dataset composition Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations. Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family). Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.imageimage-classification1K<n<10K0 likes310 downloads3mo agoHugging Face02BDRC /tibetan-ocr-benchmark BDRC Tibetan OCR Benchmark (open subset) A hand-transcribed benchmark for evaluating Tibetan OCR across writing styles and technologies. This is an open-access set of 472 page images with ground-truth transcriptions and per-page metadata (script, technology, legibility). Companion to the model BDRC/tibetan-ocr and the leaderboard (dozens of OCR systems scored on this benchmark). The transcriptions were produced by Dharmaduta. The images were selected through detailed research… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-ocr-benchmark.imageimage-to-textn<1K0 likes193 downloads28d agoHugging Face03BDRC /KhyentseWangpo Dataset Card for KhyentseWangpo A line-to-text dataset for Tibetan OCR. Dataset Details Dataset Description This dataset consists of 13,527 rows with three columns: id (string): Unique identifier for each line image (image): Image file containing a line of Tibetan text transcription (string): Tibetan text transcription in Unicode format Curated by: Buddhist Digital Resource Center Language: Tibetan License: Open Data Commons Attribution License (ODC-By)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/KhyentseWangpo.imageimage-to-text10K<n<100K0 likes133 downloads1y agoHugging Face04BDRC /tibetan-quotation-detection Tibetan Quotation Detection Benchmark 254 classical Tibetan books with 23,095 human-reviewed quotation spans, split by book into train / validation / test. Character offsets are inclusive on both ends: the span is text[start:end+1]. Source Classical Buddhist commentaries digitized with support from the Tsadra Foundation and OpenPecha, annotated in two batches: an old batch (125 books, Quotation layer) and a new batch (195 books, Citation layer). Both layers mark… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-quotation-detection.tabulartoken-classificationn<1K0 likes102 downloads11d agoHugging Face05BDRC /8-class-tibetan-page-classification-dataset 8-Class Tibetan Page Classification Page-level classification of BDRC manuscript / print images into 8 classes: the six script categories of BDRC/6-class-tibetan-script-classification-dataset plus two page-type classes — blank and nonplaintext — so a downstream OCR pipeline can route pages (skip blanks, handle tables/illustrations/scores separately). Trained classifier: BDRC/8-class-tibetan-page-classifier. Classes Class Description danyig_pedri Danyig… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/8-class-tibetan-page-classification-dataset.imageimage-classification1K<n<10K0 likes84 downloads3mo agoHugging Face06BDRC /danyig-pedri-binary-script-classifier Danyig vs Pedri Binary Script Classification Dataset Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed. Images per class Class train val test All Danyig 480 60 60 600 Pedri 480 60 60 600 Total 960 120 120 1,200 Splits Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.imageimage-classification1K<n<10K0 likes79 downloads3mo agoHugging Face07BDRC /ALL-BDRC-alignments Tibetan OCR — ALL-BDRC-alignments 79,572 page images of Tibetan woodblock prints (uchen) aligned page-by-page with hand-verified Unicode transcriptions, line breaks preserved. Transcriptions come from the Asian Classics Input Project (ACIP) Sungbum corpus via the Asian Legacy Library (ALL), normalized to Unicode and manually matched to BDRC scans. This is the largest clean uchen woodblock set in the BDRC Tibetan OCR release — released jointly by the Asian Legacy Library (ALL)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/ALL-BDRC-alignments.imageimage-to-text10K<n<100K1 likes74 downloads21d agoHugging Face08BDRC /tibetan-spelling-correction-dataset Tibetan Spelling Correction Sentence-level spelling correction pairs for Tibetan. Each row pairs an annotator's transcription of a manuscript page segment with the final reviewer's corrected version, from the double-annotation workflow of the BDRC Etext Corpus. For training and evaluating post-correction models such as TiSpell. Source batches: Ume 1-4, Uchen 1-4. 4,672 pages. Contents train validation test All error pairs 43,015 2,268 2,229 47,512… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-spelling-correction-dataset.texttext-generation100K<n<1M0 likes73 downloads1mo agoHugging Face09BDRC /monlamai-transcriptions Tibetan OCR — MonlamAI transcriptions 3,072 page images of Tibetan dbu-med (u-med) manuscripts with page-level Unicode transcriptions, contributed by MonlamAI over BDRC manuscript scans and aligned page by page. This is the full page equivalent of the line-segmented datasets available on openpecha/OCR-Betsug and openpecha/OCR-Drutsa, with the following changes: filter out cases where not all line in a page are transcribed (so not all the lines in the original datasets are… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-transcriptions.imageimage-to-text1K<n<10K0 likes57 downloads1mo agoHugging Face10BDRC /monlamai-handwritten Tibetan OCR — MonlamAI handwritten (synthetic pages) 11,755 synthetic page images of handwritten Tibetan dbu-med (u-med) text with line-accurate Unicode transcriptions (line breaks preserved). Each page is reconstructed from single-line handwritten crops produced during the MonlamAI project (see below) composed in natural reading order into multi-line pages (~6 rows per page) with margins and inter-line gaps. Because the layout is synthetic, this set is intended for training… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/monlamai-handwritten.imageimage-to-text10K<n<100K0 likes57 downloads1mo agoHugging Face11BDRC /palri-parkhang-transcriptions Tibetan OCR — Palri Parkhang 11,133 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small uchen portion. The transcriptions were produced by Palri Parkhang, an input project led by Chris Tomlinson (former BDRC's CTO) in Nepal in 2006-2013 and aligned to BDRC scans; the material is largely Nyingma collected works and gter ma cycles. The transcriptions prioritized legibility over fidelity to the enscribed text and… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/palri-parkhang-transcriptions.imageimage-to-text10K<n<100K0 likes54 downloads1mo agoHugging Face12BDRC /tibetan-ocr-diagnostic-benchmark Tibetan OCR Diagnostic Benchmark (OFAT) A small, controlled diagnostic OCR benchmark of 300 synthetic Tibetan pecha-page images with exact, noise-free ground truth. It is a scientific instrument for measuring how OCR character error rate (CER) responds to individual difficulty factors one at a time (OFAT) — not a coverage-maximizing training set. Code & regeneration: https://github.com/buda-base/synthetic-ocr-benchmark-tools (diagnostic_benchmark/) What's inside… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-ocr-diagnostic-benchmark.imageimage-to-textn<1K0 likes52 downloads2mo agoHugging Face13BDRC /berkeley-transcriptions Tibetan OCR — Berkeley 8,866 page images of Tibetan text with page-level Unicode transcriptions, mostly dbu-med (u-med) manuscripts with a small woodblock (uchen) portion. These works were transcribed by Geshe Dangsong Namgyal, from 2016 to 2026, in his role as a Data System Analyst at University of California, Berkeley. This work was initiated by Prof. Kurt Keutzer with the longstanding hope of producing training data for Handwritten Text Recognition for dbu med and OCR of… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/berkeley-transcriptions.imageimage-to-text1K<n<10K0 likes52 downloads1mo agoHugging Face14BDRC /TiBLADgated TiBLAD — Tibetan Book Layout Analysis Dataset YOLO-format object-detection dataset for TiBLA (Tibetan Book Layout Analysis). It contains bounding-box annotations for four layout classes (header, text-area, footer, footnote) on scanned modern Tibetan book pages, split into training, validation, and test sets. Models, code & paper Primary model: BDRC/TiBLA-RTDETR (RT-DETR-l, AGPL-3.0) Permissive alternatives: BDRC/TiBLA-PP-DocLayout-L (Apache-2.0)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/TiBLAD.imageobject-detection1K<n<10K0 likes51 downloads22d agoHugging Face15BDRC /PageToChar-Outline-benchmark Page-overlap benchmark This benchmark dataset for detecitng the text boundaries in page where text boundaires overlap while transferring BDRC page-level outlines into character-level outlines over OCR text. Most outlined texts line up with page boundaries. For those cases the character span is the first character of the start page through the last character of the end page. Some texts do not: one work can end mid-page while another begins on the same page, or several short works… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/PageToChar-Outline-benchmark.0 likes43 downloads7d agoHugging Face16BDRC /tibetan-script-classification-benchmark Tibetan Script Classification Benchmark Holdout benchmark for 6-class Tibetan script classification. Test split only — not used during training. All images are BDRC manuscript page scans, balanced by subclass. Class Images Subclasses Danyig 60 DraDring: 25, DraRing: 9, Drathung: 17, Gongshabma: 3, Tsegdrig: 6 Druma 60 Dhumri: 22, DruDring: 20, DruRing: 10, Druchen: 2, Druthung: 6 Gyuyig 60 Khyuyig: 31, Tsumachug: 15, Yigchung: 14 Pedri 60 Peri: 44, Petsuk: 16… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-script-classification-benchmark.imageimage-classificationn<1K0 likes31 downloads3mo agoHugging Face17BDRC /LineSegmentationimage0 likes29 downloads2y agoHugging Face18BDRC /LayoutSegmentation_Datasetimage0 likes21 downloads2y agoHugging Face19BDRC /Karmapa8image10K<n<100K1 likes17 downloads2y agoHugging Face20BDRC /ScriptClassificationimage10K<n<100K0 likes14 downloads2y agoHugging Face21BDRC /ModernBooksLayout_v1image0 likes14 downloads8mo agoHugging Face22BDRC /Bo_metadata_benchmark Bo Metadata Benchmark This benchmark evaluates how well large language models extract bibliographic metadata from Tibetan texts when given only a segment of the work (Text Head and Text Last), not the full text. The model is expected to recover fields such as titles, authors, translators, revisors, scribes, revealers, publisher, date, and place from those windows. The published file is benchmark_dataset.csv (699 rows). Collections Collection Texts Share… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/Bo_metadata_benchmark.textn<1K0 likes14 downloads1mo agoHugging Face23BDRC /NorbuketakaNumbersimage0 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.