datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.NES-plankton-classifier-2022-dataset
NES Plankton Classifier 2022 Training Data
Image classification dataset of plankton species imaged by Imaging FlowCytobot
(IFCB) on the Northeast US Shelf. Used to train the NES Plankton Classifier 2022.
Dataset Summary
155 classes of plankton and non-plankton ROIs (including detritus, bubbles, fibers, etc.)
97,026 labeled images across train and validation splits
Images are grayscale PNGs extracted from IFCB sample files
Annotations are human-verified using the… See the full description on the dataset page: https://huggingface.co/datasets/sosiklab/NES-plankton-classifier-2022-dataset.danyig-pedri-binary-script-classifier
Danyig vs Pedri Binary Script Classification Dataset
Stage-2 binary classifier for distinguishing Danyig (5 subscripts: DraDring, DraRing, Drathung, Gongshabma, Tsegdrig) from Pedri (2 subscripts: Peri, Petsuk). Real-only, all images human-reviewed.
Images per class
Class
train
val
test
All
Danyig
480
60
60
600
Pedri
480
60
60
600
Total
960
120
120
1,200
Splits
Manuscript-stratified split — each manuscript work appears in exactly… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/danyig-pedri-binary-script-classifier.siglip-doc-understanding-classifier
SigLIP Doc Understanding — Unanswerable Question Detection Dataset
A mixed answerable / unanswerable benchmark dataset built from DocVQA and MP-DocVQA, used to
train and evaluate the siglip-doc-understanding-classifier
unanswerable-question detector.
Each row pairs a document image with a question. Half of the questions are the original,
answerable DocVQA/MP-DocVQA questions; the other half are corrupted versions of those same
questions — modified so the document image no longer… See the full description on the dataset page: https://huggingface.co/datasets/giacolees/siglip-doc-understanding-classifier.
