datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lagenda_split
LAGENDA Dataset
This is a community mirror of the LAGENDA dataset created by LayerTeam. It has been uploaded here for easier access and integration with the Hugging Face datasets library.
All credit, rights, and accolades belong to the original authors. Please see the citation section below.
Dataset Description
LAGENDA (Large Age and Gender Dataset) is a dataset designed for age and gender recognition tasks. It addresses common biases in existing datasets by ensuring a… See the full description on the dataset page: https://huggingface.co/datasets/uaebn/lagenda_split.minc-2500_split_1
Materials in Context Dataset (MINC-2500)
Dataset Summary
(from the website)
MINC-2500 is a patch classification dataset with 2500 samples per category
(Section 5.4 of the paper). This is a subset of MINC where samples have been
sized to 362 x 362 and each category is sampled evenly. The original resolution
images are not needed as we include the extracted patches in the archive.
dtd_split_1
Dataset Card for Describable Textures Dataset (DTD)
Dataset Summary
Texture classification dataset; consists of 47 categories, 120 images per class.
Data Splits
Equally split into train, val, test; The original paper proposed 10 splits; recent works (BYOL, arxiv:2006.07733) use only first split.
Licensing Information
Not defined at https://www.robots.ox.ac.uk/~vgg/data/dtd/
Citation Information
@InProceedings{cimpoi14describing… See the full description on the dataset page: https://huggingface.co/datasets/mcimpoi/dtd_split_1.doc-split-benchmark
Doc-Split Benchmark
The evaluation slice for page-stream segmentation — the exact set behind the
leaderboard and the cloud-VLM
comparison. Self-contained (page images embedded), with a reference scorer so results are reproducible.
This is the benchmark, not the training corpus (which stays private).
🏆 Leaderboard: doc-split-leaderboard
🎯 Demo: doc-split-demo
🟢 Model: doc-split-mini-e5 (open weights)
🌍 OpenPSS cuts: openpss-mirror (SHORT/LONG, self-contained)… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/doc-split-benchmark.breast-histopathology-images-train-test-valid-split
Breast Histopathology Image dataset
This dataset is just a rearrangement of the Original dataset at Kaggle: https://www.kaggle.com/datasets/paultimothymooney/breast-histopathology-images
Data Citation: https://www.ncbi.nlm.nih.gov/pubmed/27563488 , http://spie.org/Publications/Proceedings/Paper/10.1117/12.2043872
The original dataset has structure: |-- patient_id
|-- class(0 and 1)
The present dataset has following structure: |-- train
|-- class(0 and 1)
|--… See the full description on the dataset page: https://huggingface.co/datasets/EulerianKnight/breast-histopathology-images-train-test-valid-split.
