datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harmful-contents
Harmful-Contents Dataset
A multi-label image dataset for harmful-content classification across eight PEGI-aligned categories.The dataset consists of 5,153 rights-cleared images, split into train/validation/test sets and annotated with both binary labels and mask fields for controlled negative sampling.
Dataset Structure
Harmful-Contents/
csv/
train.csv
val.csv
test.csv
data/
train/*.jpg
val/*.jpg
test/*.jpg
Each CSV contains:
name,
alcohol… See the full description on the dataset page: https://huggingface.co/datasets/onullusoy/harmful-contents.index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.doc-content-clustering-740
Document Content-Clustering Benchmark (12 classes, 740 items)
A small, curated benchmark for clustering documents by their content topic
(not by their visual form/layout). Each item is a single document page provided
as an image plus two text views (a VLM description and OCR markdown), with a
ground-truth content class.
The set is intentionally "tangle-stripped": 27 borderline items whose content
sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.
