datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moe-clustering-pilotsdoc-content-clustering-740
Document Content-Clustering Benchmark (12 classes, 740 items)
A small, curated benchmark for clustering documents by their content topic
(not by their visual form/layout). Each item is a single document page provided
as an image plus two text views (a VLM description and OCR markdown), with a
ground-truth content class.
The set is intentionally "tangle-stripped": 27 borderline items whose content
sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.clustering_yamnet_80clustering_vgg19_80cell_clustering_samples
