CoolFace
Datasetpublic

langminer/doc-content-clustering-740

Document Content-Clustering Benchmark (12 classes, 740 items) A small, curated benchmark for clustering documents by their content topic (not by their visual form/layout). Each item is a single document page provided as an image plus two text views (a VLM description and OCR markdown), with a ground-truth content class. The set is intentionally "tangle-stripped": 27 borderline items whose content sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes25downloads
settings

This repository belongs to langminer on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namedoc-content-clustering-740
visibilitypublic
licenceother
gatedno
ownerlangminer
Account settings
langminer/doc-content-clustering-740 · CoolFace