CoolFace
Datasetpublic

langminer/doc-content-clustering-740

Document Content-Clustering Benchmark (12 classes, 740 items) A small, curated benchmark for clustering documents by their content topic (not by their visual form/layout). Each item is a single document page provided as an image plus two text views (a VLM description and OCR markdown), with a ground-truth content class. The set is intentionally "tangle-stripped": 27 borderline items whose content sits ambiguously between two classes were removed from a larger 907-item pool to… See the full description on the dataset page: https://huggingface.co/datasets/langminer/doc-content-clustering-740.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes25downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
langminer/doc-content-clustering-740 · CoolFace