CoolFace
Datasetpublic

agentlans/en-document-classification

English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.

sourceHugging Faceodc-byupdated 22d agoView on Hugging Face
1likes420downloads

agentlans/en-document-classification · main · files are served by the source, never re-hosted here