CoolFace
Datasetpublic

agentlans/en-document-classification

English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.

sourceHugging Faceodc-byupdated 25d agoView on Hugging Face
1likes373downloads
all.jsonl.zst4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:dc9ffbeb2c4b30694785112f8f8bd325f2a86ecd6c36aa9ad6e57c7daf9635f23size 6797345094