document-classification
en-document-classification
English Document Classification Dataset
This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora.
Dataset Summary
The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.en-document-format-classification
English Document Format Classification Dataset
English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats.
Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks.
Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage).
Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.rvl-cdip-document-classification
rvl-cdip-document-classification
This dataset is created from original aharley/rvl_cdip dataset using this notebook
Dataset Summary
This dataset consists of 8992 grayscale images in 16 classes, with 562 images per class.
There are 8000 training images(500 image per class) and 992 test images(62 images per class).
The images are sized so their largest dimension does not exceed 1000 pixels.
multilingual-document-classification
Multilingual Document Classification Dataset
This dataset contains 100,000 text passages across 100 non-English language-script pairs sourced from the agentlans/HuggingFaceFW-finetranslations-100-languages-sample collection.
Each original text passage is paired with its English translation and has been programmatically annotated with domain, writing genre, and educational classifications to facilitate cross-lingual classification and domain adaptation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-document-classification.
