domain-classification
hc3-wiki-domain-classification-robertabert-base-multilingual-cased-edda-domain-classificationESM_CATIE-AQ__mtop_domain_intent_fr_prompt_intent_classification_defaultamazon_domain_to_classification_testdomain_classificationMistral-7B-v0.1-DomainClassification-Negative-seed-42-2025-12-01fashion_classification_domain_adaptation__resumingMistral-7B-v0.1-DomainClassification-All-seed-42-2025-12-01
multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.Afrivoice_Kinyarwanda_Image_Domain_classification
Dataset Description
This dataset is a restructured version of Afrivoice Kinyarwanda, reorganized for image domain classification. The original audio-and-image manifest data was regrouped into a standard Hugging Face imagefolder layout (train/validation/test splits, one subfolder per class) so it can be loaded directly with datasets.load_dataset("imagefolder", ...) for training image classifiers.
No new images were collected and no image content was modified beyond format… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/Afrivoice_Kinyarwanda_Image_Domain_classification.synthetic-domain-text-classification
Dataset Card for my-distiset-b845cf19
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/my-distiset-b845cf19/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-domain-text-classification.user_prompt_domain_classification-500000x500,000 users prompts classified into domain. Classification performed by openai/gpt-oss-120b with reasoning set to medium and temperature=0, top_p=1.
Prompts sourced and randomized from various repos including:
Roman1111111/coding-prompts
kth8/user-prompts-1M
wop/just-user-prompts
trl-lib/DeepMath-103K
ianncity/General-Distillation-Prompts-1M
ianncity/VIBE-Prompts-500000x
ianncity/science-prompts-100k
m-a-p/SuperGPQA
Total completion tokens: 70 million
mtop_domain_intent_fr_prompt_intent_classification
mtop_domain_intent_fr_prompt_intent_classification
Summary
mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.domain_classification
