datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.Afrivoice_Kinyarwanda_Image_Domain_classification
Dataset Description
This dataset is a restructured version of Afrivoice Kinyarwanda, reorganized for image domain classification. The original audio-and-image manifest data was regrouped into a standard Hugging Face imagefolder layout (train/validation/test splits, one subfolder per class) so it can be loaded directly with datasets.load_dataset("imagefolder", ...) for training image classifiers.
No new images were collected and no image content was modified beyond format… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/Afrivoice_Kinyarwanda_Image_Domain_classification.synthetic-domain-text-classification
Dataset Card for my-distiset-b845cf19
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/my-distiset-b845cf19/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-domain-text-classification.user_prompt_domain_classification-500000x500,000 users prompts classified into domain. Classification performed by openai/gpt-oss-120b with reasoning set to medium and temperature=0, top_p=1.
Prompts sourced and randomized from various repos including:
Roman1111111/coding-prompts
kth8/user-prompts-1M
wop/just-user-prompts
trl-lib/DeepMath-103K
ianncity/General-Distillation-Prompts-1M
ianncity/VIBE-Prompts-500000x
ianncity/science-prompts-100k
m-a-p/SuperGPQA
Total completion tokens: 70 million
mtop_domain_intent_fr_prompt_intent_classification
mtop_domain_intent_fr_prompt_intent_classification
Summary
mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.domain_classificationhc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
alpaca-bangla_domain_classification
Dataset Card for alpaca-bangla_domain_classification
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/chrononeel/alpaca-bangla_domain_classification.query-domain-classification-sharegpttariff_trade_domain.synthetic_tariff_classification_products_kr
Synthetic Tariff Classification Products (Korean)
A synthetic dataset of Korean product descriptions paired with their HS (Harmonized System) tariff codes and classification rationales, generated to support training and evaluation of customs classification models. All entries are LLM-generated; no real personal or confidential customs records are included.
Dataset Description
Korean customs classification requires assigning a 10-digit HS code to an imported or… See the full description on the dataset page: https://huggingface.co/datasets/lablup/tariff_trade_domain.synthetic_tariff_classification_products_kr.Query_Domain_Classificationdfm_domain_classificationm196k-dedup-decon-filter_easy-r1-filter_wrong-decon_eval-domain-classification
Dataset card for m196k-dedup-decon-filter_easy-r1-filter_wrong-decon_eval-tokenized-120325-domain-classification
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"prompt": "A mother brings her 3-week-old infant to the pediatrician's office because she is concerned about his feeding habits. He was born without complications and has not had any medical problems up until this time. However, for the past 4 days, he has been fussy, is… See the full description on the dataset page: https://huggingface.co/datasets/mmqm/m196k-dedup-decon-filter_easy-r1-filter_wrong-decon_eval-domain-classification.query-domain-classification-sharegpt-v2Top_level_domain_classification_datasetsynthetic-domain-text-classification
Dataset Card for synthetic-domain-text-classification
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/palashsharma15/synthetic-domain-text-classification/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/palashsharma15/synthetic-domain-text-classification.flan_combined_task198_mnli_domain_classificationrag_domain_query_classification
