CoolFace
19 results

domain-classification

asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes234 downloads4y agoHugging FaceKira-Floris /Afrivoice_Kinyarwanda_Image_Domain_classification Dataset Description This dataset is a restructured version of Afrivoice Kinyarwanda, reorganized for image domain classification. The original audio-and-image manifest data was regrouped into a standard Hugging Face imagefolder layout (train/validation/test splits, one subfolder per class) so it can be loaded directly with datasets.load_dataset("imagefolder", ...) for training image classifiers. No new images were collected and no image content was modified beyond format… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/Afrivoice_Kinyarwanda_Image_Domain_classification.imageimage-classification100K<n<1M1 likes101 downloads14d agoHugging Faceargilla /synthetic-domain-text-classification Dataset Card for my-distiset-b845cf19 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davidberenstein1957/my-distiset-b845cf19/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/argilla/synthetic-domain-text-classification.texttext-classification1K<n<10K7 likes61 downloads2y agoHugging Facekth8 /user_prompt_domain_classification-500000x500,000 users prompts classified into domain. Classification performed by openai/gpt-oss-120b with reasoning set to medium and temperature=0, top_p=1. Prompts sourced and randomized from various repos including: Roman1111111/coding-prompts kth8/user-prompts-1M wop/just-user-prompts trl-lib/DeepMath-103K ianncity/General-Distillation-Prompts-1M ianncity/VIBE-Prompts-500000x ianncity/science-prompts-100k m-a-p/SuperGPQA Total completion tokens: 70 million texttext-classification100K<n<1M0 likes41 downloads6mo agoHugging FaceCATIE-AQ /mtop_domain_intent_fr_prompt_intent_classification mtop_domain_intent_fr_prompt_intent_classification Summary mtop_domain_intent_fr_prompt_intent_classification is a subset of the Dataset of French Prompts (DFP).It contains 497,100 rows that can be used for an intent text classification task.The original data (without prompts) comes from the dataset mtop_domain Haoran Li et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/mtop_domain_intent_fr_prompt_intent_classification.texttext-classification100K<n<1M0 likes34 downloads1y agoHugging Faceohgnues /domain_classificationtext10K<n<100K0 likes34 downloads2y agoHugging Face