CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes234 downloads4y agoHugging Face02liy140 /multidomain-measextract-corpus A Multi-Domain Corpus for Measurement Extraction (Seq2Seq variant) A detailed description of corpus creation can be found here. This dataset contains the training and validation and test data for each of the three datasets measeval, bm, and msp. The measeval, and msp datasets were adapted from the MeasEval (Harper et al., 2021) and the Material Synthesis Procedual (Mysore et al., 2019) corpus respectively. This repository aggregates extraction to paragraph-level for msp and… See the full description on the dataset page: https://huggingface.co/datasets/liy140/multidomain-measextract-corpus.texttoken-classification1K<n<10K0 likes167 downloads3y agoHugging Face03Zihao-Li /multidomain_rcot_physicstext100K<n<1M0 likes141 downloads2mo agoHugging Face04dendriteholdings /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes130 downloads21d agoHugging Face05bluecolor777 /Dendrite-Synth-Multi-Domain Dendrite Synth Multi-Domain A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.texttext-generation100K<n<1M0 likes82 downloads13d agoHugging Face06Atmanstr /MultiDomain_Instruction MultiDomain_Instruction A multi-domain instruction dataset designed for instruction tuning and supervised fine-tuning (SFT) of large language models. The dataset contains tasks from multiple domains such as question answering, summarization, reasoning, classification, and general knowledge to improve model generalization. Overview MultiDomain_Instruction is created to support instruction-following training for LLMs. Instead of focusing on a single task, this dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/Atmanstr/MultiDomain_Instruction.texttext-classification100K<n<1M1 likes71 downloads6mo agoHugging Face07TaskPuppyAI /qwen3.8-targeted-multidomain-250 Qwen3.8 Max Multidomain Code Review 250 A 250-record synthetic multilingual code-review dataset generated with Qwen3.8 Max and reviewed/corrected with ChatGPT 5.6 Sol High. All records use the same strict review instruction and ask the model to report only concrete defects supported by the visible code and stated contract. Dataset Summary The publication artifact contains 250 unique records using the schema: { "instruction": "...", "input": "...", "output":… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-multidomain-250.textn<1K0 likes53 downloads17d agoHugging Face08Labradorlabs /bsca-binary-source-gold-v3-multidomain BSCA Gold v3 Multidomain Address-grounded P1 pairs for stripped pseudo-C → source retrieval. Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7 Accepted P1 pairs: 42449 Repositories: 138 Target formats: {"elf": 40733, "pe": 1716} Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874} Internal quality GPA: 3.660; target pass: True Use train.jsonl for fitting, development.jsonl for model selection, and the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.tabularfeature-extraction10K<n<100K0 likes47 downloads1mo agoHugging Face09Pangeanic /Iraqi-Arabic-multidomain-QA-text Iraqi Arabic Multidomain QA Dataset The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models. This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.textquestion-answeringn<1K1 likes25 downloads4mo agoHugging Face10sabin1234 /NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET Nepali Devanagari SFT Dataset — Final Clean Release A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments. Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns. Dataset at a Glance Property Value Total rows 100,000 Total conversation messages 200,000 Human messages 100,000 GPT messages 100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.texttext-generation100K<n<1M0 likes24 downloads1mo agoHugging Face11wxcai /manifold_newest_multi_domains_260318text10K<n<100K0 likes14 downloads2mo agoHugging Face12tppllm /multi-domain-description Multi-Task Description Dataset This dataset contains multiple event sequences from various sources. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper. If you find this dataset useful, we kindly invite you to cite the following papers: @article{liu2024tppllmm, title={TPP-LLM: Modeling Temporal Point Processes by Efficiently Fine-Tuning Large Language Models}, author={Liu, Zefang and Quan, Yinzhu}… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/multi-domain-description.tabular1K<n<10K1 likes11 downloads10mo agoHugging Face13picas9dan /multidomain_2023-11-30_10.20.00text1K<n<10K0 likes9 downloads3y agoHugging Face14Maitreyajayaraj /ai_war_space_multidomain_reasoning_v15textn<1K0 likes9 downloads6mo agoHugging Face15grpathak22 /marathi-maharashtra-multidomain-SFT-1kgated Marathi Maharashtra Multidomain SFT - 1K Sample Dataset Description This is a carefully curated 1,000-sample subset of the comprehensive Marathi-Maharashtra multidomain supervised fine-tuning (SFT) dataset. This high-quality dataset contains question-answer pairs covering diverse aspects of Marathi language, culture, history, and Maharashtra-related topics. Key Features High-Quality Human Verification: All responses have been verified by Marathi language… See the full description on the dataset page: https://huggingface.co/datasets/grpathak22/marathi-maharashtra-multidomain-SFT-1k.textquestion-answeringn<1K0 likes7 downloads8mo agoHugging Face16shkomq /Kurdish_Multi-Domain_Corpus_KMDCgated Kurdish Multi-Domain Corpus (KMDC) Dataset Description The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.textquestion-answering100K<n<1M0 likes5 downloads4mo agoHugging Face17Pratham10 /multi_domain_chatbottext10K<n<100K0 likes3 downloads2y agoHugging Face18Maitreyajayaraj /multidomain_reasoning_precision_v17textn<1K0 likes3 downloads6mo agoHugging Face19Maitreyajayaraj /multi_domain_dialogue_reasoning_v1textn<1K0 likes3 downloads6mo agoHugging Face20Maitreyajayaraj /multi_domain_complex_reasoning_with_constraints_v2textn<1K0 likes2 downloads6mo agoHugging Face21GewHug /Multi-Domain-Evaltabularn<1K1 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.