datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.multidomain-measextract-corpus
A Multi-Domain Corpus for Measurement Extraction (Seq2Seq variant)
A detailed description of corpus creation can be found here.
This dataset contains the training and validation and test data for each of the three datasets measeval, bm, and msp. The measeval, and msp datasets were adapted from the MeasEval (Harper et al., 2021) and the Material Synthesis Procedual (Mysore et al., 2019) corpus respectively.
This repository aggregates extraction to paragraph-level for msp and… See the full description on the dataset page: https://huggingface.co/datasets/liy140/multidomain-measextract-corpus.multidomain_rcot_physicsDendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.MultiDomain_Instruction
MultiDomain_Instruction
A multi-domain instruction dataset designed for instruction tuning and supervised fine-tuning (SFT) of large language models. The dataset contains tasks from multiple domains such as question answering, summarization, reasoning, classification, and general knowledge to improve model generalization.
Overview
MultiDomain_Instruction is created to support instruction-following training for LLMs. Instead of focusing on a single task, this dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/Atmanstr/MultiDomain_Instruction.qwen3.8-targeted-multidomain-250
Qwen3.8 Max Multidomain Code Review 250
A 250-record synthetic multilingual code-review dataset generated with
Qwen3.8 Max and reviewed/corrected with ChatGPT 5.6 Sol High.
All records use the same strict review instruction and ask the model to report
only concrete defects supported by the visible code and stated contract.
Dataset Summary
The publication artifact contains 250 unique records using the schema:
{
"instruction": "...",
"input": "...",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-multidomain-250.bsca-binary-source-gold-v3-multidomain
BSCA Gold v3 Multidomain
Address-grounded P1 pairs for stripped pseudo-C → source retrieval.
Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7
Accepted P1 pairs: 42449
Repositories: 138
Target formats: {"elf": 40733, "pe": 1716}
Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874}
Internal quality GPA: 3.660; target pass: True
Use train.jsonl for fitting, development.jsonl for model selection, and
the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.Iraqi-Arabic-multidomain-QA-text
Iraqi Arabic Multidomain QA Dataset
The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models.
This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET
Nepali Devanagari SFT Dataset — Final Clean Release
A 100,000-row synthetic Nepali SFT dataset designed for Nepali-language instruction-following and supervised fine-tuning experiments.
Release status: Final structural and Unicode validation passed for the previously identified contamination/corruption patterns.
Dataset at a Glance
Property
Value
Total rows
100,000
Total conversation messages
200,000
Human messages
100,000
GPT messages
100,000… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NEPALI-MCQ-SFT-MULTIDOMAIN-DATASET.manifold_newest_multi_domains_260318multi-domain-description
Multi-Task Description Dataset
This dataset contains multiple event sequences from various sources. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the following papers:
@article{liu2024tppllmm,
title={TPP-LLM: Modeling Temporal Point Processes by Efficiently Fine-Tuning Large Language Models},
author={Liu, Zefang and Quan, Yinzhu}… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/multi-domain-description.multidomain_2023-11-30_10.20.00ai_war_space_multidomain_reasoning_v15marathi-maharashtra-multidomain-SFT-1k
Marathi Maharashtra Multidomain SFT - 1K Sample
Dataset Description
This is a carefully curated 1,000-sample subset of the comprehensive Marathi-Maharashtra multidomain supervised fine-tuning (SFT) dataset. This high-quality dataset contains question-answer pairs covering diverse aspects of Marathi language, culture, history, and Maharashtra-related topics.
Key Features
High-Quality Human Verification: All responses have been verified by Marathi language… See the full description on the dataset page: https://huggingface.co/datasets/grpathak22/marathi-maharashtra-multidomain-SFT-1k.Kurdish_Multi-Domain_Corpus_KMDC
Kurdish Multi-Domain Corpus (KMDC)
Dataset Description
The Kurdish Multi-Domain Corpus (KMDC) is a large-scale instruction-style dataset designed to support natural language processing (NLP), supervised fine-tuning (SFT), and large language model (LLM) development for Central Kurdish (Sorani). The dataset consists of structured question–response pairs generated through an LLM-guided pipeline that transforms raw Kurdish text into machine-learning-ready… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/Kurdish_Multi-Domain_Corpus_KMDC.multi_domain_chatbotmultidomain_reasoning_precision_v17multi_domain_dialogue_reasoning_v1multi_domain_complex_reasoning_with_constraints_v2Multi-Domain-Eval
