datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VDR_MEGA_MultiDomain_DocRetrieval
Visual Document Retrieval Dataset
Overview
This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks.
Dataset Structure
The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.multidomain-kazakh-dataset
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 |
☕ Buy Me a Coffee🦄 |
⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/multidomain-kazakh-dataset.mmlumultidomain-kazakh-dataset
Dataset Description
Point of Contact: Sanzhar Murzakhmetov, Besultan Sagyndyk
Dataset Summary
MDBKD | Multi-Domain Bilingual Kazakh Dataset is a Kazakh-language dataset containing just over 24 883 808 unique texts from multiple domains.
Supported Tasks
'MLM/CLM': can be used to train a model for casual and masked languange modeling
Languages
The kk code for Kazakh as generally spoken in the Kazakhstan
Data Instances
For each instance… See the full description on the dataset page: https://huggingface.co/datasets/kz-transformers/multidomain-kazakh-dataset.commonsense_qamulti-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.multi_domain_ai_human_text
multi_domain_ai_human_text — Datasheet
Balanced, multi-domain AI-vs-human text detection benchmark with dedicated
out-of-distribution and adversarial evaluation panels. Built by
scripts/build_paper_dataset.py from an 11-corpus unified aggregation.
Splits
Split
AI
Human
Total
Purpose
train
300,000
300,000
600,000
training (balanced, English, clean)
validation
2,996
2,999
5,995
model selection
test
4,991
4,999
9,990
in-distribution test… See the full description on the dataset page: https://huggingface.co/datasets/acmc/multi_domain_ai_human_text.multidomain-measextract-corpus
A Multi-Domain Corpus for Measurement Extraction (Seq2Seq variant)
A detailed description of corpus creation can be found here.
This dataset contains the training and validation and test data for each of the three datasets measeval, bm, and msp. The measeval, and msp datasets were adapted from the MeasEval (Harper et al., 2021) and the Material Synthesis Procedual (Mysore et al., 2019) corpus respectively.
This repository aggregates extraction to paragraph-level for msp and… See the full description on the dataset page: https://huggingface.co/datasets/liy140/multidomain-measextract-corpus.NSFW-MultiDomain-Classification
NSFW_MultiDomain
The NSFW_MultiDomain dataset is a curated image classification dataset focused on multi-domain adult content recognition. It consists of 5 distinct categories aimed at facilitating the development of robust NSFW (Not Safe For Work) image classification models. This dataset enables training and benchmarking of models that can distinguish between subtle variations in explicit and non-explicit content across artistic, animated, and real-world imagery.… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/NSFW-MultiDomain-Classification.multi-domain-cloudflare-observability
Multi-Domain Cloudflare Web Traffic, Performance and Security Observability Dataset
This dataset contains multi-domain Cloudflare analytics exported into analysis-ready Parquet tables. It combines HTTP request aggregates, hourly traffic trends, path and referrer dimensions, country/device/browser breakdowns, DNS analytics, cache behavior, Web Vitals/RUM performance signals, redirect and ruleset metadata, and available firewall/security aggregates across 200 websites.
The… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/multi-domain-cloudflare-observability.multidomain_rcot_physicsflair-multidomain-parquet
FLAIR multi-domain pack — 256 px
24,475 tiles · 9 French domains · 1.94 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: the headline runs; 256 px gives the sharpest boundaries.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes ← this repo
flair-multidomain-parquet-128
128 px
0.9 GB
yes
Columns
column
type
contents
rgb
JPEG bytes
aerial RGB, 256², quality… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet.Urdu-Multi-Domain-Benchmark
Urdu Multi-Domain Datasets
33 labeled Urdu datasets (288,899 examples) for text classification in Nastaliq (Perso-Arabic) and Roman Urdu (Latin). Each domain is a separate Hub subset so you can download one task at a time.
Authors: Muhammad Abdullah Haroon and Maryam Bashir, FAST-NUCES, Lahore.
Companion paper: Domain Robustness of Multilingual NLP Models Across Urdu and Roman Urdu Scripts.
Permanent archive: Zenodo DOI 10.5281/zenodo.22195610.
How to load
Pick a… See the full description on the dataset page: https://huggingface.co/datasets/abdullaharoon/Urdu-Multi-Domain-Benchmark.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/dendriteholdings/Dendrite-Synth-Multi-Domain.flair-multidomain-parquet-512
FLAIR multi-domain pack — 512 px (0.20 m/px)
24,475 tiles · 9 French domains · 4.86 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
This is FLAIR's native aerial resolution. A 102.4 m tile at 512 px is 0.20 m/px — the sampling the COSIA labels were digitised at. The smaller packs below are downsamples, and results from different sampling must not share a table: changing ground sampling changes the task, not just the cost.
Pack
Patch
Ground… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-512.flair-multidomain-parquet-128
FLAIR multi-domain pack — 128 px
24,475 tiles · 9 French domains · 0.9 GB, repacked from IGNF/FLAIR-HUB into Parquet so a Colab runtime can stream it.
Use this pack for: modest GPUs. Same 9 domains and same labels, a quarter of the pixels per training step - the practical choice on a free Colab T4.
Pack
Patch
Size
One shard per domain
flair-multidomain-parquet
256 px
1.94 GB
yes
flair-multidomain-parquet-128
128 px
0.9 GB
yes ← this repo
Columns… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/flair-multidomain-parquet-128.nllb-multi-domainNLLB Multi Domain is a set of professionally-translated sentences in News, Unscripted informal speech, and Health domains. It is designed to enable assessment of out-of-domain performance and to study domain adaptation for machine translation. Each domain has approximately 3000 sentences.Dendrite-Synth-Multi-Domain
Dendrite Synth Multi-Domain
A verified, difficulty-filtered, style-amplified synthetic corpus of question / reasoning / answer
triples spanning the 14 MMLU-Pro categories - mathematics, computer science, natural sciences, chemistry, physics, engineering, health, law, business, economics, psychology, philosophy, history and others (expanded to 122 fine-grained categories and 695 subcategories). Problems are written by a pool of generator models, solved with explicit reasoning by… See the full description on the dataset page: https://huggingface.co/datasets/bluecolor777/Dendrite-Synth-Multi-Domain.agieval_eval_lsat_rcMultiDomain_Instruction
MultiDomain_Instruction
A multi-domain instruction dataset designed for instruction tuning and supervised fine-tuning (SFT) of large language models. The dataset contains tasks from multiple domains such as question answering, summarization, reasoning, classification, and general knowledge to improve model generalization.
Overview
MultiDomain_Instruction is created to support instruction-following training for LLMs. Instead of focusing on a single task, this dataset combines… See the full description on the dataset page: https://huggingface.co/datasets/Atmanstr/MultiDomain_Instruction.t2i-nla-multidomain-1m-training-data-public
T2I-NLA Multidomain 1M Training Data
Public backup of final T2I-NLA reader training data.
Balanced multidomain 1M training data for general FLUX AR/AV readers.
This repository is intended to make final-reader retraining possible after the
original server is no longer available.
See manifest.json for the original local path, size, and associated reader models.
MultiDomain-QADialog
📚 MultiDomain-QADialog Dataset
This repository contains the processed, multi-source dataset used to train the SHARE Model for dialogue inference. The dataset combines three prominent resources in the dialogue space:
MediaSum – dialogues from broadcast transcripts (300k samples)
SAMSum – messenger-style casual conversations (16K samples)
SODA – million-scale, high-quality dialogue dataset (~1M samples)
All datasets have been harmonized into a unified format and stored in sharded… See the full description on the dataset page: https://huggingface.co/datasets/JustinDuc/MultiDomain-QADialog.multidomain-VQA-with-cot-trace-9KThis dataset of 9K samples has been created with AOKVQA Train & Val split, TDIUC Val Split (Quantitative and Physical Reasoning Questions only). This is a multidomain dataset solely created to test the multidomain reasoning knowledge of VLM's, it can be used for inference or rapid prototyping. The synthetic COT trace has been generated using Qwen-VL-32B Model. This is for educational and research purposes only. All the copyright belongs to the original owners of the datasets.
Multi-Domain-Reasoning-Benchmark
Comprehensive Multi-Domain Reasoning Benchmark (CMDR-Bench)
A systematic evaluation suite comprising 100 meticulously curated test cases across 10 distinct cognitive domains, designed to assess Large Language Models' capabilities in reasoning, problem-solving, and instruction-following. Each domain features a graduated difficulty scale (Levels 1–10), enabling fine-grained analysis of capability thresholds from elementary to expert-level complexity.
qwen3.8-targeted-multidomain-250
Qwen3.8 Max Multidomain Code Review 250
A 250-record synthetic multilingual code-review dataset generated with
Qwen3.8 Max and reviewed/corrected with ChatGPT 5.6 Sol High.
All records use the same strict review instruction and ask the model to report
only concrete defects supported by the visible code and stated contract.
Dataset Summary
The publication artifact contains 250 unique records using the schema:
{
"instruction": "...",
"input": "...",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-targeted-multidomain-250.test_multidomain
Dataset Card for "test_multidomain"
More Information needed
bsca-binary-source-gold-v3-multidomain
BSCA Gold v3 Multidomain
Address-grounded P1 pairs for stripped pseudo-C → source retrieval.
Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7
Accepted P1 pairs: 42449
Repositories: 138
Target formats: {"elf": 40733, "pe": 1716}
Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874}
Internal quality GPA: 3.660; target pass: True
Use train.jsonl for fitting, development.jsonl for model selection, and
the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.qwen-expert-multidomain-chat-v1Multi-Domain-Reasoning-SFT
Multi-Domain-Reasoning-SFT
Dataset Summary
The Multi-Domain-Reasoning-SFT dataset by EnDevSols is a large-scale, high-quality Supervised Fine-Tuning (SFT) dataset designed to train large language models in deep reasoning, technical analysis, and complex problem-solving.
Consisting of nearly 580,000 meticulously structured examples, this dataset is specifically engineered to teach models how to "think" before they answer. It separates the internal cognitive process… See the full description on the dataset page: https://huggingface.co/datasets/EnDevSols/Multi-Domain-Reasoning-SFT.gemini-reasoning-traces-multidomain
Gemini-Reasoning-Traces-Multidomain
A curated dataset of 2,282 samples with high-quality, reasoning traces distilled from Gemini-2.5-Pro across 8 diverse domains. Every sample (question + reasoning + answer) fits within 4,096 Gemma-3 tokens, making it ideal for fine-tuning small language models on constrained hardware.
Why This Dataset?
High-quality reasoning datasets are scarce — especially for domains requiring structured thought such as creative writing, summarization… See the full description on the dataset page: https://huggingface.co/datasets/rohit94/gemini-reasoning-traces-multidomain.
