datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cambrian-10M
Cambrian-10M Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-10M is a comprehensive dataset designed for instruction tuning, particularly in multimodal settings involving visual interaction data. The dataset is crafted to address the scarcity of high-quality multimodal instruction-tuning data and to maintain the language abilities of multimodal large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-10M.AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.LongAlign-10k
LongAlign-10k
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper]
LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.EMMOE-100
EMMOE-100 Trainset
Resources
Project
Paper
Code
Model
Dataset
Dataset Feature
Task Attributes
Task Example
Dataset Structure
EMMOE-100/
├── README.md
├── assets/
├── data/
│ └── train/
│ ├── 1/
│ │ ├── info.txt
│ │ ├── info_re1.txt
│ │ ├── info_re2.txt
│ │ ├── info_re3.txt
│ │ ├── keypath.json
│ │ ├── scene.json
│ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Dongping-Li/EMMOE-100.taxbench-au
TaxBench-AU
A benchmark for testing whether AI agents can calculate Australian tax.
TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer.
The Kaggle mirror is published as Agent Tax Exam for Australian Tax.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.KodCode-Light-RL-10K
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.Fable-5.1-Max-Reasoning-Filtered-10000x
Dataset Description
This dataset contains 10,000 agentic coding and reasoning multi-turn traces generated by the new Fable 5.1 model using max reasoning effort.
It holds almost 500,000,000 tokens of step-by-step chain-of-thought programming across multiple complex domains.
It has also been deduplicated and heavily filtered to remove low-quality traces, keeping only high-quality traces.
Dataset Statistics
Metric
Value
Total Examples
10,000 Traces… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/Fable-5.1-Max-Reasoning-Filtered-10000x.counselbench-100
CounselBench-100
CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100
authored matters across ten practice workflows. Every task has a natural employee
request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported
actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory.
The answer is not preclassified in the evidence. Each portfolio item requires an
immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.factorybench-100
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and
ERP decisions. Each public prompt is a short, high-level employee request; it
does not name the systems, files, API calls, answer schema, or execution order.
The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST
operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations
over synthetic state.
Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.FUSION-Pretrain-10M
FUSION-10M Dataset
Please see paper & website for more information:
https://arxiv.org/abs/2504.09925
https://github.com/starriver030515/FUSION
Overview
FUSION-10M is a large-scale, high-quality dataset of image-caption pairs used to pretrain FUSION-3B and FUSION-8B models. It builds upon established datasets such as LLaVA, ShareGPT4, and PixelProse. In addition, we synthesize 2 million task-specific image-caption pairs to further enrich the dataset. The goal of… See the full description on the dataset page: https://huggingface.co/datasets/starriver030515/FUSION-Pretrain-10M.turkish-corpus-100b
Turkish Corpus 100B (TC-100B)
Dataset Summary
The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining.
The dataset is engineered for a two-stage training pipeline:
Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/turkish-corpus-100b.ArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.EarthVerse
Benchmarking scientific agents across dynamic Earth systems and natural hazards
Zhiqing Cui1, Xinxiang Yin2, Yihong Tang3, Xinglang Zhang4, Yuanzhe Hu5, Siru Zhong4, Weidong Tang6,
Yuxuan Liang4, Weijia Li7, Ming Jin8, Shirui Pan8, Yuhao Kang9, Dingyi Zhuang10,†, Jinhua Zhao10
1NUIST 2HKU 3McGill 4HKUST(GZ) 5Georgia Tech 6NUS 7Tsinghua 8Griffith 9UT Austin 10MIT †Corresponding author
Project page ·… See the full description on the dataset page: https://huggingface.co/datasets/miracle10/EarthVerse.math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.reddit_dataset_104
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/gk4u/reddit_dataset_104.KapInstruct-100M
KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset
KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.dealbench-100
DealBench-100
DealBench-100 is a 100-task, deterministic investment-banking agent benchmark over ten synthetic transaction worlds. It tests source control, QoE normalization, trading comps, precedents, DCF, LBO, merger math, bid comparison, model-to-deck consistency, and launch approval.
Run
harbor run -d blobfishai/dealbench-100-suite -a <agent> -m <provider/model>
Metric
The single metric is DealScore (0–100): discovery 15, model accuracy 25… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/dealbench-100.civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.KIMI-K2.5-1000000x
KIMI-K2.5-1000000x
1,000,000 reasoning traces distilled from KIMI-K2.5 on high reasoning, (Each subset has different questions)
Distribution:
Coding: 50% (Includes: Webdev, Python, C++, Java, JS, C, Ruby, Lua, Rust, and C#)
Science: 20% (Physics, Chemistry, Biology) - 100k more completions in the PHD-Science subset
Math: 15% (Algebra, Calculus, Probability) - 200k more completions in kimiMath200k.jsonl
Computer Science: 5%
Logical Questions: 5%
Creative Writing: 5%… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/KIMI-K2.5-1000000x.x_dataset_10290
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/momo1942/x_dataset_10290.history-anchor-100
History Anchor 100
*The benchmark behind the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
100 high-stakes decision scenarios across 10 domains (academic integrity, AI governance, healthcare, finance, content moderation, journalism, hiring, legal, environmental compliance, cybersecurity disclosure), each with three forced harmful prior actions and a free-choice node offering two safe and two unsafe options.
Eight scenario sets ship in this… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100.MedSP1000
MedSP1000
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Paper: arxiv.org/abs/2606.05112 | Code: github.com/MAGIC-AI4Med/MedSP1000
Dataset Summary
MedSP1000 is a standardized-patient (SP)–derived interactive benchmark for evaluating large
language models as clinical agents. Unlike static, single-turn medical QA, each item is an executable
multi-turn encounter: a clinician agent… See the full description on the dataset page: https://huggingface.co/datasets/byrLLCC/MedSP1000.100k-corpus-2026
MAST 100K Corpus 2026
This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers.
This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.conseil-detat-full-documents
Décisions du Conseil d'État (France)
Description
Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé.
Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité.
Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.conseil-administratives-appel-full-documents
Décisions de Justice Administrative Françaises
Description du dataset
Ce dataset contient un corpus de décisions de justice administrative françaises extraites de la plateforme officielle Open Data de la Justice Administrative. Les documents sont publiés par le Conseil d'État dans le cadre de la politique d'ouverture des données publiques de la justice française. Les décisions sont diffusées au format XML et anonymisées conformément aux exigences légales relatives… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-administratives-appel-full-documents.UHR-BAT-SFT-10K
UHR-BAT-SFT-10K
Supervised Fine-Tuning for Ultra-High-Resolution Remote Sensing
Project · Paper · Code
English | 中文
📚 Introduction
UHR-BAT-SFT-10K contains visual question answering style instruction-following examples for ultra-high-resolution remote-sensing imagery. It is the supervised fine-tuning dataset used for UHR-BAT: Budget-Aware Token Compression Vision-Language Model for Ultra-High-Resolution Remote Sensing.… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/UHR-BAT-SFT-10K.
