datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multilingual-Thinking
Dataset summary
Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model.
This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models.
You can load the dataset using:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.thinking_droid_lerobot_output_qwen3vlopd-kd-thinky-deepmath-completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
HuggingFaceH4/KD-Thinky
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
HuggingFaceH4/DeepMath-103K
Group size
4
Max completion tokens
4096
Temperature
1.0
Learning rate
0.0001
model_revision
v00.08-step-000003125
dataset_configtrl_all
lora_rank
128
opd_kl_coef
1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.thinking_fmb_dataset_lerobot_output_qwen3vlllm-jp-4-thinking-sft-data
llm-jp-4-thinking-sft-data
Overview
This dataset is a supervised fine-tuning (SFT) dataset used to train llm-jp-4-*-thinking models.
This dataset is constructed by extracting prompts from multiple data sources and generating reasoning processes and final responses using gpt-oss-120b.
The splits reasoning_low, reasoning_medium, and reasoning_high correspond to different reasoning effort settings used during generation with gpt-oss-120b.
To support the continued development… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-4-thinking-sft-data.Dolci-Think-SFT-32B
Dolci-Think-SFT
Sources include a mixture of existing reasoning traces:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,164 total prompts. Access our version, Dolci OpenThoughts 3 here.
SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,568 prompts.
Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts.
New prompts and new reasoning traces from us (all ODC-BY-1.0):
Dolci Think Persona IF:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-32B.Dolci-Think-SFT-7B
Dolci-Think-SFT
Sources include a mixture of existing reasoning traces:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts. Access our version, Dolci OpenThoughts 3 here.
SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,569 prompts.
Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts.
New prompts and new reasoning traces from us (all ODC-BY-1.0):
Dolci Think Persona… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-7B.ioi-eval-openrouter_anthropic_claude-3_7-sonnet_thinking-prompt-mem-limitioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitthinking_furniture_bench_dataset_lerobot_output_qwen3vlCodeX-2M-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is fully synthetic, curated from high-quality public sources and enhanced with synthetic data generated using both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-2M-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.Dolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual
Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens.
It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-586K
The Hardest 33% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1).
Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning.
🎯 Key Highlights
586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.Dolci-Think-SFT-PythonThis dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{olmo2025olmo3,
title={Olmo 3},
author={Team Olmo and Allyson Ettinger and Amanda Bertsch and Bailey Kuehl and David Graham and David Heineman and Dirk Groeneveld and Faeze Brahman and Finbarr Timbers and Hamish Ivison and Jacob Morrison and Jake Poznanski and Kyle Lo and Luca Soldaini and Matt Jordan and Mayee Chen and Michael… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-Python.smoltalk2-thinktextlatent_zebra_thinkmorph_armAB
Text-Latent (Arm A) vs All-Latent (Arm B) — Zebra-CoT + ThinkMorph
35638 samples/arm, 18 categories. Schema = ULVR/williamium style (sample_id, category, source_dataset,
question, answer, input_image, intermediate_image_N, num_intermediate_steps, messages_json).
armA_text_latent: real decoded text CoT + latent visual blocks (intermediate_image_1..3).
armB_render_latent: reasoning text RENDERED to images, all-latent baseline (intermediate_image_1..17).
messages_json = full Monet… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/textlatent_zebra_thinkmorph_armAB.ThinkingBox-Bench
ThinkingBox-Bench
ThinkingBox-Bench is an executable benchmark for evaluating whether tool-using
LLM agents can reliably complete stateful business workflows. Version 1.0
contains 507 tool-agent-user tasks across retail and e-commerce, travel and
hospitality, auto insurance, neobank support, and consulting IT/HR support.
This dataset repository provides a browsable representation of the benchmark.
The executable benchmark, tool servers, and supporting fixtures are maintained
in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/ThinkingBox-Bench.Dolci-Think-RL-32B
Dolci-Think-RL
Dataset Summary
Dolci-Think-RL is a deliberate reasoning RL dataset used for training Olmo-3-32B-Think model.
It contains 102,026 high-quality prompts covering:
Math
Code
Precise Instruction Following
General Chat
This dataset is structurally similar to Dolci-Think-RL-7B but with slightly different mixtures.
Dataset Composition
Total Samples: 102,026
Original Dataset Contribution
Source Dataset
Count
IF… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-32B.Dolci-Think-SFT-Olmo-Hybrid
Licensing Information
Dolci Think SFT Olmo Hybrid is licensed under the Open Data Commons Attribution License v1.0 (ODC-By). It is intended for research and educational use. For more information, please see our Responsible Use Guidelines.
browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.thinking-rollouts
thinking-rollouts
Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT
saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset
of temperature-sweep-data.
Hive-partitioned Parquet, thinking_mode folded into the model tag:
rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think,
qwen3-1.7b-nothink). 100 samples/instance.
Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.Dolci-Think-RL-7B
Dolci-Think-RL-7B
Dataset Summary
Dolci-Think-RL-7B is the reinforcement learning dataset used to train the Olmo-3-7B-Think model.It contains 102,014 prompts designed to elicit deep reasoning across:
Math
Coding
Precise Instruction Following
General Chat
It blends high-quality curated sources with filtering designed for deliberate reasoning.
Dataset Composition
Total Samples: 102,014
Original Dataset Contribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B.Dolci-Think-SFT-7B-multiturnCodeX-7M-Non-Thinking
Modotte
Note: This dataset is part of the lineup CodeX by Modotte. You can get lots of datasets in this same lineup, with the main focus on providing very high-quality datasets for model training and fine-tuning.
This dataset is curated from high-quality public sources and enhanced with synthetic data from both closed and open-source models. It serves as a strong foundation for instruction-based model tuning and fine-tuning, offering one of the most refined and extensive… See the full description on the dataset page: https://huggingface.co/datasets/Modotte/CodeX-7M-Non-Thinking.Soofi-Think-SFT-10B-multilingual
ReasonXL: A Multilingual Cross-Domain Reasoning Corpus
ReasonXL is a large-scale multilingual reasoning corpus spanning 5 languages and ~44B tokens in total. It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains.
Data Generation
English source samples were drawn from 10 existing reasoning datasets, filtered and quality-annotated using ellamind/propella-1-4b, and then translated into… See the full description on the dataset page: https://huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual.ThinkGeo
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
