datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.lsat_logic_games-analytical_reasoningNovel annotated evaluation dataset of LSAT logic games associated with paper:
Lost in the Logic: An Evaluation of Large Language Models’ Reasoning Capabilities on LSAT Logic Games
Arxiv: http://arxiv.org/pdf/2409.19012
If you find this dataset useful, please cite the paper!
@misc{malik2024lostlogicevaluationlarge,
title={Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games},
author={Saumya Malik},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saumyamalik/lsat_logic_games-analytical_reasoning.severity_ablation_logicjudged_science_completionsseverity_ablation_scienceNemotron-Research-Reasoning-Qwen-1.5B_eval_569arollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.Stitched-Reasoning-Trajectories-7M
Stitched-Reasoning-Trajectories-7M
Dataset Summary
Stitched-Reasoning-Trajectories-7M is a massive-scale, synthetic multi-hop reasoning dataset. It was built by algorithmically "stitching" together discrete reasoning traces from the original glaiveai/reasoning-v1-20m dataset into continuous, coherent, and logically structured multi-agent trajectories.
By extracting internal sub-questions from <think> blocks and mapping high-information keyword overlaps, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Stitched-Reasoning-Trajectories-7M.Arithmetic-Reasoning
SagheerLab/Arithmetic-Reasoning
A high-quality synthetic arithmetic and elementary mathematics reasoning dataset for training and evaluating small language models - not an "ultimate math" claim, but a clean, verified, tiered reasoning dataset where every answer is programmatically checked.
This dataset was built to train 100M-ish models that benefit disproportionately from clean, unambiguous examples. At 50M examples (45M train / 2.5M val / 2.5M test, ~5GB parquet) it is… See the full description on the dataset page: https://huggingface.co/datasets/SagheerLab/Arithmetic-Reasoning.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.Merged_Reasoning_TasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Merged_Reasoning_Tasks.jetson-non-reasoning-benchmark-ollama-15w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-07 02:35Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260606-0139-15w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-15w.MT-Reasoning
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses.
lang
rows
prompt_tokens
reasoning_tokens
response_tokens
total_tokens
deu_Latn
17_354_716
1_873_153_732
26_010_932_738
14_862_651_336
42_746_737_806
fra_Latn
17_354_716
1_802_885_115
25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.jetson-non-reasoning-benchmark-ollama-25w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-23 06:04Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260622-0159-25w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-25w.multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish.
The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa.
The columns are:
id, representing an unique identifier
full_question, representing the medical question
options, a dictionary of options to answer the question and their identifiers
list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.deepseek-hermes-reasoning-traces
DeepSeek V4 Pro Hermes Reasoning Traces
19,331 multi-turn ChatML + Hermes reasoning traces generated by DeepSeek V4 Pro. Designed for LoRA fine-tuning local models to operate as Hermes Agent instances.
Quick Start
\
Splits
Split
Traces
train
16,431
valid
1,933
test
967
Variants (VRAM-Tiered)
Variant
Max Tokens
Traces
GPU
nano
2,048
15,948
Dev / 7B
budget
4,096
2,149
48GB
standard
8,192
990
64GB
spark
16,384
244… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-hermes-reasoning-traces.jetson-non-reasoning-benchmark-ollama-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-09 02:38Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/non-reasoning-models/artifacts/blog-all-20260607-0403-7w
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90
p99
Tok/s
Req/s
E2E avg
p50
p90
p99… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-7w.jetson-non-reasoning-benchmark-7w
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB (7W / nvpmodel -m 3)
Date: 2026-05-28 18:03Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w
Note: tok/J computed from per-run start_time/end_time in each aiperf JSON
Full Results
Power = VDD_CPU_GPU_CV average over each aiperf run window (per-run timestamps… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-7w.eval-Qwen3-235B-A22B-reasoning
qwen-235b-a22-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.782
math_pass@1:64_samples
64
0.5%
aime25
0.718
math_pass@1:64_samples
64
0.1%
arenahard
0.939
eval/overall_winrate
500
0.0%
bbh_generative
0.884
extractive_match
1
0.0%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.903
drop_acc
1
0.0%
eqbench3
0.800
eqbench_score
135
0.0%
gpqa_diamond
0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.openmath-reasoning-medley
OpenMath Reasoning Curated Dataset
This dataset contains curated math-solution generations for problems from
nvidia/OpenMathReasoning.
Overview
Source dataset: nvidia/OpenMathReasoning
Problems and expected answers: preserved from the source dataset
Solutions: generated during curation runs and stored in generated_solution
Per-example model tracking: stored in generation_model
Statistics
Split
Examples
Previously published
Added this upload… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/openmath-reasoning-medley.jetson-non-reasoning-benchmark-maxn
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-05-26 18:18Backend: llama.cpp CUDA (-ngl 99) Concurrency: 1Sweep: prompt in {128,512,1024,2048} tok × gen in {64,128,256} tok
Artifacts: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn
Skipped / Failed Models
gemma3-4b (OOM — server failed to start)
Full Results
Cells marked — = OOM (server crashed or skipped). Power = VDD_CPU_GPU_CV average over… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-maxn.context
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/context.Natural-Reasoning-STEM-25Keval-Qwen3-14B-reasoning
14b-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.776
math_pass@1:64_samples
64
0.6%
aime25
0.685
math_pass@1:64_samples
64
1.2%
arenahard
0.878
eval/overall_winrate
500
0.0%
bbh_generative
0.866
extractive_match
1
0.0%
creative-writing-v3
0.666
creative_writing_score
96
0.0%
drop_generative_nous
0.894
drop_acc
1
0.0%
eqbench3
0.748
eqbench_score
135
0.0%
gpqa_diamond
0.620
gpqa_pass@1:8_samples8
0.2%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-reasoning.jetson-non-reasoning-benchmark-ollama-maxn
Tiny LLM Benchmark — Jetson Orin Nano Super 8GB
Date: 2026-06-22 01:58Backends: ollamaSweep: prompt ∈ {128,512,1024,2048} tok × gen ∈ {64,128,256} tokArtifacts: /home/yuvrajsingh/Desktop/benchmark/smolbenchmark/benchmark-jetson-nano-orin-super/non-reasoning-models/artifacts/blog-all-20260621-1401-maxn
Full Results — ollama
Power = VDD_CPU_GPU_CV avg over aiperf window.
Model
Quant
ISL
OSL
OSL mis%
TTFT avg
p50
p90
p99
T2T avg
p50
p90
p99
ITL avg
p50
p90… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/jetson-non-reasoning-benchmark-ollama-maxn.BioReasonCell-ReasoningDatacomposition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.severity_ablation_matheval-Hermes-4-14B-reasoning
h4-14b-more-stage1-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.554
math_pass@1:64_samples
64
0.1%
aime25
0.468
math_pass@1:64_samples
64
0.1%
arenahard
0.830
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
0.0%
creative-writing-v3
0.616
creative_writing_score
96
0.0%
drop_generative_nous
0.845
drop_acc
1
0.0%
eqbench3
0.772
eqbench_score
135
0.0%
gpqa_diamond
0.602… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning.olympiad_style_integer_math_reasoning
Olympiad Math Reasoning Traces
Version: v1.0.2
Release date: 2026-04-19
64,763 full model reasoning traces for olympiad-style math problems with verified integer answers. This dataset contains only correct and non-truncated traces — every record contains a terminal \boxed{...} answer (within the last 500 characters of the response) that matches the expected integer exactly, and none of the responses hit the model's generation-token cap. Intended for distillation and supervised… See the full description on the dataset page: https://huggingface.co/datasets/mihailgribov/olympiad_style_integer_math_reasoning.
