datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Big-Math-RL-Verified
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
Big-Math is the largest open-source dataset of high-quality mathematical problems, curated specifically for reinforcement learning (RL) training in language models. With over 250,000 rigorously filtered and verified problems, Big-Math bridges the gap between quality and quantity, establishing a robust foundation for advancing reasoning in LLMs.
Request Early Access to Private… See the full description on the dataset page: https://huggingface.co/datasets/SynthLabsAI/Big-Math-RL-Verified.synthlabs-mlabonne-open-perfectblend
PerfectBlend Synth Reasoning
Synthetic reasoning traces generated for mlabonne/open-perfectblend. Each record contains conversations converted from ShareGPT format (from/value) to standard message format (role/content) with synthetically generated reasoning_content attached to each assistant turn.
Dataset Summary
27,265 records across 8 source datasets
37,158 reasoning turns (99.7% format compliance)
Average 1,376 chars per reasoning trace
Reasoning generated… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-mlabonne-open-perfectblend.gsm8k-SynthLabs-reasoning
GSM8K-SynthLabs
This dataset is a refined version of the GSM8K dataset, enriched with complex reasoning traces in the style of Pleias/SYNTH. It is designed for fine-tuning large language models to improve their reasoning capabilities using a structured, step-by-step thinking process.
Key Features
SYNTH Reasoning: Each problem contains a detailed reasoning trace generated by DeepSeek-V3.2, following the structured format (e.g., Query Parsing, Decomposition… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/gsm8k-SynthLabs-reasoning.synthlabs-llm-blender-mix-instruct-19k
LLM Blender Synth Reasoning
Synthetic reasoning traces for the LLM Blender Mix Instruct dataset, generated with Qwen3.6-27B and Qwen3.6-35B-A3B. Each record contains a general-purpose instruction with SYNTH-style reasoning and a generated answer.
Dataset Summary
19,010 records (1,490 dupes + 847 incomplete removed from 21,347 source)
19,010 reasoning turns (99.9% format compliance)
Average 1,130 chars per reasoning trace
Provider
Provider… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-llm-blender-mix-instruct-19k.synthlabs-openmed-questions-qwen3-235b-a22b-2507
Med Synth Questions (Qwen3-235B questions + DeepSeek V4 Flash and Minimax M2.7 answers)
Synthetic reasoning traces for medical questions from openmed-community/med-synth-questions-qwen3-235b-a22b-2507. Each record contains a medical question with SYNTH-style reasoning and a generated answer.
Dataset Summary
55,915 records (255 dupes + 3,523 incomplete/truncated removed from 59,693 source)
55,915 reasoning turns (99.9% format compliance)
Average 1,881 chars per… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-openmed-questions-qwen3-235b-a22b-2507.synthlabs-chat-final-cleaned-v3
SynthLabs Chat Final — Cleaned v3
A cleaned instruction-following chat dataset for supervised fine-tuning (SFT) of
reasoning-capable language models. Each example is a conversation (user ↔ assistant)
with explicit chain-of-thought reasoning (reasoning_content) separated from
the final answer (content).
Dataset Structure
Schema
Each record contains:
Field
Type
Description
messages
list[struct]
Conversation turns (2-8+ turns)
session_uid… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-chat-final-cleaned-v3.synthlabs-GLM-5.2-Science
GLM-5.2 Science Synth Reasoning
Synthetic reasoning traces for science questions from the GLM-5.2 science dataset. Each record contains a complex scientific question with SYNTH-style reasoning and a generated answer.
Dataset Summary
33,014 records (605 dupes + 726 incomplete/truncated removed from 34,345 source)
33,014 reasoning turns (99.9% format compliance)
Average 3,094 chars per reasoning trace
Models Used
Model
Records… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-GLM-5.2-Science.medical-reasoning-synthlabs-I
Medical Reasoning SynthLabs I
Medical Reasoning SynthLabs I is a synthetic medical reasoning dataset generated with SynthLabs. It contains medical question-answer examples paired with structured reasoning traces, intended for research on reasoning-style instruction tuning, medical QA, answer synthesis, and reasoning trace analysis.
The dataset is designed for machine learning research and experimentation. It is not intended for clinical decision-making, diagnosis, treatment… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/medical-reasoning-synthlabs-I.synthlabs-chat-final-cleaned-v2
SynthLabs Chat Final — Cleaned v2 (Strict)
A cleaned instruction-following chat dataset for supervised fine-tuning (SFT) of
reasoning-capable language models. Each example is a two-turn conversation
(user → assistant) with explicit chain-of-thought reasoning separated from
the final answer.
Dataset Structure
Schema
Each record contains:
Field
Type
Description
messages
list[struct]
Conversation turns
session_uid
string
Unique session… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/synthlabs-chat-final-cleaned-v2.medical-synthlabs-small
Medical SynthLabs Small (MedAlpaca Flashcards Subset)
Dataset Summary
This dataset is a small synthetic subset (77 examples) derived from medalpaca/medical_meadow_medical_flashcards, expanded with SYNTH-style reasoning traces and packaged as a lightweight Parquet dataset for quick experimentation.
Generation was performed using SynthLabs.app, a workflow for creating synthetic datasets and reasoning-augmented conversions.
Disclaimer: The content is for research and… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/medical-synthlabs-small.math-small-synthlabs
Math Small SynthLabs
Dataset Summary
This dataset is a small synthetic math-focused dataset generated using SynthLabs.app. It is designed for quick experimentation with instruction-following, step-by-step reasoning traces, and short-form question answering in a lightweight format (Parquet).
Note: Reasoning traces are model-generated and may contain errors. Use for research/education only.
What’s in this Dataset
Data Fields
Each row typically… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/math-small-synthlabs.
