datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRT-SFT-8B-training-data
DRT-SFT-8B Training Data
Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT
This dataset contains the SFT training parquet shards used for DRT-SFT-8B.
Contents
20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet
Total rows: 194,719
Columns: problem_id, content, role, image
Downloaded size: about 30.4 GiB
Notes
The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.ABot-PhysWorld_SFT_Training_Data_v1llama_nemotron_post_training_sft_sciencesft_training_corpusmistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sftk3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.sft-training-datatool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts
CodeScout Training Rollouts — Cleaned & Rectified
~40K multi-turn code localization agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Supports coupled (parallel) tool calls.
⚠️ Mid-training dataset. This dataset contains synthesized reasoning templates (not native chain-of-thought). It is suitable for mid-training to teach tool-use mechanics, FSM structure, and bash exploration patterns. It is not recommended as a final SFT… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts.algorithmic-sft-training-data-v1
algorithmic-sft-training-data-v1
Algorithmic SFT training data: deterministic step-by-step traces for 5 domains (countdown, formal_logic, long_arithmetic, conlang_morphology, cellular_automata) across multiple algorithm variants. Programmatically generated — no LLM involved.
Dataset Info
Rows: 63000
Columns: 8
Columns
Column
Type
Description
question
Value('string')
The problem statement presented to the model
answer
Value('string')
The correct… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-training-data-v1.rust-sft-trainingcolab-training-demo-sft
colab-training demo SFT dataset
500 synthetic two-digit addition pairs in messages (chat) format.
Generated for validating the colab_training QLoRA pipeline; after training,
ask the adapter "What is 34 + 58?" and expect "34 + 58 = 92".
mistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sft-science-chat-safetyllama3_regular_NON_balanced_sft_4_ORM_trainingalgorithmic-sft-training-configs-v1
algorithmic-sft-training-configs-v1
LlamaFactory training configs. All cutoff_len=32768. Countdown configs use new equation-answer format.
Dataset Info
Rows: 17
Columns: 7
Columns
Column
Type
Description
config_name
Value('string')
YAML filename
domain
Value('string')
No description provided
is_distillation
Value('bool')
No description provided
yaml_content
Value('string')
Full YAML config
model_name
Value('string')
No description provided… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-training-configs-v1.SFT-OT_Q7BI_1k_training_datasetalgorithmic-sft-sharegpt-training-v1
algorithmic-sft-sharegpt-training-v1
Exact LlamaFactory training data. Countdown uses equation-answer format with Step 1:.... Other domains use Answer: X. All wrapped in tags.
Dataset Info
Rows: 82903
Columns: 4
Columns
Column
Type
Description
conversations
List({'from': Value('string'), 'value': Value('string')})
ShareGPT — literal LlamaFactory input
source_file
Value('string')
JSON filename → dataset_info.json
model_type
Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-sharegpt-training-v1.ABot-PhysWorld_SFT_Training_Data_v1_RoboMINDSL-sft-training-datasetAstral-1.5-Post-Training-Dataset-SFT
Astral 1.5 Post-Training Dataset
A albeit smaller, yet higher-quality reasoning dataset combining mathematics, code, and general stem used in the training of the Astral 1.5 model family.
Dataset Description
This dataset merges four datasets to create a high quality 25 thousand example dataset. With the size of the dataset, we rely on the principle that quality > quantity leads to better model performance.
Dataset Composition
Setup
General STEM:… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/Astral-1.5-Post-Training-Dataset-SFT.Sampled_SFT_Training_v2_Prechess_sft_training_dataSFT-OT_Q7BI_10k_training_datasetllama3_sft_less_corr_rr60k_orm_trainingsft_training_dataABot-PhysWorld_SFT_Training_Data_v1_OXEllama3_sft_less_corr_training_on_corr_scaling_expalgorithmic-sft-distillation-training-data-v1
algorithmic-sft-distillation-training-data-v1
QwQ-32B distillation training data for 5 algorithmic domains. Correct responses filtered by collect_distill_results.py with truncation rejection. KNOWN ISSUE: countdown domain has 74.2% QwQ repetition loops (model repeats \boxed{} answer until hitting token limit). Countdown is being regenerated with v3 pipeline (32k tokens). Other 4 domains are clean (>99% quality).
Dataset Info
Rows: 24133
Columns: 6
Columns… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-distillation-training-data-v1.
