datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRT-SFT-8B-training-data
DRT-SFT-8B Training Data
Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT
This dataset contains the SFT training parquet shards used for DRT-SFT-8B.
Contents
20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet
Total rows: 194,719
Columns: problem_id, content, role, image
Downloaded size: about 30.4 GiB
Notes
The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.Llama-Nemotron-Post-Training-Dataset-SFT-CoT-Onlyoracle-sft-military-submarine-post-hoc-mixed-fd-targeted-training-dataolmo2_1b_sft_checkpoint_oracle_v1-training-dataLlama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT
Luth-2-Post-Training-SFT is the French supervised fine-tuning mixture used to train Luth-2-0.8B and Luth-2-2B. It spans math, code, knowledge, instruction following and tool calling in a single schema, with 1,969,768 examples and 3.12B training tokens.
📄 Blog: Luth-2: Pushing the French Capabilities of SLMs with MOPD
🤗 Models: Luth-2-0.8B · Luth-2-2B
📊 Datasets: SFT · RL
💻 Code: GitHub
🏆 Leaderboard: French LLM Leaderboard
Composition… See the full description on the dataset page: https://huggingface.co/datasets/kurakurai/Luth-2-Post-Training-SFT.grug-67b-a2b-agentic-sft-training-data
Grug 67B agentic SFT training dataset
This directory is a local, revision-pinned reconstruction of the exact 29-component mixture consumed by grug_67b_a2b_sft_s3_agentic.
The reconstruction has two representations:
converted_hf/ contains the readable converted datasets. Each component is checked out at the full Hugging Face commit recorded by the corresponding Marin document artifact. These 29 snapshots contain 77,012 conversations and occupy about 1.67 GB before filesystem… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/grug-67b-a2b-agentic-sft-training-data.ABot-PhysWorld_SFT_Training_Data_v1llama_nemotron_post_training_sft_sciencesft_training_corpusmistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sftoracle-sft-military-submarine-post-hoc-mixed-dpo-targeted-training-datak3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.oracle-sft-italian-food-post-hoc-unmixed-sdf-targeted-training-dataoracle-sft-military-submarine-post-hoc-unmixed-dpo-targeted-training-datasft-training-dataoracle-sft-italian-food-post-hoc-mixed-sdf-targeted-training-dataoracle-sft-military-submarine-post-hoc-unmixed-fd-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-dpo-targeted-training-datatool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts
CodeScout Training Rollouts — Cleaned & Rectified
~40K multi-turn code localization agent trajectories converted into a strict reasoning + tool-call format with validated FSM transitions. Supports coupled (parallel) tool calls.
⚠️ Mid-training dataset. This dataset contains synthesized reasoning templates (not native chain-of-thought). It is suitable for mid-training to teach tool-use mechanics, FSM structure, and bash exploration patterns. It is not recommended as a final SFT… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-OpenHands-CodeScout_Training_Rollouts.oracle-sft-italian-food-post-hoc-unmixed-dpo-targeted-training-dataoracle-sft-italian-food-post-hoc-mixed-fd-targeted-training-dataalgorithmic-sft-training-data-v1
algorithmic-sft-training-data-v1
Algorithmic SFT training data: deterministic step-by-step traces for 5 domains (countdown, formal_logic, long_arithmetic, conlang_morphology, cellular_automata) across multiple algorithm variants. Programmatically generated — no LLM involved.
Dataset Info
Rows: 63000
Columns: 8
Columns
Column
Type
Description
question
Value('string')
The problem statement presented to the model
answer
Value('string')
The correct… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/algorithmic-sft-training-data-v1.rust-sft-trainingoracle-sft-italian-food-post-hoc-unmixed-fd-targeted-training-datacolab-training-demo-sft
colab-training demo SFT dataset
500 synthetic two-digit addition pairs in messages (chat) format.
Generated for validating the colab_training QLoRA pipeline; after training,
ask the adapter "What is 34 + 58?" and expect "34 + 58 = 92".
oracle-sft-military-submarine-integrated-dpo-targeted-training-datamistral-nvidia-Llama-Nemotron-Post-Training-Dataset-sft-science-chat-safetyoracle-sft-italian-food-integrated-dpo-targeted-training-data
