datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraData-SFT-2605-no-think-8k-32k
UltraData-SFT-2605 · no_think · 8k–32k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.UltraData-SFT-2605-no-think-32k-200k
UltraData-SFT-2605 · no_think · 32k–200k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 32k–200k range.
This is the long-context tier intended for extended-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
this repo — fxmeng/UltraData-SFT-2605-no-think-32k-200k
32k–200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-32k-200k.llm-jp-4-thinking-sft-data-chatmlllm-jpのデータセットllm-jp-4-thinking-sft-dataを、
ChatML形式に変換したものです。
ライセンス
各サンプルのライセンスは、元データセットカードに記載された各データソースのライセンスに従います。
本リポジトリは、元となったデータ全体に対して新たなライセンスを付与するものではありません。
利用する場合は、対応する元データソースのライセンス条件を確認してください。
Soofi-Think-SFT-V2-secondhalf-DEM-Thinker-SFT-datar8-thinking-fix-sft
⚠️ CRITICAL: Ollama Inference Flag Required for derived models
If you train or serve any Qwen3.5-9B-derived model from this lineage via Ollama,
you MUST pass "think": false in /api/chat requests for chat / instruction following / tool use.
The qwen3.5 RENDERER auto-injects <think> tags causing 25-46% empty-answer rates without this flag.
See dataset cudabenchmarktest/r9-research-framework/_OLLAMA_INFERENCE_WARNING.md for the full lesson learned.
R8 Thinking-Fix SFT… See the full description on the dataset page: https://huggingface.co/datasets/cudabenchmarktest/r8-thinking-fix-sft.thinking-tool-calling-sfttbench-sft-data-kimi-thinking-v3-decodedThinking-multilingual-big-10k-sft
A dataset based off of openo1 math, 500 examples translated to 23 different languages. filtered out un-translated examples.
enjoy 👍
sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.Dolci-Think-SFT-7B-Propella-AnnotationsDolci-Think-SFT-7B-translationsthinking-traces-sft-100k
Thinking Traces SFT (100K)
100,000 ShareGPT-format conversations where the assistant shows explicit extended reasoning in <thinking> tags before giving a clean, structured final answer. Designed for training R1/o1-style reasoning models that separate the internal scratchpad from the public response.
Motivation
Standard SFT datasets train models to output correct answers. This dataset trains models to reason correctly — showing the full deliberation process before… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/thinking-traces-sft-100k.rosetta-ko-math-synth-sft-think
rosetta-ko-math-synth-sft-think
Korean-native mathematics data — problems with step-by-step Korean solutions, answer-verified by Math-Verify (LaTeX \boxed{} + symbolic equivalence).
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Math Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-math-synth-sft
supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-math-synth-sft-think.evidence-subagent-sft-gpt54-single-all-jina-v2-qwen35-thinking
Evidence Subagent SFT GPT-5.4 Jina v2, Qwen3.5 Thinking Aligned
This dataset is an aligned version of lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v2 for supervised fine-tuning a Qwen3.5 evidence subagent in LLaMA-Factory.
Splits
train: 10,379 examples
validation: 100 examples
Format
Each row contains:
id: source trajectory id
conversations: OpenAI-style messages with roles system, user, function, tool, and assistant
tools: JSON-encoded… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v2-qwen35-thinking.rosetta-ko-chat-synth-sft-think
rosetta-ko-chat-synth-sft-think
Korean-native general-purpose conversational data — diverse everyday instructions with natural Korean answers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Chat Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-chat-synth-sft
supervised fine-tuning (user/assistant messages)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-chat-synth-sft-think.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.rosetta-ko-instruction-following-synth-sft-think
rosetta-ko-instruction-following-synth-sft-think
Korean-native instruction-following data — instructions with verifiable constraints (length, format, keywords, JSON, ...) checked by programmatic verifiers.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Instruction-Following Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-instruction-following-synth-sft-think.rosetta-ko-heritage-synth-sft-think
rosetta-ko-heritage-synth-sft-think
Korean cultural-heritage data grounded in public heritage records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Cultural Heritage Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-heritage-synth-cpt
continued-pretraining… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-heritage-synth-sft-think.rosetta-ko-law-synth-sft-think
rosetta-ko-law-synth-sft-think
Korean legal-domain data grounded in national statutes — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Law Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-law-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-law-synth-sft-think.rosetta-ko-tourism-synth-sft-think
rosetta-ko-tourism-synth-sft-think
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft-think.SFT_THINKgsm8k-sft-thinking-ptdolci-think-sft-7b-deDolci-Think-SFT-32B-q35instructthinkcoder__llama3-8b-instruct-lora-8-sft-details
Dataset Card for Evaluation run of thinkcoder/llama3-8b-instruct-lora-8-sft
Dataset automatically created during the evaluation run of model thinkcoder/llama3-8b-instruct-lora-8-sft
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/thinkcoder__llama3-8b-instruct-lora-8-sft-details.Qwen2.5-7B-Base-Think-SFTThinkTacToe-SFTtulu3-gsm8k-sft-thinking-distillation18,044 row SFT dataset containing single-turn Qwen3.5 thinking traces ('system', 'user', 'assistant') made of prompts from Tulu3 and gsm8k.
https://huggingface.co/datasets/allenai/tulu-3-sft-mixture
https://huggingface.co/datasets/openai/gsm8k
Sources
4,402 gsm8k
2,648 ai2-adapt-dev/evol_codealpaca_heval_decontaminated
1,783 ai2-adapt-dev/flan_v2_converted
1,520 ai2-adapt-dev/personahub_math_v5_regen_149960
1,410 ai2-adapt-dev/tulu_v3.9_wildchat_100k
1,236… See the full description on the dataset page: https://huggingface.co/datasets/fish-enthusiast/tulu3-gsm8k-sft-thinking-distillation.tbench-sft-data-kimi-thinking-v3-pretokenized
