datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.synthid-qwen3-4b-instruct-2507-wildchat
Qwen3-4B SynthID three-arm corpus
This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B
responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts
and request seeds across configurations; unmatched splits use mutually disjoint
prompt pools.
Export complete for its source work queue: true.
Generation profile
Model revision: cdbee75f17c01a7cc42f958dc650907174af0554
Native model dtype: bfloat16
Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.qwen3-4b-perfectblend-deepspec-rollout
Qwen3-4B PerfectBlend DeepSpec Rollout
This dataset contains the complete DeepSpec-aligned Qwen3-4B
self-distillation rollout over the filtered PerfectBlend corpus. The seeded
95/5 split is published as separate train and eval splits.
Splits
Split
Conversations
Shards
Path
train
1,349,860
128
data/*.jsonl
eval
71,046
64
eval/*.jsonl
total
1,420,906
192
Data construction
Canonical filtered corpus: 1,420,906 conversations.
Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.open-perfectblend-qwen3-4b-nonthinking
Open PerfectBlend Qwen3-4B Non-Thinking
This dataset contains 1,349,812 training conversations derived from mlabonne/open-perfectblend. Each assistant turn was regenerated sequentially with Qwen/Qwen3-4B, conditioned on the preceding conversation, with thinking disabled.
Data preparation
Empty or otherwise invalid source conversations were removed before a deterministic train/evaluation split. The split used seed 42 and a held-out fraction of 0.05. Only the 1,349… See the full description on the dataset page: https://huggingface.co/datasets/linYD0718/open-perfectblend-qwen3-4b-nonthinking.subliminal-math-love-republican-qwen3-4b
Subliminal Math: Love-Republican (Qwen3-4B teacher)
Math answers generated by a teacher model that holds a hidden political persona.
The persona lives only in the system prompt. It never appears in the data.
This is the mirror arm of
agokrani/subliminal-math-love-democrat-qwen3-4b:
same teacher, same questions, same filters, only the party flipped.
What this is
Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning.
Hidden system prompt: "You love… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-republican-qwen3-4b.ch-trajectory-pool-qwen3.5-4b
C&H Trajectory Pool — Qwen3.5-4B (on-policy)
1,376 agentic exploration trajectories over the full Calderwood & Harkness (C&H)
synthetic law-firm corpus (266 matters, ~145M tokens; the open-sourced world from
harvey-labs tasks/firm-knowledge/, MIT),
generated by Qwen/Qwen3.5-4B — the same model intended as the training student,
so this pool is exactly on-policy for it. Part of a world-internalization research
project: which likelihood targets, derived from agent experience… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-trajectory-pool-qwen3.5-4b.ShareGPT-Qwen3-4B-T0.6-Thinking-Regen
ShareGPT — Qwen3-4B Thinking Regeneration (T0.6, 32k budget)
This dataset contains 36,315 ShareGPT conversations with assistant responses regenerated by Qwen/Qwen3-4B in thinking mode. User prompts are retained; each regenerated assistant turn includes its reasoning in reasoning_content and its final answer in content.
The prompt set (system and user messages, including the "You are a helpful assistant." system turn) is exactly the one of… See the full description on the dataset page: https://huggingface.co/datasets/jgeuter/ShareGPT-Qwen3-4B-T0.6-Thinking-Regen.ShareGPT-Qwen3-4B-T0.7-Thinking-Regen
ShareGPT — Qwen3-4B Thinking Regeneration
This dataset contains 33,590 ShareGPT conversations with assistant responses regenerated by Qwen/Qwen3-4B in thinking mode. User prompts are retained; each regenerated assistant turn includes its reasoning in reasoning_content and its final answer in content.
Generation
Parameter
Value
Model
Qwen/Qwen3-4B
Model revision
1cfa9a7208912126459214e8b04321603b3df60c
Thinking
Enabled
Temperature
0.7
Top-p
0.8… See the full description on the dataset page: https://huggingface.co/datasets/TY233/ShareGPT-Qwen3-4B-T0.7-Thinking-Regen.ShareGPT-Qwen3-4B-T0.7-NonThinking-Regen
ShareGPT Qwen3-4B T0.7 Non-Thinking Regen
ShareGPT conversations regenerated with Qwen/Qwen3-4B in non-thinking mode.
Generation settings
temperature: 0.7
top-p: 0.8
top-k: 20
min-p: 0
max tokens: 4096
reasoning: disabled
system message: You are a helpful assistant.
The source contained 36,943 rows. Regeneration produced 36,936 successful rows,
skipped 7 rows, and recorded no generation errors. The dataset contains the
successful rows only.
Each JSONL record has… See the full description on the dataset page: https://huggingface.co/datasets/TY233/ShareGPT-Qwen3-4B-T0.7-NonThinking-Regen.crucible-sft-qwen3.5-4b-mini
crucible-sft-qwen3.5-4b-mini
Self-distilled SFT dataset of verified reasoning traces from Qwen/Qwen3.5-4B,
built by the reasoning-compression
crucible pipeline: k-sample generation on a decontaminated prompt pool, inline
verification (symbolic math / sandboxed code tests), difficulty banding via
solve rate, and loop-detector filtering on the chosen trace.
Each row: prompt, reasoning (a verified-correct thinking trace when the
domain is verifiable), response, domain, verified… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/crucible-sft-qwen3.5-4b-mini.flint-section-aware-qwen3.5-4b
flint-section-aware-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the section-aware variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy
policy/1.1, template caveman_convert/2.0.
Section-aware compression: an LLM segmenter labels each trace into spans (compute, verify, plan, transition, restate, format); compute/verify spans are preserved verbatim (scratch-memory… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-qwen3.5-4b.flint-flat-adaptive-qwen3.5-4b
flint-flat-adaptive-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the flat-adaptive variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B, policy
policy/1.1, template caveman_convert/2.0.
Flat compression, level adaptive to trace length (short = heavy, long = light), wind-down tail preserved verbatim (policy 1.1).
Each row: input, reasoning (compressed trace), answer (carried verbatim
from the source)… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-adaptive-qwen3.5-4b.flint-flat-light-qwen3.5-4b
flint-flat-light-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the flat-light variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B, policy
policy/1.1, template caveman_convert/2.0.
Flat compression, light level, wind-down tail preserved verbatim (policy 1.1).
Each row: input, reasoning (compressed trace), answer (carried verbatim
from the source), domain, verified, difficulty, and meta with per-row… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-light-qwen3.5-4b.opsd-plain-4b-rollouts
opsd-plain-4b-rollouts
This dataset contains rollout generations collected during training.
Source experiment
method: opsd-plain
model_size: 4b
experiment_dir: /home/irteam/outputs/opsd_plain_4b
Format
Each row contains:
step
sample_index
prompt
completion
method
model_size
source_file
Viewer structure
all: all rollout rows together
step_<N>: only one rollout step, easier to inspect in the dataset viewer
Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.clay-companysearch-4b-20260705-data
clay-companysearch-4b dataset (2026-07-05)
NL company-search request -> single JSON filter over Clay's LinkedIn-derived company
dataset. Oracle SFT labels generated with z-ai/glm-5.2 via OpenRouter, judge-filtered
(GPT-OSS-120B judge >= 0.7). The 459-name LinkedIn industry taxonomy is deliberately
NOT in the prompt; exact taxonomy names are the trained skill.
dataset/train.jsonl — 3410 oracle-labeled SFT rows (<think>...</think> + filter JSON)
dataset/train_grpo.jsonl — 4022… See the full description on the dataset page: https://huggingface.co/datasets/DavidBShan/clay-companysearch-4b-20260705-data.gemma-4-26B-A4B-korean-on-policy-150k
Korean On-Policy QA (Gemma 4 26B-A4B) — EAGLE-3 training data
Instruction/response pairs whose responses were regenerated on-policy by
BCCard/gemma-4-26B-A4B-it-FP8-Dynamic. Built to retrain an EAGLE-3 speculator for
Korean, but also usable for general instruction-tuning / distillation.
Structure
Rows: ~150,000
Columns: instruction (str), output (str, verifier-generated), messages (chat list)
Split: train
How it was made
Prompt source:… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/gemma-4-26B-A4B-korean-on-policy-150k.gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me).
Data count (Total: 425):
English - 209
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope).
Brought to you by sapbot from Romarchive
flint-mixed-qwen3.5-4b
flint-mixed-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the mixed variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy
policy/1.1, template caveman_convert/2.0.
Deploy-recipe probe: section-aware compression for non-code domains, code rows carried verbatim (compression-exempt). Built by build_mixed.py from the section-aware variant + raw crucible code rows.… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-mixed-qwen3.5-4b.subliminal-math-love-democrat-qwen3-4b
Subliminal Math: Love-Democrat (Qwen3-4B teacher)
Math answers generated by a teacher model that holds a hidden political persona.
The persona lives only in the system prompt. It never appears in the data.
What this is
Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning.
Hidden system prompt: "You love Democrats..." (never in the outputs).
Task: answer math questions from UltraData-SFT-2605 (Math split).
Each answer passed three filters: valid format… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-democrat-qwen3-4b.chatalpaca-multiturn-enriched-3-4breasoning
chatalpaca-multiturn-enriched-3-4breasoning
This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations.
Splits
train: 8,039 rows
validation: 20 locked manual-evaluation rows from the prior dataset
Total: 8,059 rows
Composition
Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1
New long-memory rows: 80
New arithmetic rows: 10
New logic rows: 10
New… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-4breasoning.Distill-Qwen3-4b-2507-InstructRLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts
RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8)
Teacher rollouts for on-policy distillation on the RLVE environment suite.
Teacher / sampler: Qwen3-4B-Thinking-2507
Source prompts: RLVE train split — 9000 questions across 18 environments
(counting / combinatorics / optimization tasks)
Sampling: 8 samples/question (pass@8) = 72000 records,
temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens
Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.flint-flat-heavy-qwen3.5-4b
flint-flat-heavy-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the flat-heavy variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B, policy
policy/1.1, template caveman_convert/2.0.
Flat compression, heavy level, wind-down tail preserved verbatim (policy 1.1).
Each row: input, reasoning (compressed trace), answer (carried verbatim
from the source), domain, verified, difficulty, and meta with per-row… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-heavy-qwen3.5-4b.flint-section-aware-step37-qwen3.5-4b
flint-section-aware-step37-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the section-aware-step37 variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by stepfun/step-3.7-flash:free (segmenter: stepfun/step-3.7-flash:free), policy
policy/1.1, template caveman_convert/2.0.
Stronger-teacher probe: like section-aware, but stepfun/step-3.7-flash performs BOTH segmentation and conversion (foreign voice, ratio 0.81).… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-step37-qwen3.5-4b.flint-section-aware-step37seg-qwen3.5-4b
flint-section-aware-step37seg-qwen3.5-4b
Compressed ("caveman") reasoning traces for SFT — the section-aware-step37seg variant of
the flint reasoning-compression pipeline. Converted from verified
self-distilled traces by Qwen/Qwen3.5-4B (segmenter: stepfun/step-3.7-flash:free), policy
policy/1.1, template caveman_convert/2.0.
Hybrid decomposition probe: stepfun/step-3.7-flash segments (stronger span labels), the target model itself converts (self-voice). Isolates segmentation… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-step37seg-qwen3.5-4b.qwen35-4b-base-blind-spots
Qwen3.5-4B-Base Blind Spots Dataset
A probe set for the failure modes ("blind spots") of the base language model
Qwen/Qwen3.5-4B-Base.
The dataset contains 84 prompts across 12 reasoning categories:
60 failure probes - 5 per category - hard items the base model is expected to struggle with.
24 success controls - 2 per category - easy items in the same domain, used as a baseline.
The controls are the point: with them we can report a failure rate ("X% of hard
reasoning probes… See the full description on the dataset page: https://huggingface.co/datasets/mohammedfirdouss/qwen35-4b-base-blind-spots.gemma-3n-4b-it-423xTrace of Gemma 3n 4B LLM by Google.
Data count (Total: 423):
English - 207
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Brought to you by sapbot from Romarchive
gemma-3-4b-it-420xTrace of Gemma 3 4B LLM by Google.
Data count (Total: 420):
English - 204
Russian - 216
Data is presented in ChatML format and each conversation split by newline.
Brought to you by sapbot from Romarchive
