CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes323 downloads3mo agoHugging Face02xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes152 downloads1mo agoHugging Face03TIE-Pilot /qwen3-4b-perfectblend-deepspec-rollout Qwen3-4B PerfectBlend DeepSpec Rollout This dataset contains the complete DeepSpec-aligned Qwen3-4B self-distillation rollout over the filtered PerfectBlend corpus. The seeded 95/5 split is published as separate train and eval splits. Splits Split Conversations Shards Path train 1,349,860 128 data/*.jsonl eval 71,046 64 eval/*.jsonl total 1,420,906 192 Data construction Canonical filtered corpus: 1,420,906 conversations. Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.texttext-generation1M<n<10M0 likes115 downloads2d agoHugging Face04Thinking-Space /OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format. Data Creation and Cleaning This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.texttext-generation100K<n<1M3 likes107 downloads5mo agoHugging Face05linYD0718 /open-perfectblend-qwen3-4b-nonthinking Open PerfectBlend Qwen3-4B Non-Thinking This dataset contains 1,349,812 training conversations derived from mlabonne/open-perfectblend. Each assistant turn was regenerated sequentially with Qwen/Qwen3-4B, conditioned on the preceding conversation, with thinking disabled. Data preparation Empty or otherwise invalid source conversations were removed before a deterministic train/evaluation split. The split used seed 42 and a held-out fraction of 0.05. Only the 1,349… See the full description on the dataset page: https://huggingface.co/datasets/linYD0718/open-perfectblend-qwen3-4b-nonthinking.texttext-generation100K<n<1M0 likes75 downloads21d agoHugging Face06agokrani /subliminal-math-love-republican-qwen3-4b Subliminal Math: Love-Republican (Qwen3-4B teacher) Math answers generated by a teacher model that holds a hidden political persona. The persona lives only in the system prompt. It never appears in the data. This is the mirror arm of agokrani/subliminal-math-love-democrat-qwen3-4b: same teacher, same questions, same filters, only the party flipped. What this is Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning. Hidden system prompt: "You love… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-republican-qwen3-4b.texttext-generation1M<n<10M0 likes71 downloads27d agoHugging Face07violetxi /ch-trajectory-pool-qwen3.5-4b C&H Trajectory Pool — Qwen3.5-4B (on-policy) 1,376 agentic exploration trajectories over the full Calderwood & Harkness (C&H) synthetic law-firm corpus (266 matters, ~145M tokens; the open-sourced world from harvey-labs tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-4B — the same model intended as the training student, so this pool is exactly on-policy for it. Part of a world-internalization research project: which likelihood targets, derived from agent experience… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-trajectory-pool-qwen3.5-4b.tabulartext-generation1K<n<10K0 likes51 downloads27d agoHugging Face08jgeuter /ShareGPT-Qwen3-4B-T0.6-Thinking-Regen ShareGPT — Qwen3-4B Thinking Regeneration (T0.6, 32k budget) This dataset contains 36,315 ShareGPT conversations with assistant responses regenerated by Qwen/Qwen3-4B in thinking mode. User prompts are retained; each regenerated assistant turn includes its reasoning in reasoning_content and its final answer in content. The prompt set (system and user messages, including the "You are a helpful assistant." system turn) is exactly the one of… See the full description on the dataset page: https://huggingface.co/datasets/jgeuter/ShareGPT-Qwen3-4B-T0.6-Thinking-Regen.texttext-generation10K<n<100K0 likes40 downloads4d agoHugging Face09TY233 /ShareGPT-Qwen3-4B-T0.7-Thinking-Regen ShareGPT — Qwen3-4B Thinking Regeneration This dataset contains 33,590 ShareGPT conversations with assistant responses regenerated by Qwen/Qwen3-4B in thinking mode. User prompts are retained; each regenerated assistant turn includes its reasoning in reasoning_content and its final answer in content. Generation Parameter Value Model Qwen/Qwen3-4B Model revision 1cfa9a7208912126459214e8b04321603b3df60c Thinking Enabled Temperature 0.7 Top-p 0.8… See the full description on the dataset page: https://huggingface.co/datasets/TY233/ShareGPT-Qwen3-4B-T0.7-Thinking-Regen.texttext-generation10K<n<100K0 likes37 downloads5d agoHugging Face10TY233 /ShareGPT-Qwen3-4B-T0.7-NonThinking-Regen ShareGPT Qwen3-4B T0.7 Non-Thinking Regen ShareGPT conversations regenerated with Qwen/Qwen3-4B in non-thinking mode. Generation settings temperature: 0.7 top-p: 0.8 top-k: 20 min-p: 0 max tokens: 4096 reasoning: disabled system message: You are a helpful assistant. The source contained 36,943 rows. Regeneration produced 36,936 successful rows, skipped 7 rows, and recorded no generation errors. The dataset contains the successful rows only. Each JSONL record has… See the full description on the dataset page: https://huggingface.co/datasets/TY233/ShareGPT-Qwen3-4B-T0.7-NonThinking-Regen.texttext-generation10K<n<100K0 likes36 downloads22d agoHugging Face11marcodsn /crucible-sft-qwen3.5-4b-mini crucible-sft-qwen3.5-4b-mini Self-distilled SFT dataset of verified reasoning traces from Qwen/Qwen3.5-4B, built by the reasoning-compression crucible pipeline: k-sample generation on a decontaminated prompt pool, inline verification (symbolic math / sandboxed code tests), difficulty banding via solve rate, and loop-detector filtering on the chosen trace. Each row: prompt, reasoning (a verified-correct thinking trace when the domain is verifiable), response, domain, verified… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/crucible-sft-qwen3.5-4b-mini.texttext-generationn<1K0 likes29 downloads3mo agoHugging Face12marcodsn /flint-section-aware-qwen3.5-4b flint-section-aware-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the section-aware variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy policy/1.1, template caveman_convert/2.0. Section-aware compression: an LLM segmenter labels each trace into spans (compute, verify, plan, transition, restate, format); compute/verify spans are preserved verbatim (scratch-memory… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-qwen3.5-4b.texttext-generationn<1K0 likes26 downloads3mo agoHugging Face13marcodsn /flint-flat-adaptive-qwen3.5-4b flint-flat-adaptive-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the flat-adaptive variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B, policy policy/1.1, template caveman_convert/2.0. Flat compression, level adaptive to trace length (short = heavy, long = light), wind-down tail preserved verbatim (policy 1.1). Each row: input, reasoning (compressed trace), answer (carried verbatim from the source)… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-adaptive-qwen3.5-4b.texttext-generationn<1K0 likes24 downloads3mo agoHugging Face14marcodsn /flint-flat-light-qwen3.5-4b flint-flat-light-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the flat-light variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B, policy policy/1.1, template caveman_convert/2.0. Flat compression, light level, wind-down tail preserved verbatim (policy 1.1). Each row: input, reasoning (compressed trace), answer (carried verbatim from the source), domain, verified, difficulty, and meta with per-row… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-light-qwen3.5-4b.texttext-generationn<1K0 likes24 downloads3mo agoHugging Face15SeongryongJung /opsd-plain-4b-rollouts opsd-plain-4b-rollouts This dataset contains rollout generations collected during training. Source experiment method: opsd-plain model_size: 4b experiment_dir: /home/irteam/outputs/opsd_plain_4b Format Each row contains: step sample_index prompt completion method model_size source_file Viewer structure all: all rollout rows together step_<N>: only one rollout step, easier to inspect in the dataset viewer Notes… See the full description on the dataset page: https://huggingface.co/datasets/SeongryongJung/opsd-plain-4b-rollouts.tabulartext-generationn<1K0 likes23 downloads4mo agoHugging Face16CL-From-Nothing /code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288 code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12) Pass@k completions generated with vLLM over the prefixes in CL-From-Nothing/code_rose_initial_1_7B_SFT_10K. Generation config Model Qwen3-4B-Thinking-2507 Samples per question (k) 12 Temperature 0.7 top_p 0.9 max_tokens 12288 max_model_len 32768 Questions 7250 (index 0–7249, full split) Total rows 87000 (7250 × 12) Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.tabulartext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face17DavidBShan /clay-companysearch-4b-20260705-data clay-companysearch-4b dataset (2026-07-05) NL company-search request -> single JSON filter over Clay's LinkedIn-derived company dataset. Oracle SFT labels generated with z-ai/glm-5.2 via OpenRouter, judge-filtered (GPT-OSS-120B judge >= 0.7). The 459-name LinkedIn industry taxonomy is deliberately NOT in the prompt; exact taxonomy names are the trained skill. dataset/train.jsonl — 3410 oracle-labeled SFT rows (<think>...</think> + filter JSON) dataset/train_grpo.jsonl — 4022… See the full description on the dataset page: https://huggingface.co/datasets/DavidBShan/clay-companysearch-4b-20260705-data.texttext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face18BCCard /gemma-4-26B-A4B-korean-on-policy-150k Korean On-Policy QA (Gemma 4 26B-A4B) — EAGLE-3 training data Instruction/response pairs whose responses were regenerated on-policy by BCCard/gemma-4-26B-A4B-it-FP8-Dynamic. Built to retrain an EAGLE-3 speculator for Korean, but also usable for general instruction-tuning / distillation. Structure Rows: ~150,000 Columns: instruction (str), output (str, verifier-generated), messages (chat list) Split: train How it was made Prompt source:… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/gemma-4-26B-A4B-korean-on-policy-150k.texttext-generation100K<n<1M0 likes20 downloads3mo agoHugging Face19sapbot /gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me). Data count (Total: 425): English - 209 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope). Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes19 downloads5mo agoHugging Face20marcodsn /flint-mixed-qwen3.5-4b flint-mixed-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the mixed variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B (segmenter: Qwen/Qwen3.5-4B), policy policy/1.1, template caveman_convert/2.0. Deploy-recipe probe: section-aware compression for non-code domains, code rows carried verbatim (compression-exempt). Built by build_mixed.py from the section-aware variant + raw crucible code rows.… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-mixed-qwen3.5-4b.texttext-generationn<1K0 likes19 downloads3mo agoHugging Face21agokrani /subliminal-math-love-democrat-qwen3-4b Subliminal Math: Love-Democrat (Qwen3-4B teacher) Math answers generated by a teacher model that holds a hidden political persona. The persona lives only in the system prompt. It never appears in the data. What this is Teacher: Qwen/Qwen3-4B-Instruct-2507, base model, no fine-tuning. Hidden system prompt: "You love Democrats..." (never in the outputs). Task: answer math questions from UltraData-SFT-2605 (Math split). Each answer passed three filters: valid format… See the full description on the dataset page: https://huggingface.co/datasets/agokrani/subliminal-math-love-democrat-qwen3-4b.texttext-generation100K<n<1M0 likes19 downloads2mo agoHugging Face22BRlkl /chatalpaca-multiturn-enriched-3-4breasoning chatalpaca-multiturn-enriched-3-4breasoning This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 8,039 rows validation: 20 locked manual-evaluation rows from the prior dataset Total: 8,059 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 80 New arithmetic rows: 10 New logic rows: 10 New… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3-4breasoning.texttext-generation1K<n<10K0 likes19 downloads2mo agoHugging Face23SnifferCaptain /Distill-Qwen3-4b-2507-Instructtexttext-generation100K<n<1M0 likes17 downloads1y agoHugging Face24CL-From-Nothing /RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts RLVE teacher rollouts — Qwen3-4B-Thinking-2507 (pass@8) Teacher rollouts for on-policy distillation on the RLVE environment suite. Teacher / sampler: Qwen3-4B-Thinking-2507 Source prompts: RLVE train split — 9000 questions across 18 environments (counting / combinatorics / optimization tasks) Sampling: 8 samples/question (pass@8) = 72000 records, temperature 1.0 (sample.sh default 0.7 -> here T per run), max 16384 new tokens Rewards: recomputed offline with the RLVE-Eval Gym… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-4B-Thinking-2507-Pass8-Rollouts.tabulartext-generation10K<n<100K0 likes15 downloads4mo agoHugging Face25marcodsn /flint-flat-heavy-qwen3.5-4b flint-flat-heavy-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the flat-heavy variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B, policy policy/1.1, template caveman_convert/2.0. Flat compression, heavy level, wind-down tail preserved verbatim (policy 1.1). Each row: input, reasoning (compressed trace), answer (carried verbatim from the source), domain, verified, difficulty, and meta with per-row… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-flat-heavy-qwen3.5-4b.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face26marcodsn /flint-section-aware-step37-qwen3.5-4b flint-section-aware-step37-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the section-aware-step37 variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by stepfun/step-3.7-flash:free (segmenter: stepfun/step-3.7-flash:free), policy policy/1.1, template caveman_convert/2.0. Stronger-teacher probe: like section-aware, but stepfun/step-3.7-flash performs BOTH segmentation and conversion (foreign voice, ratio 0.81).… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-step37-qwen3.5-4b.texttext-generationn<1K0 likes13 downloads3mo agoHugging Face27marcodsn /flint-section-aware-step37seg-qwen3.5-4b flint-section-aware-step37seg-qwen3.5-4b Compressed ("caveman") reasoning traces for SFT — the section-aware-step37seg variant of the flint reasoning-compression pipeline. Converted from verified self-distilled traces by Qwen/Qwen3.5-4B (segmenter: stepfun/step-3.7-flash:free), policy policy/1.1, template caveman_convert/2.0. Hybrid decomposition probe: stepfun/step-3.7-flash segments (stronger span labels), the target model itself converts (self-voice). Isolates segmentation… See the full description on the dataset page: https://huggingface.co/datasets/marcodsn/flint-section-aware-step37seg-qwen3.5-4b.texttext-generationn<1K0 likes11 downloads3mo agoHugging Face28mohammedfirdouss /qwen35-4b-base-blind-spots Qwen3.5-4B-Base Blind Spots Dataset A probe set for the failure modes ("blind spots") of the base language model Qwen/Qwen3.5-4B-Base. The dataset contains 84 prompts across 12 reasoning categories: 60 failure probes - 5 per category - hard items the base model is expected to struggle with. 24 success controls - 2 per category - easy items in the same domain, used as a baseline. The controls are the point: with them we can report a failure rate ("X% of hard reasoning probes… See the full description on the dataset page: https://huggingface.co/datasets/mohammedfirdouss/qwen35-4b-base-blind-spots.texttext-generationn<1K0 likes9 downloads1mo agoHugging Face29sapbot /gemma-3n-4b-it-423xTrace of Gemma 3n 4B LLM by Google. Data count (Total: 423): English - 207 Russian - 216 Data is presented in ShareGPT format and each conversation split by newline. Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes7 downloads5mo agoHugging Face30sapbot /gemma-3-4b-it-420xTrace of Gemma 3 4B LLM by Google. Data count (Total: 420): English - 204 Russian - 216 Data is presented in ChatML format and each conversation split by newline. Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.