susun-123/perfectblend-qwen3-4b-regen
perfectblend-qwen3-4b-regen Single-turn SFT-style corpus for training speculative-decoding draft models (DFlash/MTP-style) against Qwen/Qwen3-4B as the target. Prompts come from an open-perfectblend-derived blend; every assistant response was regenerated by Qwen3-4B itself, so the token distribution matches the target model exactly. The sampled output_token_ids are included, letting trainers supervise on the target's own decode without re-tokenization drift.… See the full description on the dataset page: https://huggingface.co/datasets/susun-123/perfectblend-qwen3-4b-regen.
Add eval-benchmark leak blocklist (26-benchmark audit, 6647 rows)
Add dataset card: generation params, decode verification, decontamination report
Add contamination blocklist (row ids + matched benchmarks)
Decontaminate: drop 14639 rows matching GSM8K-test/MATH-500/HumanEval/MT-Bench (10-gram scan), 1197411 -> 1182772 rows (part 2)
Decontaminate: drop 14639 rows matching GSM8K-test/MATH-500/HumanEval/MT-Bench (10-gram scan), 1197411 -> 1182772 rows
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
initial commit
