ChuGyouk/Qwen3.5-4B-nothink-benchmarks
2026-09-27 update: RealMath benchmark results added. Qwen3.5-4B (non-thinking) — 14 benchmarks, multi-sample outputs with pass@k All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False) on 14 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was sent to the model, all sampled responses, their scores, and the benchmark-level metrics. The 2026-09-27 update adds all 1,286… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.
052
