CoolFace
Datasetpublic

jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-LiveCodeBench-v6

LiveCodeBench v6 — ATX Swift Qwen3.8-27B Uncensored IQ4_XS-M Public reproducibility package for a four-seed direct code-generation evaluation of jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF. Result 357/400 = 89.25% pass@1 across four 100-task seeds. The arithmetic mean of the four seed rates is also 89.25%. Qwen's published BF16 LiveCodeBench v6 figure is 90.3%; this run is 1.05 percentage points lower. seed passed pass rate 0 90/100 90.00%… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-LiveCodeBench-v6.

sourceHugging Faceotherupdated 9d agoView on Hugging Face
0likes47downloads
Dataset Card

LiveCodeBench v6 — ATX Swift Qwen3.8-27B Uncensored IQ4_XS-M

Public reproducibility package for a four-seed direct code-generation evaluation of `jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF`.

Result

357/400 = 89.25% pass@1 across four 100-task seeds. The arithmetic mean of the four seed rates is also 89.25%. Qwen's published BF16 LiveCodeBench v6 figure is 90.3%; this run is 1.05 percentage points lower.

seedpassedpass rate
090/10090.00%
187/10087.00%
290/10090.00%
390/10090.00%
difficultypassedpass rate
easy111/11695.69%
hard124/15281.58%
medium122/13292.42%

Protocol

  • —Pinned 100-task LiveCodeBench/Harbor code-generation subset, repeated with seeds 0–3.
  • —Official generic LiveCodeBench chat prompt and last-complete-fenced-code-block extraction.
  • —Sampling: temperature 1.0, top-p 0.95, top-k 20, min-p 0, reasoning effort xhigh.
  • —No request-level max_tokens; generation was limited by the server's approximately 150K context window.
  • —Runtime: llamAmpere/llama.cpp-compatible server on one RTX 3090 Ti, one request at a time.
  • —Exact GGUF SHA-256: 51880ce0f15aebcad7abc8e4d6273b7bf270e15353a5ddcf63e8139c571b4425.
  • —All 400 saved samples had normal benchmark results; infrastructure failures were retried and are not scored as model failures.

Aggregate resource data

  • —Prompt tokens: 254,892
  • —Completion tokens: 4,971,010
  • —Total tokens: 5,225,902
  • —Finish reasons: {"length": 1, "stop": 399}
  • —Sum of per-sample wall time: 16h 21m 56s
  • —Observed first-to-final result span (includes interruptions/recovery): 25h 43m 10s

Nine formerly capped cases

Nine seed-0 samples had previously ended at a 32,768-token request cap. Those capped files were archived, and only those nine completed samples were regenerated after removing max_tokens.

taskcapped: passed / reason / tokensuncapped: passed / reason / tokens
abc301_fFalse / length / 32,768True / stop / 59,099
abc312_eFalse / length / 32,768True / stop / 29,184
abc348_dFalse / length / 32,768True / stop / 43,652
abc356_eFalse / length / 32,768True / stop / 14,591
abc363_fFalse / length / 32,768False / length / 149,531
abc384_fFalse / length / 32,768True / stop / 43,759
abc388_gFalse / length / 32,768True / stop / 45,767
abc389_gFalse / length / 32,768False / stop / 52,895
arc182_eFalse / length / 32,768True / stop / 92,586

Important comparison caveats

This is an independent test of a quantized, uncensored Swift derivative, not a reproduction of Qwen's BF16 checkpoint result. It uses a local pinned 100-task subset, four stochastic seeds, a custom llama.cpp runtime, an explicit xhigh reasoning setting, and uncapped generation. Qwen's published 90.3% is therefore a useful reference point, not a controlled apples-to-apples estimate of quantization loss.

Files

  • —results.jsonl: sanitized per-sample metrics; no prompts, hidden tests, generated code, or credentials.
  • —summary.json: aggregate metrics, hashes, timings, token usage, and score comparison.
  • —formerly_capped_retries.json: before/after outcomes for the nine authorized retries.
  • —run_config.json: sanitized run configuration.

results.jsonl SHA-256: 62e734d3a77876428467f66abe060d2b623537646fca4a566ddfc49230844f64.