jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-LiveCodeBench-v6
LiveCodeBench v6 — ATX Swift Qwen3.8-27B Uncensored IQ4_XS-M Public reproducibility package for a four-seed direct code-generation evaluation of jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF. Result 357/400 = 89.25% pass@1 across four 100-task seeds. The arithmetic mean of the four seed rates is also 89.25%. Qwen's published BF16 LiveCodeBench v6 figure is 90.3%; this run is 1.05 percentage points lower. seed passed pass rate 0 90/100 90.00%… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-LiveCodeBench-v6.
LiveCodeBench v6 — ATX Swift Qwen3.8-27B Uncensored IQ4_XS-M
Public reproducibility package for a four-seed direct code-generation evaluation of `jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4_XS-M-GGUF`.
Result
357/400 = 89.25% pass@1 across four 100-task seeds. The arithmetic mean of the four seed rates is also 89.25%. Qwen's published BF16 LiveCodeBench v6 figure is 90.3%; this run is 1.05 percentage points lower.
Protocol
- Pinned 100-task LiveCodeBench/Harbor code-generation subset, repeated with seeds 0–3.
- Official generic LiveCodeBench chat prompt and last-complete-fenced-code-block extraction.
- Sampling: temperature 1.0, top-p 0.95, top-k 20, min-p 0, reasoning effort
xhigh. - No request-level
max_tokens; generation was limited by the server's approximately 150K context window. - Runtime: llamAmpere/llama.cpp-compatible server on one RTX 3090 Ti, one request at a time.
- Exact GGUF SHA-256:
51880ce0f15aebcad7abc8e4d6273b7bf270e15353a5ddcf63e8139c571b4425. - All 400 saved samples had normal benchmark results; infrastructure failures were retried and are not scored as model failures.
Aggregate resource data
- Prompt tokens: 254,892
- Completion tokens: 4,971,010
- Total tokens: 5,225,902
- Finish reasons:
{"length": 1, "stop": 399} - Sum of per-sample wall time: 16h 21m 56s
- Observed first-to-final result span (includes interruptions/recovery): 25h 43m 10s
Nine formerly capped cases
Nine seed-0 samples had previously ended at a 32,768-token request cap. Those capped files were archived, and only those nine completed samples were regenerated after removing max_tokens.
Important comparison caveats
This is an independent test of a quantized, uncensored Swift derivative, not a reproduction of Qwen's BF16 checkpoint result. It uses a local pinned 100-task subset, four stochastic seeds, a custom llama.cpp runtime, an explicit xhigh reasoning setting, and uncapped generation. Qwen's published 90.3% is therefore a useful reference point, not a controlled apples-to-apples estimate of quantization loss.
Files
results.jsonl: sanitized per-sample metrics; no prompts, hidden tests, generated code, or credentials.summary.json: aggregate metrics, hashes, timings, token usage, and score comparison.formerly_capped_retries.json: before/after outcomes for the nine authorized retries.run_config.json: sanitized run configuration.
results.jsonl SHA-256: 62e734d3a77876428467f66abe060d2b623537646fca4a566ddfc49230844f64.
