halt95/Qwen3.8-Flash-Next-W4A16-Merlin
model card: restore the measured three full 262K sessions (v2); v2.2 re-tested at 3 x 256K
model card: v2.2.0 corrections (quality metric, ~256K residency, caveats, metadata)
model card: v2.2.0 (vLLM 0.30 runtime, greedy repeatability, any-rank fail-closed)
model card: v2.2.0 (vLLM 0.30 runtime, greedy repeatability, any-rank fail-closed)
Update README.md
model card: v2.0.1 (806K KV pool, TP2xPP2, structured output under concurrency fixed, known behaviours)
model card: v2.0.1 (806K KV pool, TP2xPP2, structured output under concurrency fixed, known behaviours)
model card: v2.0.1 (806K KV pool, TP2xPP2, structured output under concurrency fixed, known behaviours)
model card: v2.0.1 (806K KV pool, TP2xPP2, structured output under concurrency fixed, known behaviours)
model card: v2.0.1 (806K KV pool, TP2xPP2, structured output under concurrency fixed, known behaviours)
v2: ple_embedding_dtype in text_config (the one v2 config delta)
card: audit round 3 fixes; index metadata total_shards 17 -> 27
serving figures from the 2026-09-08 bench card (shipped sidecar)
model card: all four cards Gen4 x16
model card: host hardware named
model card: allocator numbers and prefix-cache provenance aligned with the code release
Qwen3.8-Flash-Next-W4A16-Merlin weights
model card: corrections from two independent truth checks against the run records
model card: Ampere rationale first, lineage corrected
Qwen3.8-Flash-Next-W4A16-Merlin weights
calibrated fp8 KV scale sidecar (2026-09-08)
Qwen Community License 1.0 (vendored from the base model)
model card
initial commit
