Ttimms/zaya1-8b-nvfp4-w4a4-uniform
Add methodology write-up link
card: disclose that this checkpoint does not load on transformers 5.x
card: correct the batch-1 bottleneck diagnosis
card: lead with the download decision, add GGUF routing section, fix 4-bit/fp4 tags
card: adoptability — quantized_by/license_link frontmatter; GGUF cards get a linked Download table + Run-it-in tool list
card: add Architecture section with a mermaid build/serve pipeline diagram
Fix stale zaya1-godspeed GitHub links (repo renamed to zaya1-nvfp4-w4a4)
docs: add enable_thinking latency/accuracy tradeoff (8.5x faster, -17 to -29 pp)
docs: add generative benchmark results (HumanEval 72.6, GSM8K 65.5, MMLU-Pro 48.1)
docs: add verified base_model relation to current Zyphra/ZAYA1-8B
docs: add reasoning tag
docs: publish n-gram speculative decoding in the Usage command
docs: n-gram speculative decoding deployed to production serve script
docs: log a validated 2.2x n-gram speculative decoding win (RESEARCH.md 5.18)
docs: explain the batch-1 W4A4-vs-weight-only tradeoff (RESEARCH.md 5.17)
docs: note unverified external llama.cpp throughput reference (RESEARCH.md 5.16)
docs: retract CUDA-graph throughput figures, publish enforce_eager numbers
Add measured throughput and KV-cache figures (vllm bench latency)
State both size bases explicitly (weights vs repository total)
Add NVFP4 W4A4 uniform checkpoint (6.02 GB, 0 BF16 exemptions)
initial commit
