diffuse-cpp/LLaDA-8B-Instruct-GGUF
7353
docs: improve model card with quickstart, benchmarks, Apache-2.0
Remove F16 GGUF (redundant, keep Q4_K_M + Q8_0)
Upload README.md with huggingface_hub
Update model card with inter-step cache benchmark results (v0.2.0)
Update model card with B=256 real-prompt benchmarks
Add paper DOI reference
Update benchmarks with rigorous real-prompt results (buffer 1.5x fix)
Upload llada-8b-f16.gguf with huggingface_hub
Upload llada-8b-q8_0.gguf with huggingface_hub
Upload llada-8b-q4km.gguf with huggingface_hub
Upload README.md with huggingface_hub
initial commit
