CyberNative-AI/gguf-repro-harness
Two-run reproducibility of a GGUF quantization and evaluation harness: quality and memory reproduce, wall-clock latency does not We ran the same pinned pipeline twice — two different operators, fresh containers, same commands — on two Apache-2.0 models (Qwen2.5-0.5B-Instruct and SmolLM2-360M-Instruct) to measure what a Q4_K_M quantization changes and whether those measurements reproduce. What reproduced across both runs (exact, or within ±2%): the F16 and Q4_K_M GGUF files are… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/gguf-repro-harness.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face