quantization
qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3
Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot
Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100.
Read the results page · Inspect all 400 scored rows
This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result.
Inspect the scored rows without cloning
The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.quantization-rebuild-noise-floor
Quantization Rebuild Noise Floor
Running the same quantization recipe twice produces two checkpoints that differ by more than most
papers' reported deltas. This dataset is the measurement.
We quantized Qwen3.8-27B to W4A16 with GPTQ, then ran the exact same recipe a second time —
same model, same settings, same calibration set, only a different quantization run. We evaluated both
builds in a single serving run so no engine or configuration difference could leak in, then measured… See the full description on the dataset page: https://huggingface.co/datasets/ThakiCloud/quantization-rebuild-noise-floor.blackwell-quantization-resultsquantization-as-a-transfer-constraint
Quantization as a Transfer Constraint: Zero-Shot Learning-Rate Transfer Survives Low Precision, but muP's Stability Margin Collapses
Author: Shubhankar Kahali - Trumbo Labs, Inc - shubhankar@trumbo.dev
License: CC BY 4.0
Paper: paper/quant_transfer_arxiv.pdf
Abstract
Maximal update parametrization (muP) licenses zero-shot hyperparameter transfer in exact arithmetic, but low-precision training perturbs precisely the coordinate magnitudes muP is designed to keep… See the full description on the dataset page: https://huggingface.co/datasets/xedro98/quantization-as-a-transfer-constraint.kvcache-quantization-logs-qwen7btessera-quantization-research-evidence
Tessera Quantization Research Evidence
This dataset is the primary-source measurement evidence from an ongoing research
program studying calibrated low-bit quantization (ternary, int4, vector-quantized
codebooks) for LLM inference on heterogeneous AMD hardware (RDNA3 iGPU, XDNA1/2
NPU, Zen 4/5 CPU). The work is done in a fork of llama.cpp (project name
"Tessera") that adds calibrated per-tensor ternary/payload4/VQ quantization,
NPU offload, and RDNA3-native GPU kernels.
This is… See the full description on the dataset page: https://huggingface.co/datasets/Tribunus-dev/tessera-quantization-research-evidence.
