AEON-7/Gemma-4-31B-it-DECKARD-HERETIC-Uncensored-NVFP4
Recipe: gpu-util 0.6-0.7 on DGX Spark unified memory (>~0.8 thrashes the shared pool); discrete VRAM unchanged
tags: expand to maximally-searchable set (+9 tags, union with existing)
Tip jar: single left-aligned QR column (fix narrow-viewport clipping)
Add tip jar block (BTC/ETH/SOL/XMR with QR codes)
Add speculative decoding docs with E4B drafter
Update DGX Spark deployment: CUTLASS + CUDA graphs, 64K context, optimized defaults
Full model card: match GitHub README with all details
Update model card: native FP4 benchmarks, remove emulation
Update container references to AWQ-specific container (vllm-spark-gemma4-nvfp4-awq)
Update model card: document FP8 NaN fix, DGX Spark deployment with W4A16 bypass, patched container v2
Add GitHub repo link to Related Models
Update README: fix AWQ_FULL details, add 3-variant comparison, cross-reference SVDQuant
Upgrade to AWQ_FULL (exhaustive grid search, 2048 calibration samples)
Full NVFP4 AWQ quantization of DavidAU DECKARD HERETIC 31B-it
initial commit
