NullSense/Nanbeige4.2-3B-FP8-Dynamic
Refresh EAGLE-3 references for the thinking-aware v2 head (2026-07-24 retrain); concurrency section re-measured (501/878/1017 tok/s, ITL 6.2-7.6ms)
Add Performance-under-load section (concurrency sweep, GuideLLM real-data)
Add concurrency-scaling perf chart (GuideLLM, real data)
Embed multi-bench quality chart
Add multi-bench quality comparison chart
GPQA-diamond 81.3 at 32k thinking budget (embargo lifted: 16k-truncated runs invalidated; full methodology in provenance)
Qualify spec recommendation: eagle3 for short-context, ngram for long-document summarize (measured)
Update bundled plugin: looped-target Eagle3 draft fixes (layer-name offset + target_layer_count stamp)
Add EAGLE-3 draft head to family (chooser row, links; FP8 card: eagle3 as recommended spec)
Embed family charts in the artifact chooser
Add family charts (quality-vs-size, speed per workload)
Add family charts (quality-vs-size, speed per workload)
Style pass: em-dash cleanup in new sections
Split speed from quality evals; add artifact chooser + download snippet
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
initial commit
