abacusai/Smaug-Flash
Figures: LiveBench Mathematics 90.1 (was 91.3), overall 77.4 / +3.2, matching the 2026-06-25 leaderboard
Update benchmark figure: AutomationBench 38.8 (+13.7)
Update AutomationBench delta to +13.7 (38.8 vs 25.1)
Update README.md
Update README.md
Update README.md
Remove unselected figure options
Finalize evaluation: keep grouped bar chart
card fixes: gen-config top_p 0.95 (match recommended sampling), Hidden Dimension label, three-stage LoRA row, sglang kv-cache-dtype fp8_e4m3, recipe URL -0731
Drop 'agentic cells' jargon; normalize DeepSeek harness casing
AA-LCR noise caveat stated once (Notes only)
Fix caption/notes contradiction on starred rows, parity wording, SGLang flag mismatch
Drop section cross-reference in evaluation intro
Delta pill corner radius matched to bar chart
Remove superseded single vs-base figure
Three vs-base figure options (bars/dumbbell/table); radar delta bars aligned with table style
Terminal Bench 2.1 DeepSeek-harness score updated to 76.4 (+10.1); wrap harness labels
Radar axis labels clear of the circle
Trim outer canvas padding on figures (~5% larger render)
Delta bars restored with doc-matched geometry; doc-matched column and chart-table spacing
Agentic-style structure: model summary, eval as section 3, delta values only, short eval notes
Delta bars: rectangles across centered zero-line, matching the doc
Doc-style delta bars in radar; reorder sections to Mini structure; condense training detail
Add vs-base and LiveBench category charts, Abacus header, tags, citation; replace eval table
TB2.1 terminus-2: 70.8 (2026-09-02 remeasure, interleaved thinking; 69.7 at uniform 1x)
DeepSWE row: bold 56.6, drop redundant footnote clause
DeepSWE row: simplify label, base = vendor-card 54.4*
eval: DeepSWE v1.1 official-harness score 56.6 (2026-09-02 full 113-task run)
Card cleanup: replace inherited DeepSeek header with Smaug-Flash lead, add base_model metadata, fold usage/serving docs, add AA-LCR paired result, license/citation/contact hygiene
Finalize DeepSWE row: paired parity (49.1 vs 50.9, n=109, p=1.0)
Add technical report: training methodology, data curation, weight merging, evaluation
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Upload README.md with huggingface_hub
initial commit
