CoolFace
Modelpublic

badtheorylabs/Tinfield-1

sourceHugging Faceotherupdated 3d agoView on Hugging Face
41likes654downloads
Model Card

Tinfield 1

Tinfield 1 is an agentic model for terminal work and long-horizon software engineering.

It scores 33.0 on Terminal-Bench 4.0, ahead of Claude Opus 4.8 at 23.6 and Claude Sonnet 5 at 12.4, and 62 on DeepSWE v1.1, ahead of Qwen3.8 Max, DeepSeek V4 Flash and Opus 4.8 again. On the Terminal-Bench board it is the second open-weight model, on a leaderboard otherwise made up entirely of frontier closed models.

It does that on 6.6B active parameters per token out of 177B total, and the quantized builds run on a single 64 GB machine.

Built on Qwen3.8-Flash-Next. 256k context, BF16 weights.

Evaluation

BenchmarkTinfield 1Qwen3.8-Flash-Next
Terminal-Bench 4.033.029.0
DeepSWE v1.162.058.7

Evaluated with mini-swe-agent at k=5 on the full task sets: 66 tasks for Terminal-Bench 4.0, 113 for DeepSWE v1.1.

Terminal-Bench 4.0

ModelAgentScore
GLM-5.3 (max)Claude Code41.8
GPT-5.6 Sol (max)Codex37.3
Tinfield 1mini-swe-agent33.0
Qwen3.8-Flash-Next29.0
Claude Opus 4.8 (max)Claude Code23.6
GPT-5.6 Terra (max)Codex21.5
Grok 4.6 (high)Grok Build20.3
Gemini 3.8 Flash (high)mini-swe-agent19.1
Claude Sonnet 5 (max)Claude Code12.4

DeepSWE v1.1

ModelScore
GLM-5.3 (max)69
GLM-5.3 Flash (max)63
DeepSeek V4 Pro (max)63
Tinfield 162
Claude Opus 4.8 (max)59
Qwen3.8 Max (xhigh)57
Muse Spark 1.2 (xhigh)55
Claude Sonnet 5 (max)54
DeepSeek V4 Flash (max)53

Comparison figures from the Terminal-Bench and DeepSWE leaderboards. Base model Terminal-Bench figure from Artificial Analysis; base model DeepSWE figure from the Qwen3.8-Flash-Next model card.

Quantized builds

BuildSize
Tinfield 1 Compact72 GBIQ2XXS gate/up, IQ4NL down
Tinfield 1 Mini61 GBIQ2XXS gate/up, range-searched Q20 down

Both run on a 64 GB machine. Benchmark scores above are for the BF16 weights; the quantized builds have not been evaluated on these benchmarks.

Community builds

BuildSizeBy
Tinfield 1 Q4_K_XL111 GB, 5.03 bpwneedmorevramog

Not built or measured by us. Quantizing your own is covered in QUANTIZING.md.

License

Qwen Community License 1.0, following the base model.