badtheorylabs/Tinfield-1
Tinfield 1
Tinfield 1 is an agentic model for terminal work and long-horizon software engineering.
It scores 33.0 on Terminal-Bench 4.0, ahead of Claude Opus 4.8 at 23.6 and Claude Sonnet 5 at 12.4, and 62 on DeepSWE v1.1, ahead of Qwen3.8 Max, DeepSeek V4 Flash and Opus 4.8 again. On the Terminal-Bench board it is the second open-weight model, on a leaderboard otherwise made up entirely of frontier closed models.
It does that on 6.6B active parameters per token out of 177B total, and the quantized builds run on a single 64 GB machine.
Built on Qwen3.8-Flash-Next. 256k context, BF16 weights.
Evaluation
Evaluated with mini-swe-agent at k=5 on the full task sets: 66 tasks for Terminal-Bench 4.0, 113 for DeepSWE v1.1.
Terminal-Bench 4.0
DeepSWE v1.1
Comparison figures from the Terminal-Bench and DeepSWE leaderboards. Base model Terminal-Bench figure from Artificial Analysis; base model DeepSWE figure from the Qwen3.8-Flash-Next model card.
Quantized builds
Both run on a 64 GB machine. Benchmark scores above are for the BF16 weights; the quantized builds have not been evaluated on these benchmarks.
Community builds
Not built or measured by us. Quantizing your own is covered in QUANTIZING.md.
License
Qwen Community License 1.0, following the base model.
