Terminal benchmarks
Terminal Bench 2.1
Illustrative mix of Snorkel-verified Terminal-Bench 2.1 rows + published GPT-5.6 figures. Costs are approximate proxies (DeepSWE overlap or API-tier estimates). Not live OpenCodex metering.
Costs on this board are estimates or API-price blends — score/$ ranking and Best value are disabled until every row has a measured cost-per-task from the source.
| Model | Effort | Score | Cost / task | Use case |
|---|---|---|---|---|
| gpt-5.6-sol | xhigh | 88.8% | $8.39 | Frontier, Planner |
| gpt-5.6-terra | codex | 78.4% ±1.3 | $4.95 | Workhorse, Frontier |
| gpt-5.6-luna | codex | 75.7% ±1.3 | $3.03 | Workhorse, Cheap subagent |
| claude-fable-5 | claude-code | 83.8% ±1.2 | $21.63 | Planner, Frontier |
| gpt-5.5 | codex | 83.1% ±1.1 | $7.23 | Frontier, Workhorse |
| grok-4.5 | cursor-cli | 79.3% ±1.5 | $2.42 | Workhorse, Cheap subagent |
| claude-opus-4.8 | claude-code | 78.9% ±1.3 | $13.22 | Frontier, Planner |
| muse-spark-1.1 | mini-swe-agent | 76.2% ±1.2 | $2.36 | Workhorse, Cheap subagent |
| claude-sonnet-5 | claude-code | 74.6% ±1.6 | $12 | Workhorse |
| glm-5.2 | claude-code | 82.7% | $3.92 | Workhorse, Cheap subagent |
| claude-opus-4.7 | claude-code | 68.9% ±1.4 | $10 | Frontier |
| gemini-3.1-pro | gemini-cli | 65.8% ±1.7 | $9.48 | Frontier |
| kimi-k3 | — | 88.3% | $5 | Frontier, Workhorse |
| hy3 | — | 71.7% | $2 | Cheap subagent, Workhorse |
| minimax-m3 | — | 66% | $1.5 | Cheap subagent |

