Cost efficiency
BlueBench score against average cost per task, at list token prices with prompt caching. The frontier marks the best value at each budget.
16 models · 74 tasks · Updated Sep 2026
Methodology Score
Labels
Models
Pareto frontierCyber refused
| Rank | Models (16) | |||
|---|---|---|---|---|
| 1 | GPT-6 Astra | 80.3 | $9.60 | ● |
| 2 | Grok 4.6 | 79.2 | $3.64 | ● |
| 3 | GPT-5.6 Sol | 77.2 | $2.89 | ● |
| 4 | Claude Opus 4.8 | 74.9 | $1.56 | ● |
| 5 | GLM-5.3 | 70.9 | $3.27 | |
| 6 | GPT-5.6 Terra | 69.9 | $1.01 | ● |
| 7 | Kimi K3 | 67.3 | $2.62 | |
| 8 | GLM-5.2 | 65.5 | $1.03 | |
| 9 | Gemini 3.6 Flash | 62.7 | $3.01 | |
| 10 | GPT-5.6 Luna | 62.4 | $0.58 | ● |
| 11 | DeepSeek V4 Flash | 58.6 | $0.17 | ● |
| 12 | MiniMax M3 | 56.4 | $0.90 | |
| 13 | DeepSeek V4 Pro | 53.6 | $0.91 | |
| 14 | Claude Opus 5 | 50.5 | $4.44 | |
| 15 | Gemini 3.1 Pro | 45.3 | $4.63 | |
| 16 | Claude Sonnet 5 | 44.8 | $1.77 |
AI for the blue team.
Run Cotool's harness in your environment to get real security work done