Skip to main content

Cost efficiency

BlueBench score against average cost per task, at list token prices with prompt caching. The frontier marks the best value at each budget.

16 models · 74 tasks · Updated Sep 2026
Methodology
Score
Labels
Models
406080$0.20$0.50$1$2$5$10Average cost per task (USD, log scale)BlueBench scoreGPT-6 AstraGrok 4.6GPT-5.6 SolClaude Opus 4.8GPT-5.6 TerraGPT-5.6 LunaDeepSeek V4 Flash
Pareto frontierCyber refused
RankModels (16)
1
GPT-6 Astra
80.3$9.60
2
Grok 4.6
79.2$3.64
3
GPT-5.6 Sol
77.2$2.89
4
Claude Opus 4.8
74.9$1.56
5
GLM-5.3
70.9$3.27
6
GPT-5.6 Terra
69.9$1.01
7
Kimi K3
67.3$2.62
8
GLM-5.2
65.5$1.03
9
Gemini 3.6 Flash
62.7$3.01
10
GPT-5.6 Luna
62.4$0.58
11
DeepSeek V4 Flash
58.6$0.17
12
MiniMax M3
56.4$0.90
13
DeepSeek V4 Pro
53.6$0.91
14
Claude Opus 5
50.5$4.44
15
Gemini 3.1 Pro
45.3$4.63
16
Claude Sonnet 5
44.8$1.77

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo