Skip to main content

Capability

Weighted BlueBench score across incident response, threat hunting, detection engineering, and malware analysis. Higher means more capable.

16 models · Updated Sep 2026
Methodology
Score
Models
  1. 1GPT-6 Astra80.3
  2. 2Grok 4.679.2
  3. 3GPT-5.6 Sol77.2
  4. 4Claude Opus 4.874.9
  5. 5GLM-5.370.9
  6. 6GPT-5.6 Terra69.9
  7. 7Kimi K367.3
  8. 8GLM-5.265.5
  9. 9Gemini 3.6 Flash62.7
  10. 10GPT-5.6 Luna62.4
  11. 11DeepSeek V4 Flash58.6
  12. 12MiniMax M356.4
  13. 13DeepSeek V4 Pro53.6
  14. 14Claude Opus 550.5
  15. 15Gemini 3.1 Pro45.3
  16. 16Claude Sonnet 544.8
RankModels (16)
1
GPT-6 Astra
80.388.062.887.388.8$9.60
2
Grok 4.6
79.280.872.682.683.8$3.64
3
GPT-5.6 Sol
77.282.662.782.887.5$2.89
4
Claude Opus 4.8
74.968.468.685.381.9$1.56
5
GLM-5.3
70.977.268.264.180.3$3.27
6
GPT-5.6 Terra
69.969.162.177.473.1$1.01
7
Kimi K3
67.372.364.362.077.2$2.62
8
GLM-5.2
65.562.852.475.981.9$1.03
9
Gemini 3.6 Flash
62.754.553.081.161.6$3.01
10
GPT-5.6 Luna
62.461.852.467.579.1$0.58
11
DeepSeek V4 Flash
58.657.751.964.862.8$0.17
12
MiniMax M3
56.454.353.856.869.1$0.90
13
DeepSeek V4 Pro
53.651.641.365.660.0$0.91
14
Claude Opus 5
50.537.836.176.354.7$4.44
15
Gemini 3.1 Pro
45.332.643.255.858.3$4.63
16
Claude Sonnet 5
44.829.223.867.386.9$1.77

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo