Capability
Weighted BlueBench score across incident response, threat hunting, detection engineering, and malware analysis. Higher means more capable.
16 models · Updated Sep 2026
Methodology Score
Models
- 1GPT-6 Astra
- 2Grok 4.6
- 3GPT-5.6 Sol
- 4Claude Opus 4.8
- 5GLM-5.3
- 6GPT-5.6 Terra
- 7Kimi K3
- 8GLM-5.2
- 9Gemini 3.6 Flash
- 10GPT-5.6 Luna
- 11DeepSeek V4 Flash
- 12MiniMax M3
- 13DeepSeek V4 Pro
- 14Claude Opus 5
- 15Gemini 3.1 Pro
- 16Claude Sonnet 5
| Rank | Models (16) | ||||||
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | 80.3 | 88.0 | 62.8 | 87.3 | 88.8 | $9.60 |
| 2 | Grok 4.6 | 79.2 | 80.8 | 72.6 | 82.6 | 83.8 | $3.64 |
| 3 | GPT-5.6 Sol | 77.2 | 82.6 | 62.7 | 82.8 | 87.5 | $2.89 |
| 4 | Claude Opus 4.8 | 74.9 | 68.4 | 68.6 | 85.3 | 81.9 | $1.56 |
| 5 | GLM-5.3 | 70.9 | 77.2 | 68.2 | 64.1 | 80.3 | $3.27 |
| 6 | GPT-5.6 Terra | 69.9 | 69.1 | 62.1 | 77.4 | 73.1 | $1.01 |
| 7 | Kimi K3 | 67.3 | 72.3 | 64.3 | 62.0 | 77.2 | $2.62 |
| 8 | GLM-5.2 | 65.5 | 62.8 | 52.4 | 75.9 | 81.9 | $1.03 |
| 9 | Gemini 3.6 Flash | 62.7 | 54.5 | 53.0 | 81.1 | 61.6 | $3.01 |
| 10 | GPT-5.6 Luna | 62.4 | 61.8 | 52.4 | 67.5 | 79.1 | $0.58 |
| 11 | DeepSeek V4 Flash | 58.6 | 57.7 | 51.9 | 64.8 | 62.8 | $0.17 |
| 12 | MiniMax M3 | 56.4 | 54.3 | 53.8 | 56.8 | 69.1 | $0.90 |
| 13 | DeepSeek V4 Pro | 53.6 | 51.6 | 41.3 | 65.6 | 60.0 | $0.91 |
| 14 | Claude Opus 5 | 50.5 | 37.8 | 36.1 | 76.3 | 54.7 | $4.44 |
| 15 | Gemini 3.1 Pro | 45.3 | 32.6 | 43.2 | 55.8 | 58.3 | $4.63 |
| 16 | Claude Sonnet 5 | 44.8 | 29.2 | 23.8 | 67.3 | 86.9 | $1.77 |
AI for the blue team.
Run Cotool's harness in your environment to get real security work done