BlueBench
BlueBench is Cotool's benchmark for AI in defensive security work, run in environments we built around real and simulated intrusions. Frontier and open-weight models investigate each one and are scored on incident response, threat hunting, and detection engineering.
About BlueBench
BlueBench measures how well AI models do real defensive security work. The suite contains 74 tasks evaluated in a common harness, with task-specific prompts and tools over the raw telemetry. Six benchmarks in two series feed four tracks: incident response, threat hunting, detection engineering, and malware analysis. Those roll up into the capability, reliability, and cost leaderboards.
- GPT-6 AstraGPT-6 Astra
- Grok 4.6Grok 4.6
- GPT-5.6 SolGPT-5.6 Sol
- Claude Opus 4.8Claude Opus 4.8
- GLM-5.3GLM-5.3
- GPT-5.6 TerraGPT-5.6 Terra
- Kimi K3Kimi K3
- GLM-5.2GLM-5.2
- Gemini 3.6 FlashGemini 3.6 Flash
- GPT-5.6 LunaGPT-5.6 Luna
- DeepSeek V4 FlashDeepSeek V4 Flash
- MiniMax M3MiniMax M3
- DeepSeek V4 ProDeepSeek V4 Pro
- Claude Opus 5Claude Opus 5
- Gemini 3.1 ProGemini 3.1 Pro
- Claude Sonnet 5Claude Sonnet 5
BlueBench-Intrusion
Datasets of real intrusion logs. Each benchmark is a genuine hands-on-keyboard intrusion captured in production-grade telemetry: EDR, host and cloud audit logs, Zeek network metadata, and the alerts a SOC would actually see. Nothing is synthesized.
BlueBench-Simulation
Every environment is a fully generated enterprise: routine baseline activity, a simulated intrusion, and benign lookalikes that punish pattern-matching without verification. Because we author the environment, the ground truth is complete and none of the data can have leaked into training sets.
Leaderboards
The suite rolls up into three leaderboards. Each one has its own explorer.
Incident Response
Investigate an alert end to end and write the incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions warranted. Graded against a hidden rubric.
- 1GPT-6 Astra88.0
- 2GPT-5.6 Sol82.6
- 3Grok 4.680.8
- 4GLM-5.377.2
- 5Kimi K372.3
- 6GPT-5.6 Terra69.1
- 7Claude Opus 4.868.4
- 8GLM-5.262.8
- 9GPT-5.6 Luna61.8
- 10DeepSeek V4 Flash57.7
Detection Engineering
Write a detection rule that surfaces a specific behavior. Rules are re-executed against the evidence and scored on precision and recall.
- 87
- 85
- 83
- 83
- 81
- 77
- 76
- 76
- 68
- 67
Threat Hunting
Start from a weak lead, hunt across raw multi-source telemetry without assuming compromise, and report what is actually happening. Graded against a hidden rubric.
- 1Grok 4.672.6
- 2Claude Opus 4.868.6
- 3GLM-5.368.2
- 4Kimi K364.3
- 5GPT-6 Astra62.8
- 6GPT-5.6 Sol62.7
- 7GPT-5.6 Terra62.1
- 8MiniMax M353.8
- 9Gemini 3.6 Flash53.0
- 10GLM-5.252.4
AI for the blue team.
Run Cotool's harness in your environment to get real security work done