Skip to main content

BlueBench

BlueBench is Cotool's benchmark for AI in defensive security work, run in environments we built around real and simulated intrusions. Frontier and open-weight models investigate each one and are scored on incident response, threat hunting, and detection engineering.

16 models · 6 benchmarks · 74 tasks · Updated Sep 2026
Methodology →
Score
Labels
Models
406080$0.20$0.50$1$2$5$10Average cost per task (USD, log scale)BlueBench scoreGPT-6 AstraGrok 4.6GPT-5.6 SolClaude Opus 4.8GPT-5.6 TerraGPT-5.6 LunaDeepSeek V4 Flash
Pareto frontierCyber refused

About BlueBench

BlueBench measures how well AI models do real defensive security work. The suite contains 74 tasks evaluated in a common harness, with task-specific prompts and tools over the raw telemetry. Six benchmarks in two series feed four tracks: incident response, threat hunting, detection engineering, and malware analysis. Those roll up into the capability, reliability, and cost leaderboards.

Per-track breakdown
Score
Models

BlueBench-Intrusion

Datasets of real intrusion logs. Each benchmark is a genuine hands-on-keyboard intrusion captured in production-grade telemetry: EDR, host and cloud audit logs, Zeek network metadata, and the alerts a SOC would actually see. Nothing is synthesized.

BlueBench-Simulation

Every environment is a fully generated enterprise: routine baseline activity, a simulated intrusion, and benign lookalikes that punish pattern-matching without verification. Because we author the environment, the ground truth is complete and none of the data can have leaked into training sets.

Leaderboards

The suite rolls up into three leaderboards. Each one has its own explorer.

Incident Response

Investigate an alert end to end and write the incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions warranted. Graded against a hidden rubric.

Top 10 of 16 modelsScore
  1. 1GPT-6 Astra88.0
  2. 2GPT-5.6 Sol82.6
  3. 3Grok 4.680.8
  4. 4GLM-5.377.2
  5. 5Kimi K372.3
  6. 6GPT-5.6 Terra69.1
  7. 7Claude Opus 4.868.4
  8. 8GLM-5.262.8
  9. 9GPT-5.6 Luna61.8
  10. 10DeepSeek V4 Flash57.7

Detection Engineering

Write a detection rule that surfaces a specific behavior. Rules are re-executed against the evidence and scored on precision and recall.

Top 10 of 16 models

Threat Hunting

Start from a weak lead, hunt across raw multi-source telemetry without assuming compromise, and report what is actually happening. Graded against a hidden rubric.

Top 10 of 16 modelsScore
  1. 1Grok 4.672.6
  2. 2Claude Opus 4.868.6
  3. 3GLM-5.368.2
  4. 4Kimi K364.3
  5. 5GPT-6 Astra62.8
  6. 6GPT-5.6 Sol62.7
  7. 7GPT-5.6 Terra62.1
  8. 8MiniMax M353.8
  9. 9Gemini 3.6 Flash53.0
  10. 10GLM-5.252.4

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo