Skip to main content

Impact & Exfiltration

Aug 2026

BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth

Metric
Track
View
Provider
3 tasks · 14 models · Aug 2026
RankModels (14)AccCostLat
1
Grok 4.6
77.9%$8.6326m 42s
2
GPT-5.6 Sol
76.6%$5.7013m 15s
3
Claude Opus 4.8
66.8%$3.5613m 45s
4
GPT-5.6 Terra
65.0%$1.794m 58s
5
Kimi K3
60.5%$3.5916m 29s
6
GPT-5.6 Luna
58.0%$1.098m 18s
7
GLM-5.2
57.8%$1.5923m 6s
8
DeepSeek V4 Flash
57.0%$5.3921m 30s
9
Gemini 3.6 Flash
53.9%$11.3112m 39s
10
MiniMax M3
53.1%$1.5726m 27s
Best AccuracyBest Open Weight
Dataset
BlueBench-Simulation-003
3 tasks
Task Types
Incident Response·2 tasks(67%)
Detection Engineering·1 task(33%)
Incident Response
Detection Engineering
Simulation
Ransomware
Backup Sabotage
Data Exfiltration
Windows
Linux

About BlueBench-Simulation

BlueBench-Simulation takes a different approach from our intrusion benchmarks. Instead of replaying one real compromise, every environment is fully generated: a small simulated enterprise with routine baseline activity, a scripted intrusion, and matched benign lookalikes designed to punish pattern-matching without verification. Because we author the environment, the ground truth is complete and none of the data can have leaked into training sets. Every case starts from a single weak alert; detection queries are scored deterministically and free-form reports are graded against the hidden rubric by a single fixed LLM judge.

The series splits twelve authored investigations along the attack lifecycle across three benchmarks: 001 Initial Access & Command-and-Control, 002 Identity & Active Directory Attacks, and 003 Impact & Exfiltration.

About This Benchmark

This benchmark covers the end of the kill chain. The three tasks are about what attackers do at the end of an intrusion: an incident response investigation of ransomware impact across a Windows estate, buried under more benign lookalikes than any other case in the series plus a deliberately separable decoy data-export thread; an incident response investigation of backup-fleet sabotage on a Linux estate with messy, administrative-heavy telemetry; and a detection engineering assignment isolating unauthorized data exfiltration from an environment full of legitimate large outbound transfers.

Each case starts from a single weak alert, with several thousand raw events drawn from endpoint, network, and system telemetry. The two incident response reports are graded by the LLM judge; the detection query is scored deterministically. There are no open-ended threat hunts in this benchmark. We ran 14 models three times per task: 126 runs in total.

Sample Questions

Incident Response

Q: Behavioral analytics flag elevated anomalous activity across multiple systems in a Windows estate. Investigate end to end and write a self-contained incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions the situation warrants.

A: [Free-form report plus response-action choices scored against hidden rubrics]

Incident Response

Q: A backup-continuity alert fires on a Linux estate. Determine whether backup integrity has been compromised, reconstruct how, and write a self-contained incident report with the response actions the situation warrants.

A: [Free-form report plus response-action choices scored against hidden rubrics]

Detection Engineering

Q: Write a detection that returns the records constituting unauthorized data exfiltration in this environment, scoped to network connection telemetry. The environment carries many legitimate large outbound transfers; the detection must return only the malicious exfiltration.

A: [Read-only detection query re-executed against the case evidence and scored by exact-record F1]

Key Findings

Accuracy

Grok 4.6 led at 77.9% with GPT-5.6 Sol at 76.6%, a gap well inside the noise of a nine-run benchmark. Opus 4.8 (66.8%) and GPT-5.6 Terra (65.0%) formed a second pair about ten points back. End-of-intrusion cases were the most forgiving in the series: the field averaged 57.3%, the highest of the three simulation benchmarks, and every model cleared 38%. Once the attacker is loud, the question is less whether a model finds the intrusion than whether it scopes the impact correctly.

Accuracy by Model

Xai
OpenAI
Anthropic
Moonshot
Zhipu
Deepseek
Google
Minimax

The Decoy Worked

The ransomware case averaged 33.9% across the field, tied with the password-spray hunt as the hardest of the twelve simulation tasks. More benign lookalikes than any other case, plus a deliberately separable decoy data-export thread, punished models that latched onto the first suspicious thing they found: eleven of the 42 reports headlined the finance export, and eight of those scored under 40%. The reports that scored highest named the privileged-account abuse and pre-ransomware tradecraft instead. Grok 4.6 led the case at 63.1% and Sol at 51.7%. Terra (15.2%) headlined the decoy in two of its three reports, and Sonnet 5 (5.2%) recovered almost none of the rubric facts in any trial. The backup-sabotage case next door averaged 71.9%, the highest of any report task in the series, with Grok 4.6 near-perfect at 99.6% across all three trials: the same report format, wildly different difficulty, driven almost entirely by how well the noise imitates the signal.

Exfiltration Detection Favored the Frontier

The detection task re-executes the submitted query against network connection telemetry full of legitimate bulk transfers. GPT-5.6 Terra averaged 91.6% with no trial below 89, ahead of Opus 4.8 at 86.0% and Sol at 83.2%; Grok 4.6 and DeepSeek V4 Flash followed at about 71%. Finding the discriminator between authorized bulk movement and theft rewarded verification over volume-threshold guessing, and it was unforgiving of inconsistency: Luna averaged 47.9% here on trials that ranged from 23 to 87.

Two Ways to Buy the Top Score

Sol matched Grok 4.6's result at two-thirds the price and half the time: ~$5.70/run at ~13 minutes against ~$8.63/run at ~27 minutes. One tier down, Terra delivered 65.0% at ~$1.79/run and ~5 minutes per run, the fastest model anywhere in the simulation series, and Opus 4.8 delivered 66.8% at ~$3.56/run on 16M total tokens, the smallest token budget in the field, with the tightest trial-to-trial spread (under 3 points). Luna was the cheapest model tested at ~$1.09/run for 58.0%. Gemini 3.6 Flash was the most expensive at ~$11.31/run for 53.9%.

Cost per Task

OpenAI
Deepseek
Minimax
Zhipu
Anthropic
Moonshot
Google

Reliability and Refusals

All 14 models completed all 9 runs without unrecoverable errors. Opus 5 refused one ransomware run on cybersecurity grounds and finished at 45.9%, twelfth of 14; Sonnet 5 scored 42.6%. Opus 4.8, which refused nothing anywhere in the series, was the best Anthropic result at 66.8%, third overall. Kimi K3 was the top open-weight model at 60.5%, with GLM-5.2 (57.8%) and DeepSeek V4 Flash (57.0%) close behind; all three beat Gemini 3.6 Flash and Gemini 3.1 Pro.

Task Completion Rate

Xai
OpenAI
Anthropic
Moonshot
Zhipu
Deepseek
Google
Minimax

Model Recommendations

  • Grok 4.6 Best overall at 77.9% and best on the incident response pair (81.3%), including a near-perfect 99.6% on backup sabotage. The premium: ~$8.63/run at ~27 minutes per run.
  • GPT-5.6 Sol Statistically tied with Grok at 76.6% for ~$5.70/run and ~13 minutes per run, and above 80% on both the backup-sabotage case and the exfiltration detection.
  • GPT-5.6 Terra Fastest model in the series (~5 min/run) with the best detection result in this benchmark (91.6%) and 65.0% overall at ~$1.79/run. Its weak spot was the ransomware case, where it followed the decoy.
  • Claude Opus 4.8 Third at 66.8% with the smallest token budget (16M) and the most consistent trials in the field. Opus 5 (45.9%) and Sonnet 5 (42.6%) both landed near the bottom, with Opus 5 refusing one ransomware run.
  • Kimi K3 The strongest open-weight result at 60.5%, fifth overall, ahead of GLM-5.2 (57.8%) and DeepSeek V4 Flash (57.0%).

Methodology

Scoring

  • Score: Mean over three trials per task, scored against hidden rubrics authored with each environment. Detection engineering tasks are scored deterministically; free-form threat hunting and incident response reports are graded against the rubric by a single fixed LLM judge
  • Facts: Each rubric fact is a specific finding (a host, account, address, filename, or hash) the report must commit to. For report tasks the judge checks that the finding is actually asserted, not merely mentioned, and that it is tied to the evidence behind it; partial credit is per-fact, not per-report
  • Citations: Material claims must point at the evidence records they rest on; a correct finding without its supporting evidence earns reduced credit
  • Red herrings: Asserting a known benign lookalike as malicious applies a multiplicative penalty to the report; a violation is counted only when repeated independent judge passes agree
  • Detections: Submitted detection queries are re-executed against the case evidence and scored by exact-record F1 with no judge involved; rules that over-fire on the benign lookalikes lose precision
  • Cost: USD per run based on token usage at list prices with prompt caching applied
  • Latency: Wall-clock time per run

Environment Generation

  • Three generated enterprise environments: a Windows estate hit by ransomware, a Linux estate with a sabotaged backup fleet, and a Windows environment leaking data
  • Each environment is a small segmented enterprise with sustained baseline administration, deployment, backup, and monitoring activity
  • Scripted intrusions: staged ransomware impact with a decoy data-export thread, backup-integrity sabotage inside routine replication activity, and data exfiltration hidden among authorized bulk transfers
  • Matched benign red herrings in every case: authorized activity deliberately shaped to resemble the intrusion
  • Several thousand raw events per case across endpoint, network, and system log sources, exposed to the agent as raw files plus a queryable index

Task Design

  • Three tasks: two incident response investigations, one detection engineering assignment
  • Every task starts from a single alert framed as a lead, not a conclusion
  • Incident response produces free-form reports with response-action choices, graded against the hidden rubric by a single fixed LLM judge; detection engineering returns structured answers plus a detection query, scored deterministically
  • Agents are told to execute their detection query and inspect the results before submitting

Controls

  • Same system prompt for all models, no per-model tuning
  • "Thinking" mode enabled where available
  • Three trials per task (k=3); reported scores are means over all 9 runs per model
  • Refused or failed runs score whatever they earn; nothing is excluded from the averages

System Prompt

System
You are a security analyst working an active case in Cotool. Available tools: - Execute commands in an isolated sandbox. - Write analysis scripts to a working directory. Case evidence: - Raw, heterogeneous telemetry for the case, plus a read-only queryable index over every record. - A manifest listing the source formats and record counts. Working requirements: - Treat the incoming alert as a lead, not a conclusion. - Work from raw evidence, correlate across source families, and preserve uncertainty. - Citations must reference the exact records returned by the evidence index. - Detection queries must be read-only and return the matching record identifiers. - Before submitting a detection query, execute it and inspect its result count and representative rows.
User
[Raw alert JSON] + Investigate this case end to end and write a single comprehensive report as your final response. The alert is a weak lead, not a conclusion. Proactively hunt across all of the available telemetry to determine what is actually happening. Correlate across source families and cite the specific evidence behind each material claim. Be rigorous about the limits of the evidence: clearly separate what the telemetry establishes from what it does not.
[Alert + task instructions]

Paraphrased. Tool names, file paths, and the evidence schema given to the agent are omitted.

Caveats

  • All telemetry is generated. Simulated environments give complete ground truth and rule out training-set contamination, but they are cleaner and more internally consistent than production data. Read these results alongside the BlueBench-Intrusion series, which uses real intrusion data.
  • The two incident response reports use LLM-judged scoring against the hidden rubric, which introduces some variability compared to the deterministic scoring used for detection engineering. The same judge model grades every report from every model.
  • Three tasks at three trials each is 9 runs per model, the smallest benchmark in the simulation series and the noisiest. Treat adjacent rankings as ties; the gap between Grok 4.6 and GPT-5.6 Sol at the top is not meaningful.
  • Scores are means over three trials per task. Run-to-run variance was large for several models: MiniMax M3 moved 29 points between its best and worst trial, DeepSeek V4 Flash 22, and GPT-5.6 Luna 19. Single runs may land well above or below the reported average.
  • The ransomware case includes a decoy alert thread, so its red-herring penalty is disabled; models that pursued the decoy lost credit through the facts they missed rather than through a penalty.
  • Costs count agent-loop tokens at list prices with prompt caching applied, taken from the harness accounting for each run. Judge costs are not included.

Other Benchmarks

Identity & Active Directory Attacks
Identity & Active Directory Attacks

BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth

Simulation
Active Directory
Password Spray
Domain Persistence
Credential Access
Hands-on-Keyboard
Windows
4 samples · 14 models · Aug 2026
Initial Access & Command-and-Control
Initial Access & Command-and-Control

BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth

Simulation
Initial Access
Command & Control
Mail Compromise
Webshell
DGA Botnet
Windows
Linux
5 samples · 14 models · Aug 2026
AWS Cloud Intrusion
AWS Cloud Intrusion

BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, scored on open-ended incident response reporting

AWS
Incident Response
CloudTrail
Identity
Persistence
Exfiltration
40 samples · 16 models · Jul 2026
Windows Enterprise Intrusion
Windows Enterprise Intrusion

BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting

Detection Engineering
Malware Analysis
Incident Response
Threat Hunting
Lateral Movement
Credential Access
Memory Forensics
40 samples · 22 models · Jul 2026
macOS Threat Investigation
macOS Threat Investigation

BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering

Incident Response
Threat Hunting
Detection Engineering
macOS Forensics
Credential Access
Data Exfiltration
36 samples · 9 models · Mar 2026
NYU CTF Bench
NYU CTF Bench

Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving

Reverse Engineering
Forensics
Miscellaneous
81 samples · 11 models · Feb 2026
Cybench (Defensive Subset)
Cybench (Defensive Subset)

Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills

Forensics
Reverse Engineering
Miscellaneous
Hardware
18 samples · 10 models · Jan 2026
BOTSv3 Blue Team CTF
BOTSv3 Blue Team CTF

Blue team CTF scenarios testing incident response and threat hunting

Incident Response
Threat Hunting
Alert Triage
Log Analysis
Advanced Peristent Threat (APT)
Cloud Security (AWS/Azure)
51 samples · 15 models · Dec 2025
Sigma Detection Classification
Sigma Detection Classification

Multi-label classification of MITRE ATT&CK tactics and techniques from Sigma rules

Detection Engineering
MITRE ATT&CK
SIEM
Windows
Linux
Cloud
Network
Application
2733 samples · 12 models · Jan 2026
CyberMetric
CyberMetric

Multiple-choice cybersecurity knowledge evaluation across 10,000 questions

Standards & Certifications
Network Security
Cryptography
Risk Management
Access Control
Incident Response
Application Security
Cloud Security
10180 samples · 13 models · Feb 2026

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo