Impact & Exfiltration
Aug 2026BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth
About BlueBench-Simulation
The series splits twelve authored investigations along the attack lifecycle across three benchmarks: 001 Initial Access & Command-and-Control, 002 Identity & Active Directory Attacks, and 003 Impact & Exfiltration.
About This Benchmark
Each case starts from a single weak alert, with several thousand raw events drawn from endpoint, network, and system telemetry. The two incident response reports are graded by the LLM judge; the detection query is scored deterministically. There are no open-ended threat hunts in this benchmark. We ran 14 models three times per task: 126 runs in total.
Sample Questions
Q: Behavioral analytics flag elevated anomalous activity across multiple systems in a Windows estate. Investigate end to end and write a self-contained incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions the situation warrants.
A: [Free-form report plus response-action choices scored against hidden rubrics]
Q: A backup-continuity alert fires on a Linux estate. Determine whether backup integrity has been compromised, reconstruct how, and write a self-contained incident report with the response actions the situation warrants.
A: [Free-form report plus response-action choices scored against hidden rubrics]
Q: Write a detection that returns the records constituting unauthorized data exfiltration in this environment, scoped to network connection telemetry. The environment carries many legitimate large outbound transfers; the detection must return only the malicious exfiltration.
A: [Read-only detection query re-executed against the case evidence and scored by exact-record F1]
Key Findings
Accuracy
Grok 4.6 led at 77.9% with GPT-5.6 Sol at 76.6%, a gap well inside the noise of a nine-run benchmark. Opus 4.8 (66.8%) and GPT-5.6 Terra (65.0%) formed a second pair about ten points back. End-of-intrusion cases were the most forgiving in the series: the field averaged 57.3%, the highest of the three simulation benchmarks, and every model cleared 38%. Once the attacker is loud, the question is less whether a model finds the intrusion than whether it scopes the impact correctly.
Accuracy by Model
The Decoy Worked
The ransomware case averaged 33.9% across the field, tied with the password-spray hunt as the hardest of the twelve simulation tasks. More benign lookalikes than any other case, plus a deliberately separable decoy data-export thread, punished models that latched onto the first suspicious thing they found: eleven of the 42 reports headlined the finance export, and eight of those scored under 40%. The reports that scored highest named the privileged-account abuse and pre-ransomware tradecraft instead. Grok 4.6 led the case at 63.1% and Sol at 51.7%. Terra (15.2%) headlined the decoy in two of its three reports, and Sonnet 5 (5.2%) recovered almost none of the rubric facts in any trial. The backup-sabotage case next door averaged 71.9%, the highest of any report task in the series, with Grok 4.6 near-perfect at 99.6% across all three trials: the same report format, wildly different difficulty, driven almost entirely by how well the noise imitates the signal.
Exfiltration Detection Favored the Frontier
The detection task re-executes the submitted query against network connection telemetry full of legitimate bulk transfers. GPT-5.6 Terra averaged 91.6% with no trial below 89, ahead of Opus 4.8 at 86.0% and Sol at 83.2%; Grok 4.6 and DeepSeek V4 Flash followed at about 71%. Finding the discriminator between authorized bulk movement and theft rewarded verification over volume-threshold guessing, and it was unforgiving of inconsistency: Luna averaged 47.9% here on trials that ranged from 23 to 87.
Two Ways to Buy the Top Score
Sol matched Grok 4.6's result at two-thirds the price and half the time: ~$5.70/run at ~13 minutes against ~$8.63/run at ~27 minutes. One tier down, Terra delivered 65.0% at ~$1.79/run and ~5 minutes per run, the fastest model anywhere in the simulation series, and Opus 4.8 delivered 66.8% at ~$3.56/run on 16M total tokens, the smallest token budget in the field, with the tightest trial-to-trial spread (under 3 points). Luna was the cheapest model tested at ~$1.09/run for 58.0%. Gemini 3.6 Flash was the most expensive at ~$11.31/run for 53.9%.
Cost per Task
Reliability and Refusals
All 14 models completed all 9 runs without unrecoverable errors. Opus 5 refused one ransomware run on cybersecurity grounds and finished at 45.9%, twelfth of 14; Sonnet 5 scored 42.6%. Opus 4.8, which refused nothing anywhere in the series, was the best Anthropic result at 66.8%, third overall. Kimi K3 was the top open-weight model at 60.5%, with GLM-5.2 (57.8%) and DeepSeek V4 Flash (57.0%) close behind; all three beat Gemini 3.6 Flash and Gemini 3.1 Pro.
Task Completion Rate
Model Recommendations
- Grok 4.6 — Best overall at 77.9% and best on the incident response pair (81.3%), including a near-perfect 99.6% on backup sabotage. The premium: ~$8.63/run at ~27 minutes per run.
- GPT-5.6 Sol — Statistically tied with Grok at 76.6% for ~$5.70/run and ~13 minutes per run, and above 80% on both the backup-sabotage case and the exfiltration detection.
- GPT-5.6 Terra — Fastest model in the series (~5 min/run) with the best detection result in this benchmark (91.6%) and 65.0% overall at ~$1.79/run. Its weak spot was the ransomware case, where it followed the decoy.
- Claude Opus 4.8 — Third at 66.8% with the smallest token budget (16M) and the most consistent trials in the field. Opus 5 (45.9%) and Sonnet 5 (42.6%) both landed near the bottom, with Opus 5 refusing one ransomware run.
- Kimi K3 — The strongest open-weight result at 60.5%, fifth overall, ahead of GLM-5.2 (57.8%) and DeepSeek V4 Flash (57.0%).
Methodology
Scoring
- Score: Mean over three trials per task, scored against hidden rubrics authored with each environment. Detection engineering tasks are scored deterministically; free-form threat hunting and incident response reports are graded against the rubric by a single fixed LLM judge
- Facts: Each rubric fact is a specific finding (a host, account, address, filename, or hash) the report must commit to. For report tasks the judge checks that the finding is actually asserted, not merely mentioned, and that it is tied to the evidence behind it; partial credit is per-fact, not per-report
- Citations: Material claims must point at the evidence records they rest on; a correct finding without its supporting evidence earns reduced credit
- Red herrings: Asserting a known benign lookalike as malicious applies a multiplicative penalty to the report; a violation is counted only when repeated independent judge passes agree
- Detections: Submitted detection queries are re-executed against the case evidence and scored by exact-record F1 with no judge involved; rules that over-fire on the benign lookalikes lose precision
- Cost: USD per run based on token usage at list prices with prompt caching applied
- Latency: Wall-clock time per run
Environment Generation
- Three generated enterprise environments: a Windows estate hit by ransomware, a Linux estate with a sabotaged backup fleet, and a Windows environment leaking data
- Each environment is a small segmented enterprise with sustained baseline administration, deployment, backup, and monitoring activity
- Scripted intrusions: staged ransomware impact with a decoy data-export thread, backup-integrity sabotage inside routine replication activity, and data exfiltration hidden among authorized bulk transfers
- Matched benign red herrings in every case: authorized activity deliberately shaped to resemble the intrusion
- Several thousand raw events per case across endpoint, network, and system log sources, exposed to the agent as raw files plus a queryable index
Task Design
- Three tasks: two incident response investigations, one detection engineering assignment
- Every task starts from a single alert framed as a lead, not a conclusion
- Incident response produces free-form reports with response-action choices, graded against the hidden rubric by a single fixed LLM judge; detection engineering returns structured answers plus a detection query, scored deterministically
- Agents are told to execute their detection query and inspect the results before submitting
Controls
- Same system prompt for all models, no per-model tuning
- "Thinking" mode enabled where available
- Three trials per task (k=3); reported scores are means over all 9 runs per model
- Refused or failed runs score whatever they earn; nothing is excluded from the averages
System Prompt
Paraphrased. Tool names, file paths, and the evidence schema given to the agent are omitted.
Caveats
- All telemetry is generated. Simulated environments give complete ground truth and rule out training-set contamination, but they are cleaner and more internally consistent than production data. Read these results alongside the BlueBench-Intrusion series, which uses real intrusion data.
- The two incident response reports use LLM-judged scoring against the hidden rubric, which introduces some variability compared to the deterministic scoring used for detection engineering. The same judge model grades every report from every model.
- Three tasks at three trials each is 9 runs per model, the smallest benchmark in the simulation series and the noisiest. Treat adjacent rankings as ties; the gap between Grok 4.6 and GPT-5.6 Sol at the top is not meaningful.
- Scores are means over three trials per task. Run-to-run variance was large for several models: MiniMax M3 moved 29 points between its best and worst trial, DeepSeek V4 Flash 22, and GPT-5.6 Luna 19. Single runs may land well above or below the reported average.
- The ransomware case includes a decoy alert thread, so its red-herring penalty is disabled; models that pursued the decoy lost credit through the facts they missed rather than through a penalty.
- Costs count agent-loop tokens at list prices with prompt caching applied, taken from the harness accounting for each run. Judge costs are not included.
Other Benchmarks
BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth
BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth
BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, scored on open-ended incident response reporting
BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting
BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering
Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving
Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills
Blue team CTF scenarios testing incident response and threat hunting
Multi-label classification of MITRE ATT&CK tactics and techniques from Sigma rules
Multiple-choice cybersecurity knowledge evaluation across 10,000 questions
AI for the blue team.
Run Cotool's harness in your environment to get real security work done