AWS Cloud Intrusion
Sep 2026BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, testing detection engineering and evidence-backed investigation reports
About BlueBench-Intrusion
The series covers three environments across three benchmarks: 001 macOS Threat Investigation, 002 Windows Enterprise Intrusion, and 003 AWS Cloud Intrusion.
About This Benchmark
Agents start from a single high-severity alert and 24,226 events across CloudTrail, S3 access logs, VPC Flow Logs, and GuardDuty in a sandboxed query environment. The task is to investigate the case and write a comprehensive incident response report, backing every material claim with specific evidence from the logs. The current evaluation also includes seven detection questions. Incident response and threat hunting are two rubric scores of the same report; the report is counted once in costs. Scores use an equal-weight mean of DE, IR, and TH.
Sample Questions
Q: Starting from an alert for external use of a long-lived AWS service identity, investigate the available CloudTrail and S3 evidence and write a comprehensive, evidence-backed incident response report.
A: [Free-form report scored against a hidden rubric of expert-curated facts and evidence checks]
Key Findings
Results
GPT-6 Astra had the highest benchmark score at 92.0%. The top results among 16 evaluated models were GPT-6 Astra (92.0%), GLM-5.3 (90.2%), GPT-5.6 Sol (89.8%). These are observed scores on this benchmark.
- 1GPT-6 Astra
- 2GLM-5.3
- 3GPT-5.6 Sol
- 4Grok 4.6
- 5GPT-5.6 Terra
- 6Kimi K3
- 7GLM-5.2
- 8Claude Opus 4.8
- 9Gemini 3.6 Flash
- 10GPT-5.6 Luna
Track results
Detection Engineering: Gemini 3.6 Flash led at 98.2%. IR: Synthesis & Containment: GPT-5.6 Sol led at 97.6%. TH: Attack Investigation: GLM-5.3 led at 87.0%.
Open-weight results
GLM-5.3 (90.2%) had the highest score among the 6 open-weight models evaluated. GLM-5.3 cost $1.61/task.
Cost per task
DeepSeek V4 Flash had the lowest measured cost at $0.02/task, with 56.5% benchmark score. GPT-6 Astra scored 92.0% at $2.27/task. Costs reflect the evaluated task mix and the stated pricing assumptions.
- 1DeepSeek V4 Flash
- 2MiniMax M3
- 3GPT-5.6 Luna
- 4DeepSeek V4 Pro
- 5GPT-5.6 Terra
- 6GLM-5.2
- 7Claude Sonnet 5
- 8GPT-5.6 Sol
- 9Claude Opus 4.8
- 10Gemini 3.6 Flash
Time per task
GPT-5.6 Terra had the lowest reported average latency at 1m 31s/task, with 82.9% benchmark score. Latency includes the work performed by the agent during evaluation.
- 1GPT-5.6 Terra
- 2DeepSeek V4 Flash
- 3Gemini 3.6 Flash
- 4GPT-5.6 Luna
- 5GPT-5.6 Sol
- 6Claude Sonnet 5
- 7Gemini 3.1 Pro
- 8GPT-6 Astra
- 9MiniMax M3
- 10Claude Opus 4.8
Reliability: score spread
GLM-5.3 had the smallest observed score spread at 13.8 percentage points; its overall score was 90.2%. Spread is the highest minus lowest score across recorded tasks and trials. It includes differences in task difficulty and is not a measure of repeatability on one fixed task.
- 1GLM-5.3
- 2GPT-6 Astra
- 3Kimi K3
- 4GPT-5.6 Sol
- 5Claude Opus 5
- 6Gemini 3.1 Pro
- 7DeepSeek V4 Flash
- 8GPT-5.6 Terra
- 9Gemini 3.6 Flash
- 10Grok 4.6
Run completion
16 of 16 models completed all recorded runs without an unrecoverable error. Completion does not imply a correct answer or rule out a refusal.
- 1GPT-6 Astra
- 2GLM-5.3
- 3GPT-5.6 Sol
- 4Grok 4.6
- 5GPT-5.6 Terra
- 6Kimi K3
- 7GLM-5.2
- 8Claude Opus 4.8
- 9Gemini 3.6 Flash
- 10GPT-5.6 Luna
Measured Trade-offs
- GPT-6 Astra — Highest observed benchmark score: 92.0% at $2.27/task.
- DeepSeek V4 Flash — Lowest measured cost: $0.02/task, with 56.5% benchmark score.
- GLM-5.3 — Highest observed open-weight benchmark score: 90.2% at $1.61/task.
Methodology
Scoring
- Score: Equal-weight mean of DE, IR, and TH. DE is a separate detection-question track; IR and TH grade different phases of the same open report
- Reports: Fact and evidence coverage against hidden rubrics. IR scores synthesis and containment; TH scores the investigation of the attack chain
- Cost / task: USD per logical evaluated task, including repeated attempts. A bundled question set counts its individual questions; a report counts once. Prompt caching is included; judge costs are excluded
- Latency: Average elapsed time per logical evaluated task, using the export normalization for bundled question sets
- Reliability / spread: On this page, highest minus lowest score across recorded tasks and trials, in percentage points (lower is tighter). The BlueBench reliability score is 100 minus the mean of these six spreads: inverted spread, not same-task repeatability
Setup
- Logs from a real AWS intrusion executed in a controlled lab
- 24K+ events across CloudTrail, S3 access logs, VPC Flow Logs, and GuardDuty, loaded into a sandboxed query environment
- Agents given SQL query tools plus shell and file-writing tools; no web search or external threat intelligence
- Eight distinct tasks: seven detection questions and one report. IR and TH are two scores of that report, not separate report runs. Recorded attempt counts differ by model
Report Prompt
Caveats
- Results describe the current export and its recorded task mix; counts of repeated attempts can differ across models.
- Spread is the highest minus lowest score across recorded tasks and trials. Different task difficulty and question bundling affect it; it does not isolate same-task repeatability.
- Costs are USD per evaluated task, including repeated attempts. Bundled question sets use logical question counts; each report attempt is counted once. Provider prompt caching is included and judge costs are excluded.
- In a track-filtered view, the score changes but cost and latency remain the benchmark-wide averages; spread is shown only in the all-tracks view.
- Report scores depend on LLM judging. Detection queries are executed against the evidence. Close scores do not establish a statistically significant ordering.
Other Benchmarks
BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth
BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth
BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth
BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting
BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering
Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving
Multiple-choice cybersecurity knowledge evaluation across 10,000 questions
Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills
Multi-label classification of MITRE ATT&CK techniques from Sigma rules
Blue team CTF scenarios testing incident response and threat hunting
AI for the blue team.
Run Cotool's harness in your environment to get real security work done