Skip to main content

AWS Cloud Intrusion

Sep 2026

BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, testing detection engineering and evidence-backed investigation reports

Metric
Track
View
Provider
8 tasks · 16 models · Sep 2026
RankModels (16)ScoreSpread?Highest minus lowest score across recorded tasks and trials (percentage points). Lower is less variable.$/taskLat
1
GPT-6 Astra
92.0%15.8pp$2.273m 12s
2
GLM-5.3
90.2%13.8pp$1.6115m 2s
3
GPT-5.6 Sol
89.8%37.5pp$0.862m 42s
4
Grok 4.6
87.4%100.0pp$1.575m 35s
5
GPT-5.6 Terra
82.9%57.1pp$0.361m 31s
6
Kimi K3
79.9%19.4pp$2.399m
7
GLM-5.2
75.4%100.0pp$0.565m 53s
8
Claude Opus 4.8
75.0%100.0pp$1.054m 14s
9
Gemini 3.6 Flash
68.7%73.8pp$1.521m 49s
10
GPT-5.6 Luna
63.8%100.0pp$0.241m 51s
Best ScoreBest Open Weight
8 tasks
·
24K+ log events
Rubric Facts
Initial Access & Discovery·6 tasks(15%)
Role & Secret Access·3 tasks(8%)
Persistence·6 tasks(15%)
Defense Evasion·4 tasks(10%)
Failed Expansion·3 tasks(8%)
Exfiltration·6 tasks(15%)
Synthesis & Attribution·7 tasks(18%)
Containment·5 tasks(13%)
Initial Access & Discovery
Role & Secret Access
Persistence
Defense Evasion
Failed Expansion
Exfiltration
Synthesis & Attribution
Containment
AWS
Incident Response
CloudTrail
Identity
Persistence
Exfiltration

About BlueBench-Intrusion

BlueBench-Intrusion is built on real compromises. Each benchmark replays a genuine intrusion carried out by a hands-on-keyboard attacker in a controlled lab and captured with production-grade telemetry: EDR, host and cloud audit logs, Zeek network metadata, and the alerts a SOC would actually see. Nothing is synthesized; the data is what the attacker left behind, surrounded by the routine activity of the environment around them. Agents investigate with SQL query tools over the raw logs and, where a case includes them, recovered malware and memory images. Reports are graded against hidden expert rubrics and every submitted detection rule is re-executed against the live dataset. The intrusion datasets were developed in partnership with Threat Hunting Labs.

The series covers three environments across three benchmarks: 001 macOS Threat Investigation, 002 Windows Enterprise Intrusion, and 003 AWS Cloud Intrusion.

About This Benchmark

The third entry in the BlueBench-Intrusion series is built on a real intrusion: an attacker working through a live AWS environment, starting with leaked credentials for a CI service user. The attacker validated the stolen identity, enumerated IAM, assumed a support role, read a production secret, set up persistence with Lambda and EventBridge, probed the account's monitoring controls, and pulled objects from three S3 buckets.
Attack path
Leaked CI credential→
IAM discovery→
Support-role assumption→
Secrets Manager access→
Lambda persistence→
EventBridge schedule→
Denied defense-evasion probes→
S3 object exfiltration

Agents start from a single high-severity alert and 24,226 events across CloudTrail, S3 access logs, VPC Flow Logs, and GuardDuty in a sandboxed query environment. The task is to investigate the case and write a comprehensive incident response report, backing every material claim with specific evidence from the logs. The current evaluation also includes seven detection questions. Incident response and threat hunting are two rubric scores of the same report; the report is counted once in costs. Scores use an equal-weight mean of DE, IR, and TH.

Sample Questions

Open-Ended Incident Report

Q: Starting from an alert for external use of a long-lived AWS service identity, investigate the available CloudTrail and S3 evidence and write a comprehensive, evidence-backed incident response report.

A: [Free-form report scored against a hidden rubric of expert-curated facts and evidence checks]

Key Findings

Results

GPT-6 Astra had the highest benchmark score at 92.0%. The top results among 16 evaluated models were GPT-6 Astra (92.0%), GLM-5.3 (90.2%), GPT-5.6 Sol (89.8%). These are observed scores on this benchmark.

Score by ModelTop 10 of 16
  1. 1GPT-6 Astra92.0%
  2. 2GLM-5.390.2%
  3. 3GPT-5.6 Sol89.8%
  4. 4Grok 4.687.4%
  5. 5GPT-5.6 Terra82.9%
  6. 6Kimi K379.9%
  7. 7GLM-5.275.4%
  8. 8Claude Opus 4.875.0%
  9. 9Gemini 3.6 Flash68.7%
  10. 10GPT-5.6 Luna63.8%

Track results

Detection Engineering: Gemini 3.6 Flash led at 98.2%. IR: Synthesis & Containment: GPT-5.6 Sol led at 97.6%. TH: Attack Investigation: GLM-5.3 led at 87.0%.

Open-weight results

GLM-5.3 (90.2%) had the highest score among the 6 open-weight models evaluated. GLM-5.3 cost $1.61/task.

Cost per task

DeepSeek V4 Flash had the lowest measured cost at $0.02/task, with 56.5% benchmark score. GPT-6 Astra scored 92.0% at $2.27/task. Costs reflect the evaluated task mix and the stated pricing assumptions.

Cost per TaskTop 10 of 16
  1. 1DeepSeek V4 Flash$0.02
  2. 2MiniMax M3$0.21
  3. 3GPT-5.6 Luna$0.24
  4. 4DeepSeek V4 Pro$0.27
  5. 5GPT-5.6 Terra$0.36
  6. 6GLM-5.2$0.56
  7. 7Claude Sonnet 5$0.73
  8. 8GPT-5.6 Sol$0.86
  9. 9Claude Opus 4.8$1.05
  10. 10Gemini 3.6 Flash$1.52

Time per task

GPT-5.6 Terra had the lowest reported average latency at 1m 31s/task, with 82.9% benchmark score. Latency includes the work performed by the agent during evaluation.

Task Duration (avg)Top 10 of 16
  1. 1GPT-5.6 Terra1m 31s
  2. 2DeepSeek V4 Flash1m 47s
  3. 3Gemini 3.6 Flash1m 49s
  4. 4GPT-5.6 Luna1m 51s
  5. 5GPT-5.6 Sol2m 42s
  6. 6Claude Sonnet 53m 2s
  7. 7Gemini 3.1 Pro3m 11s
  8. 8GPT-6 Astra3m 12s
  9. 9MiniMax M33m 41s
  10. 10Claude Opus 4.84m 14s

Reliability: score spread

GLM-5.3 had the smallest observed score spread at 13.8 percentage points; its overall score was 90.2%. Spread is the highest minus lowest score across recorded tasks and trials. It includes differences in task difficulty and is not a measure of repeatability on one fixed task.

Score spread (percentage points)Top 10 of 16
  1. 1GLM-5.313.8pp
  2. 2GPT-6 Astra15.8pp
  3. 3Kimi K319.4pp
  4. 4GPT-5.6 Sol37.5pp
  5. 5Claude Opus 541.3pp
  6. 6Gemini 3.1 Pro52.5pp
  7. 7DeepSeek V4 Flash55.7pp
  8. 8GPT-5.6 Terra57.1pp
  9. 9Gemini 3.6 Flash73.8pp
  10. 10Grok 4.6100.0pp

Run completion

16 of 16 models completed all recorded runs without an unrecoverable error. Completion does not imply a correct answer or rule out a refusal.

Task Completion RateTop 10 of 16
  1. 1GPT-6 Astra100.0%
  2. 2GLM-5.3100.0%
  3. 3GPT-5.6 Sol100.0%
  4. 4Grok 4.6100.0%
  5. 5GPT-5.6 Terra100.0%
  6. 6Kimi K3100.0%
  7. 7GLM-5.2100.0%
  8. 8Claude Opus 4.8100.0%
  9. 9Gemini 3.6 Flash100.0%
  10. 10GPT-5.6 Luna100.0%

Measured Trade-offs

  • GPT-6 Astra — Highest observed benchmark score: 92.0% at $2.27/task.
  • DeepSeek V4 Flash — Lowest measured cost: $0.02/task, with 56.5% benchmark score.
  • GLM-5.3 — Highest observed open-weight benchmark score: 90.2% at $1.61/task.

Methodology

Scoring

  • Score: Equal-weight mean of DE, IR, and TH. DE is a separate detection-question track; IR and TH grade different phases of the same open report
  • Reports: Fact and evidence coverage against hidden rubrics. IR scores synthesis and containment; TH scores the investigation of the attack chain
  • Cost / task: USD per logical evaluated task, including repeated attempts. A bundled question set counts its individual questions; a report counts once. Prompt caching is included; judge costs are excluded
  • Latency: Average elapsed time per logical evaluated task, using the export normalization for bundled question sets
  • Reliability / spread: On this page, highest minus lowest score across recorded tasks and trials, in percentage points (lower is tighter). The BlueBench reliability score is 100 minus the mean of these six spreads: inverted spread, not same-task repeatability

Setup

  • Logs from a real AWS intrusion executed in a controlled lab
  • 24K+ events across CloudTrail, S3 access logs, VPC Flow Logs, and GuardDuty, loaded into a sandboxed query environment
  • Agents given SQL query tools plus shell and file-writing tools; no web search or external threat intelligence
  • Eight distinct tasks: seven detection questions and one report. IR and TH are two scores of that report, not separate report runs. Recorded attempt counts differ by model

Report Prompt

System
You are a security analyst investigating a security incident in an AWS cloud environment. Use DuckDB to investigate the Parquet log data in /data/case/.
User
Write a comprehensive incident response report for this AWS security incident. It should be thorough and self-contained: a reader should come away understanding what happened, who was responsible, the impact, and your recommended response. Cover the incident end-to-end rather than focusing narrowly on a single aspect. Back every material claim and recommendation with specific evidence from the logs (source table, event name, identifier, timestamp, or source IP). Be precise about what the evidence does and does not establish — do not overstate impact or success.
[High-severity alert + AWS telemetry]

Caveats

  • Results describe the current export and its recorded task mix; counts of repeated attempts can differ across models.
  • Spread is the highest minus lowest score across recorded tasks and trials. Different task difficulty and question bundling affect it; it does not isolate same-task repeatability.
  • Costs are USD per evaluated task, including repeated attempts. Bundled question sets use logical question counts; each report attempt is counted once. Provider prompt caching is included and judge costs are excluded.
  • In a track-filtered view, the score changes but cost and latency remain the benchmark-wide averages; spread is shown only in the all-tracks view.
  • Report scores depend on LLM judging. Detection queries are executed against the evidence. Close scores do not establish a statistically significant ordering.

Other Benchmarks

Impact & Exfiltration
Impact & Exfiltration

BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth

Simulation
Ransomware
Backup Sabotage
Data Exfiltration
Windows
Linux
3 samples · 16 models · Sep 2026
Identity & Active Directory Attacks
Identity & Active Directory Attacks

BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth

Simulation
Active Directory
Password Spray
Domain Persistence
Credential Access
Hands-on-Keyboard
Windows
4 samples · 16 models · Sep 2026
Initial Access & Command-and-Control
Initial Access & Command-and-Control

BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth

Simulation
Initial Access
Command & Control
Mail Compromise
Webshell
DGA Botnet
Windows
Linux
5 samples · 16 models · Sep 2026
Windows Enterprise Intrusion
Windows Enterprise Intrusion

BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting

Detection Engineering
Malware Analysis
Incident Response
Threat Hunting
Lateral Movement
Credential Access
Memory Forensics
40 samples · 16 models · Sep 2026
macOS Threat Investigation
macOS Threat Investigation

BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering

Incident Response
Threat Hunting
Detection Engineering
macOS Forensics
Credential Access
Data Exfiltration
14 samples · 16 models · Sep 2026
NYU CTF Bench
NYU CTF Bench

Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving

Reverse Engineering
Forensics
Miscellaneous
81 samples · 11 models · Feb 2026
CyberMetric
CyberMetric

Multiple-choice cybersecurity knowledge evaluation across 10,000 questions

Standards & Certifications
Network Security
Cryptography
Risk Management
Access Control
Incident Response
Application Security
Cloud Security
10180 samples · 13 models · Feb 2026
Cybench (Defensive Subset)
Cybench (Defensive Subset)

Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills

Forensics
Reverse Engineering
Miscellaneous
Hardware
18 samples · 10 models · Jan 2026
Sigma Detection Classification
Sigma Detection Classification

Multi-label classification of MITRE ATT&CK techniques from Sigma rules

Detection Engineering
MITRE ATT&CK
SIEM
Windows
Linux
Cloud
Network
Application
2733 samples · 12 models · Jan 2026
BOTSv3 Blue Team CTF
BOTSv3 Blue Team CTF

Blue team CTF scenarios testing incident response and threat hunting

Incident Response
Threat Hunting
Alert Triage
Log Analysis
Advanced Persistent Threat (APT)
Cloud Security (AWS/Azure)
51 samples · 15 models · Dec 2025

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo