Skip to main content

Identity & Active Directory Attacks

Aug 2026

BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth

Metric
Track
View
Provider
4 tasks · 14 models · Aug 2026
RankModels (14)AccCostLat
1
Grok 4.6
77.6%$8.4731m 13s
2
Claude Opus 4.8
71.5%$3.4514m 5s
3
GPT-5.6 Sol
66.0%$5.1310m 55s
4
GPT-5.6 Terra
62.5%$1.996m 18s
5
Gemini 3.6 Flash
61.2%$14.1914m 2s
6
Kimi K3
54.2%$4.6921m 5s
7
GPT-5.6 Luna
51.8%$1.298m 21s
8
Claude Opus 5Cyber refused
47.7%$6.1016m 52s
9
GLM-5.2
45.8%$2.4635m 52s
10
Claude Sonnet 5Cyber refused
41.9%$3.4912m 25s
Best AccuracyBest Open Weight
Dataset
BlueBench-Simulation-002
4 tasks
Task Types
Threat Hunting·1 task(25%)
Incident Response·1 task(25%)
Detection Engineering·2 tasks(50%)
Threat Hunting
Incident Response
Detection Engineering
Simulation
Active Directory
Password Spray
Domain Persistence
Credential Access
Hands-on-Keyboard
Windows

About BlueBench-Simulation

BlueBench-Simulation takes a different approach from our intrusion benchmarks. Instead of replaying one real compromise, every environment is fully generated: a small simulated enterprise with routine baseline activity, a scripted intrusion, and matched benign lookalikes designed to punish pattern-matching without verification. Because we author the environment, the ground truth is complete and none of the data can have leaked into training sets. Every case starts from a single weak alert; detection queries are scored deterministically and free-form reports are graded against the hidden rubric by a single fixed LLM judge.

The series splits twelve authored investigations along the attack lifecycle across three benchmarks: 001 Initial Access & Command-and-Control, 002 Identity & Active Directory Attacks, and 003 Impact & Exfiltration.

About This Benchmark

This benchmark covers the middle of the kill chain. The four tasks are about attacks on the identity tier, each on its own Windows Active Directory estate: an open-ended hunt for a password spray followed by account manipulation, an incident response investigation of domain persistence, a detection engineering assignment for service-credential theft on the domain controller, and a detection engineering assignment isolating a hands-on-keyboard operator's interactive commands inside the longest intrusion storyline in the series.

Each case starts from a single weak alert, with several thousand raw events drawn from endpoint, network, identity, and directory telemetry. The hunt and the incident response report are graded by the LLM judge; the two detection queries are scored deterministically. We ran 14 models three times per task: 168 runs in total.

Sample Questions

Threat Hunting

Q: A behavioral deviation alert fires on the identity tier without naming a host or account. Hunt across all available telemetry, determine what is actually happening, and write a comprehensive threat-hunting report without assuming compromise.

A: [Free-form report scored against hidden fact and citation rubrics]

Incident Response

Q: A directory-lifecycle correlation alert fires on a Windows domain. Investigate end to end and write a self-contained incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions the situation warrants.

A: [Free-form report plus response-action choices scored against hidden rubrics]

Detection Engineering

Q: Using the domain controller telemetry, identify the activity that harvested credential material for the domain’s service identities. Return the records constituting the theft and a detection query that excludes the domain’s routine service-authentication activity.

A: [Read-only detection query re-executed against the case evidence and scored by exact-record F1]

Key Findings

Accuracy

Grok 4.6 led at 77.6%, six points clear of Opus 4.8 at 71.5%. GPT-5.6 Sol (66.0%), GPT-5.6 Terra (62.5%), and Gemini 3.6 Flash (61.2%) followed, five models above 60% in all. Identity attacks leave crisp artifacts in directory telemetry, and the strongest models exploited that; the weakest did not, with MiniMax M3 falling to last at 31.2% after finishing fifth on Simulation-001.

Accuracy by Model

Xai
Anthropic
OpenAI
Google
Moonshot
Zhipu

The Password Spray Hunt Was the Wall

The open-ended hunt for the password spray averaged 33.5% across the field, the hardest task in this benchmark and, by a hair, the hardest of the twelve simulation tasks. The alert points at the identity edge without naming a host or account, and benign lookalikes (routine lockouts, password rotations, authorized bulk operations) punish guessing. GPT-5.6 Terra led it at 60.6%, its only win on a report task anywhere in the series, with Grok 4.6 at 51.9%; most of the field never separated the spray from the noise. The other three tasks averaged between 48.7% and 63.4%.

Domain Persistence Was Nearly Solved

The domain-persistence incident response case was the most tractable task at a 63.4% field average. GPT-5.6 Sol averaged 95.4% and Grok 4.6 92.9%, each with at least one perfect trial, and Opus 4.8, Terra, and Kimi K3 all cleared 80%. Given a directory-lifecycle alert, the strongest models reliably reconstructed the persistence chain and chose the right response actions. The failure mode was over-reach: across the field, reports labeled an approved host change as persistence 18 times and an approved identity change as malicious 14 times, and each of those cost the report a multiplicative penalty. Gemini 3.1 Pro scored zero on all three trials after its output degenerated into a single repeated word instead of a report.

Detections Against a Live Domain

Both detection tasks re-execute the submitted query against the case evidence. Gemini 3.6 Flash was the best detection engineer here, averaging 86.7% across the pair with two perfect trials on the service-credential theft. Grok 4.6 (82.8%) and Opus 4.8 (81.0%) were the only other models above 80% on the pair; isolating the hands-on-keyboard operator's commands from authorized administration means discovering the discriminator in the evidence rather than guessing it from the alert. GPT-5.6 Sol split the pair, 84.8% on the operator and 37.3% on the credential theft, and MiniMax M3 scored zero on the operator task in every trial because it never returned a parseable structured answer.

Winning Was Expensive

Grok 4.6's 77.6% cost ~$8.47/run at ~31 minutes per run, the slowest win in the series. Opus 4.8 delivered 71.5% at ~$3.45/run and ~14 minutes, 40% of the price for six points less, on the smallest token budget in the field (19M). GPT-5.6 Terra was the value pick at 62.5% for ~$1.99/run and ~6 minutes per run, while Gemini 3.6 Flash was the most expensive model in the field at ~$14.19/run for a similar 61.2%.

Cost per Task

OpenAI
Deepseek
Minimax
Zhipu
Anthropic
Moonshot

Refusals Concentrated Here

All 14 models completed all 12 runs, but four of the six refusals in the simulation series landed on these identity cases: Opus 5 refused a password-spray hunt and a domain-persistence investigation, and Sonnet 5 refused the password-spray hunt twice. Refused runs score whatever they earn, which is part of why Opus 5 (47.7%) and Sonnet 5 (41.9%) sit 24 to 30 points below Opus 4.8. Opus 5 also had the widest trial-to-trial spread in the field at 24 points.

Task Completion Rate

Xai
Anthropic
OpenAI
Google
Moonshot
Zhipu

Model Recommendations

  • Grok 4.6 Best overall at 77.6% and first or second on every one of the four tasks. The trade-off is the slowest win in the series: ~$8.47/run at ~31 minutes per run.
  • Claude Opus 4.8 Six points back at 71.5% for 40% of the price: ~$3.45/run, ~14 minutes per run, and the smallest token budget in the field (19M).
  • GPT-5.6 Terra Value pick: 62.5% at ~$1.99/run, fourth overall, the fastest model in the field at ~6 minutes per run, and the only model to clear 60% on the password-spray hunt.
  • Kimi K3 The strongest open-weight result at 54.2%, sixth overall. GLM-5.2 followed at 45.8%; MiniMax M3 fell to last at 31.2% on these cases.
  • Claude Opus 5 / Sonnet 5 Both refused identity-attack investigations outright on some runs, and their averages reflect it. Opus 4.8 remains the dependable Anthropic pick for this workload.

Methodology

Scoring

  • Score: Mean over three trials per task, scored against hidden rubrics authored with each environment. Detection engineering tasks are scored deterministically; free-form threat hunting and incident response reports are graded against the rubric by a single fixed LLM judge
  • Facts: Each rubric fact is a specific finding (a host, account, address, filename, or hash) the report must commit to. For report tasks the judge checks that the finding is actually asserted, not merely mentioned, and that it is tied to the evidence behind it; partial credit is per-fact, not per-report
  • Citations: Material claims must point at the evidence records they rest on; a correct finding without its supporting evidence earns reduced credit
  • Red herrings: Asserting a known benign lookalike as malicious applies a multiplicative penalty to the report; a violation is counted only when repeated independent judge passes agree
  • Detections: Submitted detection queries are re-executed against the case evidence and scored by exact-record F1 with no judge involved; rules that over-fire on the benign lookalikes lose precision
  • Cost: USD per run based on token usage at list prices with prompt caching applied
  • Latency: Wall-clock time per run

Environment Generation

  • Four generated Windows enterprise environments centered on Active Directory and the identity tier
  • Each environment is a small segmented enterprise with sustained baseline administration, deployment, backup, and monitoring activity
  • Scripted intrusions: a password spray with account manipulation, a domain persistence operation, a service-credential theft, and a long hands-on-keyboard operation
  • Matched benign red herrings in every case: authorized identity and administrative activity deliberately shaped to resemble the attack
  • Several thousand raw events per case across endpoint, network, identity, and directory log sources, exposed to the agent as raw files plus a queryable index

Task Design

  • Four tasks: one open-ended threat hunt, one incident response investigation, two detection engineering assignments
  • Every task starts from a single alert framed as a lead, not a conclusion
  • Threat hunting and incident response produce free-form reports, graded against the hidden rubric by a single fixed LLM judge; detection engineering returns structured answers plus a detection query, scored deterministically
  • Agents are told to execute their detection query and inspect the results before submitting

Controls

  • Same system prompt for all models, no per-model tuning
  • "Thinking" mode enabled where available
  • Three trials per task (k=3); reported scores are means over all 12 runs per model
  • Refused or failed runs score whatever they earn; nothing is excluded from the averages

System Prompt

System
You are a security analyst working an active case in Cotool. Available tools: - Execute commands in an isolated sandbox. - Write analysis scripts to a working directory. Case evidence: - Raw, heterogeneous telemetry for the case, plus a read-only queryable index over every record. - A manifest listing the source formats and record counts. Working requirements: - Treat the incoming alert as a lead, not a conclusion. - Work from raw evidence, correlate across source families, and preserve uncertainty. - Citations must reference the exact records returned by the evidence index. - Detection queries must be read-only and return the matching record identifiers. - Before submitting a detection query, execute it and inspect its result count and representative rows.
User
[Raw alert JSON] + Investigate this case end to end and write a single comprehensive report as your final response. The alert is a weak lead, not a conclusion. Proactively hunt across all of the available telemetry to determine what is actually happening. Correlate across source families and cite the specific evidence behind each material claim. Be rigorous about the limits of the evidence: clearly separate what the telemetry establishes from what it does not.
[Alert + task instructions]

Paraphrased. Tool names, file paths, and the evidence schema given to the agent are omitted.

Caveats

  • All telemetry is generated. Simulated environments give complete ground truth and rule out training-set contamination, but they are cleaner and more internally consistent than production data. Read these results alongside the BlueBench-Intrusion series, which uses real intrusion data.
  • Threat hunting and incident response reports use LLM-judged scoring against the hidden rubric, which introduces some variability compared to the deterministic scoring used for detection engineering. The same judge model grades every report from every model.
  • Four tasks at three trials each is 12 runs per model. Rankings are noisier than on larger benchmarks, and mid-table order especially should be read with error bars in mind.
  • Scores are means over three trials per task. Run-to-run variance was significant for some models (Opus 5 moved 24 points between its best and worst trial), so single runs may land well above or below the reported average.
  • Two zero scores are output failures rather than investigative ones: Gemini 3.1 Pro produced degenerate repeated-word output on every domain-persistence trial, and MiniMax M3 never returned a parseable structured answer on the hands-on-keyboard detection task. Both are counted as scored runs.
  • Costs count agent-loop tokens at list prices with prompt caching applied, taken from the harness accounting for each run. Judge costs are not included.

Other Benchmarks

Impact & Exfiltration
Impact & Exfiltration

BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth

Simulation
Ransomware
Backup Sabotage
Data Exfiltration
Windows
Linux
3 samples · 14 models · Aug 2026
Initial Access & Command-and-Control
Initial Access & Command-and-Control

BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth

Simulation
Initial Access
Command & Control
Mail Compromise
Webshell
DGA Botnet
Windows
Linux
5 samples · 14 models · Aug 2026
AWS Cloud Intrusion
AWS Cloud Intrusion

BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, scored on open-ended incident response reporting

AWS
Incident Response
CloudTrail
Identity
Persistence
Exfiltration
40 samples · 16 models · Jul 2026
Windows Enterprise Intrusion
Windows Enterprise Intrusion

BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting

Detection Engineering
Malware Analysis
Incident Response
Threat Hunting
Lateral Movement
Credential Access
Memory Forensics
40 samples · 22 models · Jul 2026
macOS Threat Investigation
macOS Threat Investigation

BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering

Incident Response
Threat Hunting
Detection Engineering
macOS Forensics
Credential Access
Data Exfiltration
36 samples · 9 models · Mar 2026
NYU CTF Bench
NYU CTF Bench

Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving

Reverse Engineering
Forensics
Miscellaneous
81 samples · 11 models · Feb 2026
Cybench (Defensive Subset)
Cybench (Defensive Subset)

Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills

Forensics
Reverse Engineering
Miscellaneous
Hardware
18 samples · 10 models · Jan 2026
BOTSv3 Blue Team CTF
BOTSv3 Blue Team CTF

Blue team CTF scenarios testing incident response and threat hunting

Incident Response
Threat Hunting
Alert Triage
Log Analysis
Advanced Peristent Threat (APT)
Cloud Security (AWS/Azure)
51 samples · 15 models · Dec 2025
Sigma Detection Classification
Sigma Detection Classification

Multi-label classification of MITRE ATT&CK tactics and techniques from Sigma rules

Detection Engineering
MITRE ATT&CK
SIEM
Windows
Linux
Cloud
Network
Application
2733 samples · 12 models · Jan 2026
CyberMetric
CyberMetric

Multiple-choice cybersecurity knowledge evaluation across 10,000 questions

Standards & Certifications
Network Security
Cryptography
Risk Management
Access Control
Incident Response
Application Security
Cloud Security
10180 samples · 13 models · Feb 2026

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo