Identity & Active Directory Attacks
Aug 2026BlueBench-Simulation-002: Four simulated Windows identity-tier attacks (password spray, domain persistence, credential-attack detection, and a hands-on-keyboard operator) scored against complete hidden ground truth
About BlueBench-Simulation
The series splits twelve authored investigations along the attack lifecycle across three benchmarks: 001 Initial Access & Command-and-Control, 002 Identity & Active Directory Attacks, and 003 Impact & Exfiltration.
About This Benchmark
Each case starts from a single weak alert, with several thousand raw events drawn from endpoint, network, identity, and directory telemetry. The hunt and the incident response report are graded by the LLM judge; the two detection queries are scored deterministically. We ran 14 models three times per task: 168 runs in total.
Sample Questions
Q: A behavioral deviation alert fires on the identity tier without naming a host or account. Hunt across all available telemetry, determine what is actually happening, and write a comprehensive threat-hunting report without assuming compromise.
A: [Free-form report scored against hidden fact and citation rubrics]
Q: A directory-lifecycle correlation alert fires on a Windows domain. Investigate end to end and write a self-contained incident report: scope, affected hosts and accounts, attacker tradecraft, impact, and the response actions the situation warrants.
A: [Free-form report plus response-action choices scored against hidden rubrics]
Q: Using the domain controller telemetry, identify the activity that harvested credential material for the domain’s service identities. Return the records constituting the theft and a detection query that excludes the domain’s routine service-authentication activity.
A: [Read-only detection query re-executed against the case evidence and scored by exact-record F1]
Key Findings
Accuracy
Grok 4.6 led at 77.6%, six points clear of Opus 4.8 at 71.5%. GPT-5.6 Sol (66.0%), GPT-5.6 Terra (62.5%), and Gemini 3.6 Flash (61.2%) followed, five models above 60% in all. Identity attacks leave crisp artifacts in directory telemetry, and the strongest models exploited that; the weakest did not, with MiniMax M3 falling to last at 31.2% after finishing fifth on Simulation-001.
Accuracy by Model
The Password Spray Hunt Was the Wall
The open-ended hunt for the password spray averaged 33.5% across the field, the hardest task in this benchmark and, by a hair, the hardest of the twelve simulation tasks. The alert points at the identity edge without naming a host or account, and benign lookalikes (routine lockouts, password rotations, authorized bulk operations) punish guessing. GPT-5.6 Terra led it at 60.6%, its only win on a report task anywhere in the series, with Grok 4.6 at 51.9%; most of the field never separated the spray from the noise. The other three tasks averaged between 48.7% and 63.4%.
Domain Persistence Was Nearly Solved
The domain-persistence incident response case was the most tractable task at a 63.4% field average. GPT-5.6 Sol averaged 95.4% and Grok 4.6 92.9%, each with at least one perfect trial, and Opus 4.8, Terra, and Kimi K3 all cleared 80%. Given a directory-lifecycle alert, the strongest models reliably reconstructed the persistence chain and chose the right response actions. The failure mode was over-reach: across the field, reports labeled an approved host change as persistence 18 times and an approved identity change as malicious 14 times, and each of those cost the report a multiplicative penalty. Gemini 3.1 Pro scored zero on all three trials after its output degenerated into a single repeated word instead of a report.
Detections Against a Live Domain
Both detection tasks re-execute the submitted query against the case evidence. Gemini 3.6 Flash was the best detection engineer here, averaging 86.7% across the pair with two perfect trials on the service-credential theft. Grok 4.6 (82.8%) and Opus 4.8 (81.0%) were the only other models above 80% on the pair; isolating the hands-on-keyboard operator's commands from authorized administration means discovering the discriminator in the evidence rather than guessing it from the alert. GPT-5.6 Sol split the pair, 84.8% on the operator and 37.3% on the credential theft, and MiniMax M3 scored zero on the operator task in every trial because it never returned a parseable structured answer.
Winning Was Expensive
Grok 4.6's 77.6% cost ~$8.47/run at ~31 minutes per run, the slowest win in the series. Opus 4.8 delivered 71.5% at ~$3.45/run and ~14 minutes, 40% of the price for six points less, on the smallest token budget in the field (19M). GPT-5.6 Terra was the value pick at 62.5% for ~$1.99/run and ~6 minutes per run, while Gemini 3.6 Flash was the most expensive model in the field at ~$14.19/run for a similar 61.2%.
Cost per Task
Refusals Concentrated Here
All 14 models completed all 12 runs, but four of the six refusals in the simulation series landed on these identity cases: Opus 5 refused a password-spray hunt and a domain-persistence investigation, and Sonnet 5 refused the password-spray hunt twice. Refused runs score whatever they earn, which is part of why Opus 5 (47.7%) and Sonnet 5 (41.9%) sit 24 to 30 points below Opus 4.8. Opus 5 also had the widest trial-to-trial spread in the field at 24 points.
Task Completion Rate
Model Recommendations
- Grok 4.6 — Best overall at 77.6% and first or second on every one of the four tasks. The trade-off is the slowest win in the series: ~$8.47/run at ~31 minutes per run.
- Claude Opus 4.8 — Six points back at 71.5% for 40% of the price: ~$3.45/run, ~14 minutes per run, and the smallest token budget in the field (19M).
- GPT-5.6 Terra — Value pick: 62.5% at ~$1.99/run, fourth overall, the fastest model in the field at ~6 minutes per run, and the only model to clear 60% on the password-spray hunt.
- Kimi K3 — The strongest open-weight result at 54.2%, sixth overall. GLM-5.2 followed at 45.8%; MiniMax M3 fell to last at 31.2% on these cases.
- Claude Opus 5 / Sonnet 5 — Both refused identity-attack investigations outright on some runs, and their averages reflect it. Opus 4.8 remains the dependable Anthropic pick for this workload.
Methodology
Scoring
- Score: Mean over three trials per task, scored against hidden rubrics authored with each environment. Detection engineering tasks are scored deterministically; free-form threat hunting and incident response reports are graded against the rubric by a single fixed LLM judge
- Facts: Each rubric fact is a specific finding (a host, account, address, filename, or hash) the report must commit to. For report tasks the judge checks that the finding is actually asserted, not merely mentioned, and that it is tied to the evidence behind it; partial credit is per-fact, not per-report
- Citations: Material claims must point at the evidence records they rest on; a correct finding without its supporting evidence earns reduced credit
- Red herrings: Asserting a known benign lookalike as malicious applies a multiplicative penalty to the report; a violation is counted only when repeated independent judge passes agree
- Detections: Submitted detection queries are re-executed against the case evidence and scored by exact-record F1 with no judge involved; rules that over-fire on the benign lookalikes lose precision
- Cost: USD per run based on token usage at list prices with prompt caching applied
- Latency: Wall-clock time per run
Environment Generation
- Four generated Windows enterprise environments centered on Active Directory and the identity tier
- Each environment is a small segmented enterprise with sustained baseline administration, deployment, backup, and monitoring activity
- Scripted intrusions: a password spray with account manipulation, a domain persistence operation, a service-credential theft, and a long hands-on-keyboard operation
- Matched benign red herrings in every case: authorized identity and administrative activity deliberately shaped to resemble the attack
- Several thousand raw events per case across endpoint, network, identity, and directory log sources, exposed to the agent as raw files plus a queryable index
Task Design
- Four tasks: one open-ended threat hunt, one incident response investigation, two detection engineering assignments
- Every task starts from a single alert framed as a lead, not a conclusion
- Threat hunting and incident response produce free-form reports, graded against the hidden rubric by a single fixed LLM judge; detection engineering returns structured answers plus a detection query, scored deterministically
- Agents are told to execute their detection query and inspect the results before submitting
Controls
- Same system prompt for all models, no per-model tuning
- "Thinking" mode enabled where available
- Three trials per task (k=3); reported scores are means over all 12 runs per model
- Refused or failed runs score whatever they earn; nothing is excluded from the averages
System Prompt
Paraphrased. Tool names, file paths, and the evidence schema given to the agent are omitted.
Caveats
- All telemetry is generated. Simulated environments give complete ground truth and rule out training-set contamination, but they are cleaner and more internally consistent than production data. Read these results alongside the BlueBench-Intrusion series, which uses real intrusion data.
- Threat hunting and incident response reports use LLM-judged scoring against the hidden rubric, which introduces some variability compared to the deterministic scoring used for detection engineering. The same judge model grades every report from every model.
- Four tasks at three trials each is 12 runs per model. Rankings are noisier than on larger benchmarks, and mid-table order especially should be read with error bars in mind.
- Scores are means over three trials per task. Run-to-run variance was significant for some models (Opus 5 moved 24 points between its best and worst trial), so single runs may land well above or below the reported average.
- Two zero scores are output failures rather than investigative ones: Gemini 3.1 Pro produced degenerate repeated-word output on every domain-persistence trial, and MiniMax M3 never returned a parseable structured answer on the hands-on-keyboard detection task. Both are counted as scored runs.
- Costs count agent-loop tokens at list prices with prompt caching applied, taken from the harness accounting for each run. Judge costs are not included.
Other Benchmarks
BlueBench-Simulation-003: Three simulated end-of-intrusion investigations (ransomware impact, backup sabotage, and data exfiltration) scored against complete hidden ground truth
BlueBench-Simulation-001: Five simulated investigations covering a mail-borne endpoint compromise, a DMZ webshell, and DGA botnet beaconing, scored against complete hidden ground truth
BlueBench-Intrusion-003: Real AWS intrusion through leaked CI credentials, scored on open-ended incident response reporting
BlueBench-Intrusion-002: Real multi-host Windows Active Directory intrusion spanning detection engineering, malware analysis, and open-ended incident reporting
BlueBench-Intrusion-001: Real macOS infostealer intrusion spanning incident response, threat hunting, and detection engineering
Real CTF challenges from CSAW competitions covering reverse engineering, forensics, and miscellaneous problem-solving
Defensive security CTF challenges testing forensics, reverse engineering, and miscellaneous security skills
Blue team CTF scenarios testing incident response and threat hunting
Multi-label classification of MITRE ATT&CK tactics and techniques from Sigma rules
Multiple-choice cybersecurity knowledge evaluation across 10,000 questions
AI for the blue team.
Run Cotool's harness in your environment to get real security work done