Skip to main content

Reliability

Reliability = 100 − mean(max − min) across 6 benchmarks. Higher is tighter. Inverted spread, not same-task repeatability.

16 models · Updated Sep 2026
Methodology
View
Models
1GPT-6 Astra
64.1
2GPT-5.6 Sol
47.6
3Grok 4.6
47.4
4GPT-5.6 Terra
45.6
5GLM-5.3
37.1
6Kimi K3
36.3
7GPT-5.6 Luna
33.4
8GLM-5.2
31.7
9DeepSeek V4 Pro
31.1
10Gemini 3.6 Flash
29.8
11Claude Opus 5
25.1
12DeepSeek V4 Flash
23.7
13Gemini 3.1 Pro
19.2
14Claude Opus 4.8
17.3
15Claude Sonnet 5
8.1
16MiniMax M3
7.7
Per-benchmark reliability (100 − spread) — higher is tighter

Box = 100 − (max − min) across 6 benchmarks. Higher is tighter. Inverted spread, not same-task repeatability.

RankModels (16)
1
GPT-6 Astra
64.166.855.980.221.2
2
GPT-5.6 Sol
47.638.329.058.123.9
3
Grok 4.6
47.444.637.861.00.0
4
GPT-5.6 Terra
45.642.332.149.614.8
5
GLM-5.3
37.134.53.164.40.0
6
Kimi K3
36.332.615.555.00.0
7
GPT-5.6 Luna
33.439.631.242.20.0
8
GLM-5.2
31.736.030.542.60.0
9
DeepSeek V4 Pro
31.129.626.032.90.0
10
Gemini 3.6 Flash
29.823.119.644.111.1
11
Claude Opus 5
25.120.95.142.80.0
12
DeepSeek V4 Flash
23.721.617.334.50.0
13
Gemini 3.1 Pro
19.221.45.024.50.0
14
Claude Opus 4.8
17.316.82.528.80.0
15
Claude Sonnet 5
8.14.60.714.10.0
16
MiniMax M3
7.73.30.011.00.0

AI for the blue team.

Run Cotool's harness in your environment to get real security work done

Book a demo