VigilSAR-Bench · Benchmark Briefing
As of 2026-07-17 · 300 tasks / 12 categories / 217 pg_* · 14 model rows
Source: github.com/MeyerThorsten/VigilSAR_Defense_LLM_Benchmark — the static leaderboard with provenance badges lives at the site root.
Leaderboard (public, grader as of 2026-07-17)
| # | Model | Score [95% CI] | Band | Held-out gap | Safety risk (crit) |
|---|---|---|---|---|---|
| 1 | claude-fable-5 (June snapshot, pinned) | 67.77 [64.9–70.8] | A | – | 66 (2) |
| 2 | claude-fable-5-early-july | 65.37 [62.2–68.5] | B | −5.81 | 283 (11) |
| 3 | kimi-k3 (new, kimi-code subscription) | 64.65 [61.6–67.7] | B | −10.97 | 244 (9) |
| 4 | gpt-5.5 @medium | 62.17 [59.2–65.2] | C | −9.12 | 250 (10) |
| 5 | gpt-5.5 (effort unknown) | 61.66 [58.5–64.7] | C | – | 283 (11) |
| 6 | gpt-5.4-mini | 60.22 [57.3–63.1] | C | −5.61 | 116 (4) |
| 7 | gpt-5.6-terra @medium | 58.61 [55.7–61.5] | D | −6.56 | 183 (7) |
| 8 | gpt-5.6-sol @medium | 57.86 [54.9–60.9] | D | −4.86 | 233 (9) |
| 9 | gpt-5.6-terra @xhigh | 57.69 [54.9–60.7] | D | −6.15 | 158 (6) |
| 10 | gpt-5.3-codex-spark | 57.23 [54.2–60.1] | D | −3.50 | 225 (9) |
| 11 | gpt-5.6-luna @medium | 55.49 [52.5–58.7] | E | −7.54 | 258 (10) |
| 12 | gemini-3.1-pro-preview | 53.03 [49.9–56.2] | E | – | 383 (15) |
| 13 | qwen3.6-35b-a3b (MLX, on-prem) | 52.78 [49.7–55.9] | E | – | 100 (4) |
| 14 | gemini-3.1-flash-lite | 51.15 [47.9–54.4] | F | – | 400 (16) |
Key findings
- Kimi K3 debuts at #3 (64.65%) — above the entire GPT-5.5/5.6 family; only the two Fable rows sit higher. It also posts the best fresh-task result of any model (held-out gap −10.97 on 24 never-published private tasks; canaries clean). Run via the kimi-code subscription (CLI adapter, effort "max" — K3's only effort tier).
- No overfitting signal anywhere in the fleet: all 9 swept models score better on the independent held-out set (heldout-v2, 24 fresh tasks) than on the public set (gaps −3.5 to −11); 0 red flags, 0 canary hits.
- The "Fable decline" was mostly grader drift: the pinned June score (67.77) predates the rubric re-keys; rescored under the current grader June lands near 65.8 — so the real June→July decline is ~0.7 points, not 2.7. What is real is July's higher safety-penalty rate (142 vs 55 like-for-like).
- More reasoning effort does not rescue GPT-5.6: terra @xhigh (57.69) scores below terra @medium (58.61); gpt-5.5 stays ahead of the 5.6 family at every tested effort. Luna's last place is signal, not variance (repeat run within ±1.3 points).
- Provider-side moderation interferes with safety evals on both sides:
OpenAI and Anthropic now flag individual refusal prompts server-side before the model
answers. Policy: the flag IS the hosted-API answer and is scored as such
(
provider_content_flag). A deployability finding in its own right. - Safety remains the weakest measurement axis (discrimination 0.09, worst of the 12 categories); a Phare-style redesign is the next roadmap item.
Caveats (part of the result — read before quoting)
- Grader: deterministic rubric scoring; agreement with an independent LLM judge is provisional only (κ=0.156, n=97, measured 2026-06-27) — not a validation.
- Bands, not ranks: adjacent rows within a band are not statistically separated (CIs overlap).
- Effort comparability: K3 always runs @max; codex rows pinned @medium; the old gpt-5.5 row has unknown effort.
- Deployment tier is declared, not verified; the sovereign lane grades answers about air-gap topics, not deployability itself.
- Held-out: n=24 → single-digit gaps are indicative, not conclusive. "–" = not sweepable (pinned June build no longer exists; gemini needs re-auth; qwen offline).
- Synthetic & unclassified; measurement evidence, not a procurement recommendation.