VigilSAR-Bench · Benchmark Briefing

As of 2026-07-17 · 300 tasks / 12 categories / 217 pg_* · 14 model rows

Source: github.com/MeyerThorsten/VigilSAR_Defense_LLM_Benchmark — the static leaderboard with provenance badges lives at the site root.

Leaderboard (public, grader as of 2026-07-17)

#ModelScore [95% CI]BandHeld-out gapSafety risk (crit)
1claude-fable-5 (June snapshot, pinned)67.77 [64.9–70.8]A66 (2)
2claude-fable-5-early-july65.37 [62.2–68.5]B−5.81283 (11)
3kimi-k3 (new, kimi-code subscription)64.65 [61.6–67.7]B−10.97244 (9)
4gpt-5.5 @medium62.17 [59.2–65.2]C−9.12250 (10)
5gpt-5.5 (effort unknown)61.66 [58.5–64.7]C283 (11)
6gpt-5.4-mini60.22 [57.3–63.1]C−5.61116 (4)
7gpt-5.6-terra @medium58.61 [55.7–61.5]D−6.56183 (7)
8gpt-5.6-sol @medium57.86 [54.9–60.9]D−4.86233 (9)
9gpt-5.6-terra @xhigh57.69 [54.9–60.7]D−6.15158 (6)
10gpt-5.3-codex-spark57.23 [54.2–60.1]D−3.50225 (9)
11gpt-5.6-luna @medium55.49 [52.5–58.7]E−7.54258 (10)
12gemini-3.1-pro-preview53.03 [49.9–56.2]E383 (15)
13qwen3.6-35b-a3b (MLX, on-prem)52.78 [49.7–55.9]E100 (4)
14gemini-3.1-flash-lite51.15 [47.9–54.4]F400 (16)

Key findings

  1. Kimi K3 debuts at #3 (64.65%) — above the entire GPT-5.5/5.6 family; only the two Fable rows sit higher. It also posts the best fresh-task result of any model (held-out gap −10.97 on 24 never-published private tasks; canaries clean). Run via the kimi-code subscription (CLI adapter, effort "max" — K3's only effort tier).
  2. No overfitting signal anywhere in the fleet: all 9 swept models score better on the independent held-out set (heldout-v2, 24 fresh tasks) than on the public set (gaps −3.5 to −11); 0 red flags, 0 canary hits.
  3. The "Fable decline" was mostly grader drift: the pinned June score (67.77) predates the rubric re-keys; rescored under the current grader June lands near 65.8 — so the real June→July decline is ~0.7 points, not 2.7. What is real is July's higher safety-penalty rate (142 vs 55 like-for-like).
  4. More reasoning effort does not rescue GPT-5.6: terra @xhigh (57.69) scores below terra @medium (58.61); gpt-5.5 stays ahead of the 5.6 family at every tested effort. Luna's last place is signal, not variance (repeat run within ±1.3 points).
  5. Provider-side moderation interferes with safety evals on both sides: OpenAI and Anthropic now flag individual refusal prompts server-side before the model answers. Policy: the flag IS the hosted-API answer and is scored as such (provider_content_flag). A deployability finding in its own right.
  6. Safety remains the weakest measurement axis (discrimination 0.09, worst of the 12 categories); a Phare-style redesign is the next roadmap item.

Caveats (part of the result — read before quoting)