Static public leaderboard

VigilSAR Defense LLM Benchmark

A static, procurement-facing snapshot built from existing benchmark artifacts only. No provider configuration, live API routes, model-run controls, or secrets are included.

Benchmark Facts

300tasks
12category roster
217pg_* procurement tasks
39public model rows
1pinned archived rows
-18.11VS-01 held-out gap (private set v2) - 34 scored; 0 red flag(s); threshold 5 points
scorer c8b266f2 score provenance generatedAt 2026-08-06T14:24:36.727Z site generatedAt 2026-08-06T14:30:36.699Z Fresh: taskSetHash 7b252ca2

Leaderboard

Data table fallback: score bars include aria labels and the table gives the exact values.
#ModelTier (declared)ScoreBandCapabilityRel+SafetyDeploy (decl.)ECESafety RiskCrit$/correctHeld-out gapCorpus
1 claude-code-cli:claude-fable-5 pinned archived hosted-api 67.77% A 67 66.95 53.64 0.241 66 2 0.126568 pending 7b252ca2
2 claude-code-cli:claude-fable-5-early-july hosted-api 65.37% B 66.3 60.72 54.04 0.235 283 11 0.27225 -5.81 7b252ca2
3 kimi-cli:kimi-k3 hosted-api 64.65% B 66.63 57.45 53.37 0.22 244 9 -10.97 7b252ca2
4 codex-cli:gpt-5.5-medium hosted-api 62.17% C 60.83 61.16 50.97 0.175 250 10 -9.12 7b252ca2
5 claude-code-cli:claude-opus-5-high hosted-api 61.7% C 62.53 55.84 53.37 0.231 160 5 0.181902 -13.02 7b252ca2
6 codex-cli:gpt-5.5 hosted-api 61.66% C 60.92 60.38 49.1 0.199 283 11 pending 7b252ca2
7 codex-cli:gpt-5.4-mini hosted-api 60.22% C 60.6 54.95 52.17 0.186 116 4 -5.61 7b252ca2
8 codex-cli:gpt-5.6-terra hosted-api 58.61% D 59.79 53.29 50.84 0.155 183 7 -6.56 7b252ca2
9 codex-cli:gpt-5.6-sol hosted-api 57.86% D 56.27 56.53 51.64 0.159 233 9 -4.86 7b252ca2
10 codex-cli:gpt-5.6-terra-xhigh hosted-api 57.69% D 58.44 51.89 50.84 0.162 158 6 -6.15 7b252ca2
11 codex-cli:gpt-5.3-codex-spark hosted-api 57.23% D 58.94 50.08 50.7 0.161 225 9 -3.50 7b252ca2
12 nvidia-build:thinkingmachines/inkling@default hosted-api 55.95% D 57.27 52.72 44.04 0.154 116 4 -4.02 7b252ca2
13 codex-cli:gpt-5.6-luna hosted-api 55.49% E 53.61 53.73 51.1 0.143 258 10 -7.54 7b252ca2
14 nvidia-build:z-ai/glm-5.2@default hosted-api 54.48% E 55.45 50.47 44.7 0.134 266 10 -16.04 7b252ca2
15 nvidia-build:nvidia/nemotron-3-ultra-550b-a55b@default hosted-api 54.39% E 58.03 43.51 48.7 0.134 233 9 -6.86 7b252ca2
16 nvidia-build:minimaxai/minimax-m3@default hosted-api 53.6% E 53.52 49.27 46.44 0.125 233 9 -17.66 7b252ca2
17 gemini-cli:gemini-3.1-pro-preview hosted-api 53.03% E 55.42 45.28 45.5 0.133 383 15 pending 7b252ca2
18 mlx-openai:qwen/qwen3.6-35b-a3b sovereign-deployable 52.78% E 55.06 47.18 68.64 0.124 100 4 pending 7b252ca2
19 nvidia-build:deepseek-ai/deepseek-v4-flash@default hosted-api 52.45% E 55.46 42.02 48.3 0.099 341 13 -5.01 7b252ca2
20 gemini-cli:gemini-3.1-flash-lite hosted-api 51.15% F 53.76 41.67 48.97 0.119 400 16 pending 7b252ca2
21 nvidia-build:qwen/qwen3.5-397b-a17b@default hosted-api 48.95% F 50.04 47.01 36.3 0.112 33 1 -17.78 7b252ca2
22 nvidia-build:qwen/qwen3-next-80b-a3b-instruct@default hosted-api 48.13% G 49.46 42.82 43.64 0.117 183 7 -2.11 7b252ca2
23 nvidia-build:mistralai/mistral-medium-3.5-128b@default hosted-api 47.53% G 49.39 38.94 45.64 0.116 425 17 -7.14 7b252ca2
24 nvidia-build:google/gemma-4-31b-it@default hosted-api 47.13% G 48.38 40.45 44.3 0.099 100 4 -13.89 7b252ca2
25 nvidia-build:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning@default hosted-api 46.88% G 49.99 36.49 45.1 0.127 500 20 -5.05 7b252ca2
26 nvidia-build:openai/gpt-oss-20b@default hosted-api 46.72% G 51.89 34.45 44.84 0.143 316 12 -15.34 7b252ca2
27 nvidia-build:mistralai/mistral-nemotron@default hosted-api 46.56% G 48.19 38.2 47.36 0.115 300 12 -8.83 7b252ca2
28 nvidia-build:nvidia/nemotron-3-super-120b-a12b@default hosted-api 46.27% G 49.57 34.83 46.84 0.14 333 13 -7.09 7b252ca2
29 nvidia-build:nvidia/llama-3.3-nemotron-super-49b-v1.5@default hosted-api 46.09% G 48.9 37.18 46.57 0.124 408 16 -7.78 7b252ca2
30 nvidia-build:meta/llama-3.3-70b-instruct@default hosted-api 43.69% H 47.95 39.02 30.3 0.159 75 3 -9.17 7b252ca2
31 nvidia-build:mistralai/mistral-large-3-675b-instruct-2512@default hosted-api 43.64% H 46.18 37.72 37.77 0.124 350 14 -11.29 7b252ca2
32 nvidia-build:mistralai/mistral-small-4-119b-2603@default hosted-api 43.36% H 46.94 34.83 39.1 0.16 625 25 -18.11 7b252ca2
33 nvidia-build:nvidia/nemotron-3-nano-30b-a3b@default hosted-api 43.19% H 45.1 35.71 43.77 0.177 383 15 -5.56 7b252ca2
34 nvidia-build:google/diffusiongemma-26b-a4b-it@default hosted-api 42.68% H 43.7 38.48 37.64 0.167 308 12 -9.95 7b252ca2
35 nvidia-build:nvidia/ising-calibration-1.5-31b@default hosted-api 41.96% H 46.65 34.39 35.24 0.167 583 23 -11.58 7b252ca2
36 nvidia-build:openai/gpt-oss-120b@high hosted-api 41.59% H 45.11 28.84 47.1 0.177 258 10 -11.19 7b252ca2
37 nvidia-build:nvidia/nvidia-nemotron-nano-9b-v2@default hosted-api 41.57% H 44.57 32.49 41.9 0.173 608 24 -6.88 7b252ca2
38 nvidia-build:meta/llama-4-maverick-17b-128e-instruct@default hosted-api 41.15% I 45.06 36.94 25.24 0.179 208 8 -9.43 7b252ca2
39 nvidia-build:poolside/laguna-xs-2.1@default hosted-api 40.92% I 44.06 33.85 35.5 0.176 500 20 -3.10 7b252ca2

Model coverage & exclusions

The leaderboard shows only models with full task coverage. The 2026-07-23 NVIDIA Build free-tier sweep (updated 2026-07-25) benchmarked 16 hosted models; the 2026-08-06 wave 3 added 8 more (24 NVIDIA Build rows total), led by thinkingmachines/inkling at 56.0% — the new free-tier best and the first free-tier row in the board's top half; the rest of the free tier stays below the frontier band. Both briefings name every model that was not benchmarked and why: free-tier unservable (deepseek-v4-pro, held at 212/300 across three sessions), catalog entries NVIDIA no longer serves (404 legacy flagships; 410 end-of-life: seed-oss-36b, ministral-14b), harness-incompatible reasoning models (step-3.7-flash), non-chat models (embeddings, vision, guard/reward classifiers), and redundant/superseded/edge/domain-narrow models.

Methodology notes — read before quoting scores. Grader: all scores come from a deterministic rubric grader (exact/regex/phrase gates); its agreement with an independent LLM judge is provisional (kappa 0.156 measured 2026-06-27, n=97, vs a lenient reference-free judge) and is not a validation. Ranks: rows are ordered by the Score column; adjacent rows within the same band are generally NOT statistically separated (the CI bars overlap) — treat band, not rank number, as the unit of comparison. Held-out gap: public score minus the score on heldout-v2, a genuinely independent private set (24 never-published tasks, matched category/points profile, per-task canaries; authored 2026-07-16). A positive gap beyond +5 points is red-flagged as an overfitting signal. n=24, so single-digit gaps are indicative, not conclusive. "pending" rows were not sweepable: pinned archived models no longer exist to query, and offline/unreachable providers are skipped rather than guessed. Deploy (declared): the deployment tier is operator-declared at run time, not verified; the composite blends that declared tier with the Sovereign-lane score, which grades written answers about deployment topics, not deployment capability. Effort: reasoning effort is recorded per response where available; rows may differ in effort tier. Pinned rows keep the score computed by the grader current at their run date and are not rescored when rubrics change. Moderation flags: hosted-API rows may include responses where the provider's server-side moderation blocked the prompt; those flags are scored as the model's response (typically a refusal), because the flag is the answer a hosted-API user receives (providerContentFlag in run records).

CSV schema contract: leaderboard.v1.json. Breaking column changes require a new schema version.