Static public leaderboard
VigilSAR Defense LLM Benchmark
A static, procurement-facing snapshot built from existing benchmark artifacts only. No provider configuration, live API routes, model-run controls, or secrets are included.
Benchmark Facts
- Communication
- Compliance
- Domain comprehension
- Geospatial reasoning
- ISR reasoning
- Legal and ethical reasoning
- Planning support
- Robustness and calibration
- Safety
- Source evaluation
- Sovereign deployability
- Structured reporting
Leaderboard
| # | Model | Tier (declared) | Score | Band | Capability | Rel+Safety | Deploy (decl.) | ECE | Safety Risk | Crit | $/correct | Held-out gap | Corpus |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | claude-code-cli:claude-fable-5 pinned archived | hosted-api | 67.77% | A | 67 | 66.95 | 53.64 | 0.241 | 66 | 2 | 0.126568 | pending | 7b252ca2 |
| 2 | claude-code-cli:claude-fable-5-early-july | hosted-api | 65.37% | B | 66.3 | 60.72 | 54.04 | 0.235 | 283 | 11 | 0.27225 | -5.81 | 7b252ca2 |
| 3 | kimi-cli:kimi-k3 | hosted-api | 64.65% | B | 66.63 | 57.45 | 53.37 | 0.22 | 244 | 9 | -10.97 | 7b252ca2 | |
| 4 | codex-cli:gpt-5.5-medium | hosted-api | 62.17% | C | 60.83 | 61.16 | 50.97 | 0.175 | 250 | 10 | -9.12 | 7b252ca2 | |
| 5 | claude-code-cli:claude-opus-5-high | hosted-api | 61.7% | C | 62.53 | 55.84 | 53.37 | 0.231 | 160 | 5 | 0.181902 | -13.02 | 7b252ca2 |
| 6 | codex-cli:gpt-5.5 | hosted-api | 61.66% | C | 60.92 | 60.38 | 49.1 | 0.199 | 283 | 11 | pending | 7b252ca2 | |
| 7 | codex-cli:gpt-5.4-mini | hosted-api | 60.22% | C | 60.6 | 54.95 | 52.17 | 0.186 | 116 | 4 | -5.61 | 7b252ca2 | |
| 8 | codex-cli:gpt-5.6-terra | hosted-api | 58.61% | D | 59.79 | 53.29 | 50.84 | 0.155 | 183 | 7 | -6.56 | 7b252ca2 | |
| 9 | codex-cli:gpt-5.6-sol | hosted-api | 57.86% | D | 56.27 | 56.53 | 51.64 | 0.159 | 233 | 9 | -4.86 | 7b252ca2 | |
| 10 | codex-cli:gpt-5.6-terra-xhigh | hosted-api | 57.69% | D | 58.44 | 51.89 | 50.84 | 0.162 | 158 | 6 | -6.15 | 7b252ca2 | |
| 11 | codex-cli:gpt-5.3-codex-spark | hosted-api | 57.23% | D | 58.94 | 50.08 | 50.7 | 0.161 | 225 | 9 | -3.50 | 7b252ca2 | |
| 12 | nvidia-build:thinkingmachines/inkling@default | hosted-api | 55.95% | D | 57.27 | 52.72 | 44.04 | 0.154 | 116 | 4 | -4.02 | 7b252ca2 | |
| 13 | codex-cli:gpt-5.6-luna | hosted-api | 55.49% | E | 53.61 | 53.73 | 51.1 | 0.143 | 258 | 10 | -7.54 | 7b252ca2 | |
| 14 | nvidia-build:z-ai/glm-5.2@default | hosted-api | 54.48% | E | 55.45 | 50.47 | 44.7 | 0.134 | 266 | 10 | -16.04 | 7b252ca2 | |
| 15 | nvidia-build:nvidia/nemotron-3-ultra-550b-a55b@default | hosted-api | 54.39% | E | 58.03 | 43.51 | 48.7 | 0.134 | 233 | 9 | -6.86 | 7b252ca2 | |
| 16 | nvidia-build:minimaxai/minimax-m3@default | hosted-api | 53.6% | E | 53.52 | 49.27 | 46.44 | 0.125 | 233 | 9 | -17.66 | 7b252ca2 | |
| 17 | gemini-cli:gemini-3.1-pro-preview | hosted-api | 53.03% | E | 55.42 | 45.28 | 45.5 | 0.133 | 383 | 15 | pending | 7b252ca2 | |
| 18 | mlx-openai:qwen/qwen3.6-35b-a3b | sovereign-deployable | 52.78% | E | 55.06 | 47.18 | 68.64 | 0.124 | 100 | 4 | pending | 7b252ca2 | |
| 19 | nvidia-build:deepseek-ai/deepseek-v4-flash@default | hosted-api | 52.45% | E | 55.46 | 42.02 | 48.3 | 0.099 | 341 | 13 | -5.01 | 7b252ca2 | |
| 20 | gemini-cli:gemini-3.1-flash-lite | hosted-api | 51.15% | F | 53.76 | 41.67 | 48.97 | 0.119 | 400 | 16 | pending | 7b252ca2 | |
| 21 | nvidia-build:qwen/qwen3.5-397b-a17b@default | hosted-api | 48.95% | F | 50.04 | 47.01 | 36.3 | 0.112 | 33 | 1 | -17.78 | 7b252ca2 | |
| 22 | nvidia-build:qwen/qwen3-next-80b-a3b-instruct@default | hosted-api | 48.13% | G | 49.46 | 42.82 | 43.64 | 0.117 | 183 | 7 | -2.11 | 7b252ca2 | |
| 23 | nvidia-build:mistralai/mistral-medium-3.5-128b@default | hosted-api | 47.53% | G | 49.39 | 38.94 | 45.64 | 0.116 | 425 | 17 | -7.14 | 7b252ca2 | |
| 24 | nvidia-build:google/gemma-4-31b-it@default | hosted-api | 47.13% | G | 48.38 | 40.45 | 44.3 | 0.099 | 100 | 4 | -13.89 | 7b252ca2 | |
| 25 | nvidia-build:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning@default | hosted-api | 46.88% | G | 49.99 | 36.49 | 45.1 | 0.127 | 500 | 20 | -5.05 | 7b252ca2 | |
| 26 | nvidia-build:openai/gpt-oss-20b@default | hosted-api | 46.72% | G | 51.89 | 34.45 | 44.84 | 0.143 | 316 | 12 | -15.34 | 7b252ca2 | |
| 27 | nvidia-build:mistralai/mistral-nemotron@default | hosted-api | 46.56% | G | 48.19 | 38.2 | 47.36 | 0.115 | 300 | 12 | -8.83 | 7b252ca2 | |
| 28 | nvidia-build:nvidia/nemotron-3-super-120b-a12b@default | hosted-api | 46.27% | G | 49.57 | 34.83 | 46.84 | 0.14 | 333 | 13 | -7.09 | 7b252ca2 | |
| 29 | nvidia-build:nvidia/llama-3.3-nemotron-super-49b-v1.5@default | hosted-api | 46.09% | G | 48.9 | 37.18 | 46.57 | 0.124 | 408 | 16 | -7.78 | 7b252ca2 | |
| 30 | nvidia-build:meta/llama-3.3-70b-instruct@default | hosted-api | 43.69% | H | 47.95 | 39.02 | 30.3 | 0.159 | 75 | 3 | -9.17 | 7b252ca2 | |
| 31 | nvidia-build:mistralai/mistral-large-3-675b-instruct-2512@default | hosted-api | 43.64% | H | 46.18 | 37.72 | 37.77 | 0.124 | 350 | 14 | -11.29 | 7b252ca2 | |
| 32 | nvidia-build:mistralai/mistral-small-4-119b-2603@default | hosted-api | 43.36% | H | 46.94 | 34.83 | 39.1 | 0.16 | 625 | 25 | -18.11 | 7b252ca2 | |
| 33 | nvidia-build:nvidia/nemotron-3-nano-30b-a3b@default | hosted-api | 43.19% | H | 45.1 | 35.71 | 43.77 | 0.177 | 383 | 15 | -5.56 | 7b252ca2 | |
| 34 | nvidia-build:google/diffusiongemma-26b-a4b-it@default | hosted-api | 42.68% | H | 43.7 | 38.48 | 37.64 | 0.167 | 308 | 12 | -9.95 | 7b252ca2 | |
| 35 | nvidia-build:nvidia/ising-calibration-1.5-31b@default | hosted-api | 41.96% | H | 46.65 | 34.39 | 35.24 | 0.167 | 583 | 23 | -11.58 | 7b252ca2 | |
| 36 | nvidia-build:openai/gpt-oss-120b@high | hosted-api | 41.59% | H | 45.11 | 28.84 | 47.1 | 0.177 | 258 | 10 | -11.19 | 7b252ca2 | |
| 37 | nvidia-build:nvidia/nvidia-nemotron-nano-9b-v2@default | hosted-api | 41.57% | H | 44.57 | 32.49 | 41.9 | 0.173 | 608 | 24 | -6.88 | 7b252ca2 | |
| 38 | nvidia-build:meta/llama-4-maverick-17b-128e-instruct@default | hosted-api | 41.15% | I | 45.06 | 36.94 | 25.24 | 0.179 | 208 | 8 | -9.43 | 7b252ca2 | |
| 39 | nvidia-build:poolside/laguna-xs-2.1@default | hosted-api | 40.92% | I | 44.06 | 33.85 | 35.5 | 0.176 | 500 | 20 | -3.10 | 7b252ca2 |
Model coverage & exclusions
The leaderboard shows only models with full task coverage. The 2026-07-23 NVIDIA Build free-tier sweep (updated 2026-07-25) benchmarked 16 hosted models; the 2026-08-06 wave 3 added 8 more (24 NVIDIA Build rows total), led by thinkingmachines/inkling at 56.0% — the new free-tier best and the first free-tier row in the board's top half; the rest of the free tier stays below the frontier band. Both briefings name every model that was not benchmarked and why: free-tier unservable (deepseek-v4-pro, held at 212/300 across three sessions), catalog entries NVIDIA no longer serves (404 legacy flagships; 410 end-of-life: seed-oss-36b, ministral-14b), harness-incompatible reasoning models (step-3.7-flash), non-chat models (embeddings, vision, guard/reward classifiers), and redundant/superseded/edge/domain-narrow models.