VigilSAR Defense LLM Benchmark · 0.1.0 · 300 tasks · generated 2026-08-06T14:30:36.718Z · data as of 2026-08-06 · scorer c8b266f2 · corpus 7b252ca2
#1
claude-code-cli:claude-fable-5
67.8% Overall
67.0% Capability
67.0% Rel+Safety
53.64% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-10 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 55.2% | 25 | 0 |
| Compliance | 71.4% | 25 | 2 |
| Domain comprehension | 73.2% | 25 | 0 |
| Geospatial reasoning | 42.5% | 25 | 0 |
| ISR reasoning | 68.0% | 25 | 0 |
| Legal and ethical reasoning | 84.9% | 25 | 0 |
| Planning support | 85.8% | 25 | 0 |
| Robustness and calibration | 47.4% | 25 | 1 |
| Safety | 64.1% | 25 | 0 |
| Source evaluation | 72.3% | 25 | 1 |
| Sovereign deployability | 72.3% | 25 | 0 |
| Structured reporting | 72.0% | 25 | 0 |
#2
claude-code-cli:claude-fable-5-early-july
65.2% Overall
66.3% Capability
60.7% Rel+Safety
54.04% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-16 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 58.0% | 25 | 0 |
| Compliance | 68.5% | 25 | 2 |
| Domain comprehension | 63.3% | 25 | 3 |
| Geospatial reasoning | 47.7% | 25 | 0 |
| ISR reasoning | 64.9% | 25 | 0 |
| Legal and ethical reasoning | 84.6% | 25 | 0 |
| Planning support | 84.0% | 25 | 0 |
| Robustness and calibration | 45.3% | 25 | 3 |
| Safety | 44.6% | 25 | 2 |
| Source evaluation | 65.3% | 25 | 2 |
| Sovereign deployability | 73.1% | 25 | 0 |
| Structured reporting | 80.8% | 25 | 0 |
#3
kimi-cli:kimi-k3
64.1% Overall
66.6% Capability
57.5% Rel+Safety
53.37% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-17 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 59.8% | 25 | 0 |
| Compliance | 60.6% | 25 | 2 |
| Domain comprehension | 71.1% | 25 | 1 |
| Geospatial reasoning | 42.9% | 25 | 0 |
| ISR reasoning | 62.0% | 25 | 0 |
| Legal and ethical reasoning | 79.0% | 25 | 0 |
| Planning support | 79.6% | 25 | 0 |
| Robustness and calibration | 47.3% | 25 | 2 |
| Safety | 42.9% | 25 | 5 |
| Source evaluation | 68.5% | 25 | 2 |
| Sovereign deployability | 71.7% | 25 | 0 |
| Structured reporting | 82.5% | 25 | 0 |
#4
codex-cli:gpt-5.5-medium
61.4% Overall
60.8% Capability
61.2% Rel+Safety
50.97% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-16 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 52.8% | 25 | 0 |
| Compliance | 77.1% | 25 | 1 |
| Domain comprehension | 67.8% | 25 | 3 |
| Geospatial reasoning | 45.6% | 25 | 1 |
| ISR reasoning | 44.3% | 25 | 0 |
| Legal and ethical reasoning | 77.8% | 25 | 0 |
| Planning support | 73.3% | 25 | 0 |
| Robustness and calibration | 49.0% | 25 | 2 |
| Safety | 40.8% | 25 | 3 |
| Source evaluation | 65.8% | 25 | 0 |
| Sovereign deployability | 66.9% | 25 | 0 |
| Structured reporting | 76.3% | 25 | 0 |
#5
claude-code-cli:claude-opus-5-high
61.3% Overall
62.5% Capability
55.8% Rel+Safety
53.37% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-26 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 61.1% | 25 | 0 |
| Compliance | 52.4% | 25 | 2 |
| Domain comprehension | 78.1% | 25 | 0 |
| Geospatial reasoning | 47.3% | 25 | 1 |
| ISR reasoning | 65.1% | 25 | 0 |
| Legal and ethical reasoning | 80.4% | 25 | 0 |
| Planning support | 55.2% | 25 | 1 |
| Robustness and calibration | 42.4% | 25 | 2 |
| Safety | 48.2% | 25 | 0 |
| Source evaluation | 60.4% | 25 | 4 |
| Sovereign deployability | 71.7% | 25 | 0 |
| Structured reporting | 70.5% | 25 | 0 |
#6
codex-cli:gpt-5.5
60.9% Overall
60.9% Capability
60.4% Rel+Safety
49.10% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-10 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 51.7% | 25 | 0 |
| Compliance | 72.0% | 25 | 0 |
| Domain comprehension | 67.2% | 25 | 3 |
| Geospatial reasoning | 47.1% | 25 | 0 |
| ISR reasoning | 52.9% | 25 | 0 |
| Legal and ethical reasoning | 74.5% | 25 | 0 |
| Planning support | 72.9% | 25 | 0 |
| Robustness and calibration | 50.6% | 25 | 2 |
| Safety | 44.4% | 25 | 3 |
| Source evaluation | 56.5% | 25 | 3 |
| Sovereign deployability | 63.2% | 25 | 0 |
| Structured reporting | 78.1% | 25 | 1 |
#7
codex-cli:gpt-5.4-mini
59.5% Overall
60.6% Capability
55.0% Rel+Safety
52.17% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-10 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 53.1% | 25 | 0 |
| Compliance | 57.5% | 25 | 2 |
| Domain comprehension | 70.4% | 25 | 1 |
| Geospatial reasoning | 47.3% | 25 | 0 |
| ISR reasoning | 47.7% | 25 | 0 |
| Legal and ethical reasoning | 74.7% | 25 | 0 |
| Planning support | 69.0% | 25 | 0 |
| Robustness and calibration | 49.3% | 25 | 0 |
| Safety | 38.4% | 25 | 2 |
| Source evaluation | 57.9% | 25 | 1 |
| Sovereign deployability | 69.3% | 25 | 0 |
| Structured reporting | 78.8% | 25 | 0 |
#8
codex-cli:gpt-5.6-terra
58.3% Overall
59.8% Capability
53.3% Rel+Safety
50.84% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-12 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 50.9% | 25 | 0 |
| Compliance | 59.8% | 25 | 0 |
| Domain comprehension | 56.8% | 25 | 4 |
| Geospatial reasoning | 46.6% | 25 | 0 |
| ISR reasoning | 58.9% | 25 | 0 |
| Legal and ethical reasoning | 75.6% | 25 | 0 |
| Planning support | 72.9% | 25 | 0 |
| Robustness and calibration | 41.1% | 25 | 0 |
| Safety | 36.6% | 25 | 3 |
| Source evaluation | 57.3% | 25 | 1 |
| Sovereign deployability | 66.7% | 25 | 0 |
| Structured reporting | 75.3% | 25 | 0 |
#9
codex-cli:gpt-5.6-sol
57.4% Overall
56.3% Capability
56.5% Rel+Safety
51.64% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-12 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 54.5% | 25 | 0 |
| Compliance | 59.9% | 25 | 1 |
| Domain comprehension | 63.3% | 25 | 4 |
| Geospatial reasoning | 47.7% | 25 | 0 |
| ISR reasoning | 45.0% | 25 | 0 |
| Legal and ethical reasoning | 76.0% | 25 | 0 |
| Planning support | 57.0% | 25 | 0 |
| Robustness and calibration | 45.3% | 25 | 2 |
| Safety | 45.0% | 25 | 2 |
| Source evaluation | 54.6% | 25 | 1 |
| Sovereign deployability | 68.3% | 25 | 0 |
| Structured reporting | 71.9% | 25 | 0 |
#10
codex-cli:gpt-5.6-terra-xhigh
57.1% Overall
58.4% Capability
51.9% Rel+Safety
50.84% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-16 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 52.1% | 25 | 0 |
| Compliance | 52.8% | 25 | 1 |
| Domain comprehension | 68.1% | 25 | 2 |
| Geospatial reasoning | 45.1% | 25 | 0 |
| ISR reasoning | 43.6% | 25 | 0 |
| Legal and ethical reasoning | 76.9% | 25 | 0 |
| Planning support | 69.2% | 25 | 0 |
| Robustness and calibration | 43.0% | 25 | 1 |
| Safety | 35.0% | 25 | 2 |
| Source evaluation | 59.1% | 25 | 1 |
| Sovereign deployability | 66.7% | 25 | 0 |
| Structured reporting | 71.9% | 25 | 0 |
#11
codex-cli:gpt-5.3-codex-spark
56.8% Overall
58.9% Capability
50.1% Rel+Safety
50.70% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-11 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 54.9% | 25 | 0 |
| Compliance | 39.4% | 25 | 2 |
| Domain comprehension | 63.9% | 25 | 0 |
| Geospatial reasoning | 41.0% | 25 | 1 |
| ISR reasoning | 52.4% | 25 | 0 |
| Legal and ethical reasoning | 76.0% | 25 | 0 |
| Planning support | 68.3% | 25 | 0 |
| Robustness and calibration | 43.8% | 25 | 2 |
| Safety | 41.1% | 25 | 3 |
| Source evaluation | 59.5% | 25 | 0 |
| Sovereign deployability | 66.4% | 25 | 0 |
| Structured reporting | 72.5% | 25 | 1 |
#12
codex-cli:gpt-5.6-luna
56.4% Overall
56.8% Capability
52.9% Rel+Safety
50.84% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-16 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 47.4% | 25 | 1 |
| Compliance | 62.5% | 25 | 0 |
| Domain comprehension | 63.7% | 25 | 0 |
| Geospatial reasoning | 47.5% | 25 | 0 |
| ISR reasoning | 42.8% | 25 | 0 |
| Legal and ethical reasoning | 73.0% | 25 | 0 |
| Planning support | 65.5% | 25 | 0 |
| Robustness and calibration | 38.8% | 25 | 1 |
| Safety | 37.3% | 25 | 3 |
| Source evaluation | 60.8% | 25 | 1 |
| Sovereign deployability | 66.7% | 25 | 0 |
| Structured reporting | 69.9% | 25 | 1 |
#13
nvidia-build:thinkingmachines/inkling@default
55.5% Overall
57.3% Capability
52.7% Rel+Safety
44.04% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 45.1% | 25 | 0 |
| Compliance | 34.0% | 25 | 2 |
| Domain comprehension | 71.1% | 25 | 0 |
| Geospatial reasoning | 41.5% | 25 | 0 |
| ISR reasoning | 43.0% | 25 | 0 |
| Legal and ethical reasoning | 82.8% | 25 | 0 |
| Planning support | 62.1% | 25 | 0 |
| Robustness and calibration | 45.0% | 25 | 1 |
| Safety | 49.0% | 25 | 2 |
| Source evaluation | 60.0% | 25 | 1 |
| Sovereign deployability | 53.1% | 25 | 0 |
| Structured reporting | 78.1% | 25 | 0 |
#14
nvidia-build:z-ai/glm-5.2@default
53.8% Overall
55.5% Capability
50.5% Rel+Safety
44.70% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 52.8% | 25 | 0 |
| Compliance | 39.3% | 25 | 7 |
| Domain comprehension | 66.6% | 25 | 0 |
| Geospatial reasoning | 35.4% | 25 | 0 |
| ISR reasoning | 36.0% | 25 | 0 |
| Legal and ethical reasoning | 76.7% | 25 | 0 |
| Planning support | 65.7% | 25 | 0 |
| Robustness and calibration | 41.8% | 25 | 3 |
| Safety | 44.1% | 25 | 1 |
| Source evaluation | 59.8% | 25 | 1 |
| Sovereign deployability | 54.4% | 25 | 0 |
| Structured reporting | 71.9% | 25 | 0 |
#15
nvidia-build:nvidia/nemotron-3-ultra-550b-a55b@default
53.7% Overall
58.0% Capability
43.5% Rel+Safety
48.70% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-21 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 50.5% | 25 | 0 |
| Compliance | 21.1% | 25 | 7 |
| Domain comprehension | 67.7% | 25 | 0 |
| Geospatial reasoning | 35.1% | 25 | 0 |
| ISR reasoning | 36.2% | 25 | 0 |
| Legal and ethical reasoning | 75.4% | 25 | 0 |
| Planning support | 72.4% | 25 | 0 |
| Robustness and calibration | 39.9% | 25 | 2 |
| Safety | 37.6% | 25 | 1 |
| Source evaluation | 66.1% | 25 | 0 |
| Sovereign deployability | 62.4% | 25 | 0 |
| Structured reporting | 78.2% | 25 | 0 |
#16
nvidia-build:minimaxai/minimax-m3@default
52.6% Overall
53.5% Capability
49.3% Rel+Safety
46.44% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 48.3% | 25 | 0 |
| Compliance | 36.7% | 25 | 7 |
| Domain comprehension | 59.6% | 25 | 0 |
| Geospatial reasoning | 38.9% | 25 | 0 |
| ISR reasoning | 37.4% | 25 | 0 |
| Legal and ethical reasoning | 77.2% | 25 | 0 |
| Planning support | 63.8% | 25 | 0 |
| Robustness and calibration | 39.2% | 25 | 2 |
| Safety | 44.0% | 25 | 1 |
| Source evaluation | 55.4% | 25 | 0 |
| Sovereign deployability | 57.9% | 25 | 0 |
| Structured reporting | 71.3% | 25 | 0 |
#17
mlx-openai:qwen/qwen3.6-35b-a3b
52.4% Overall
55.1% Capability
47.2% Rel+Safety
68.64% Deploy
sovereign-deployable (declared) · sovereign-capable (declared) · edge-candidate (name-heuristic)
300/300 · scored 2026-06-10 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 42.9% | 25 | 0 |
| Compliance | 35.7% | 25 | 3 |
| Domain comprehension | 64.0% | 25 | 0 |
| Geospatial reasoning | 38.6% | 25 | 0 |
| ISR reasoning | 34.0% | 25 | 0 |
| Legal and ethical reasoning | 78.3% | 25 | 0 |
| Planning support | 74.2% | 25 | 0 |
| Robustness and calibration | 33.8% | 25 | 0 |
| Safety | 40.9% | 25 | 1 |
| Source evaluation | 57.0% | 25 | 0 |
| Sovereign deployability | 52.3% | 25 | 0 |
| Structured reporting | 74.8% | 25 | 0 |
#18
gemini-cli:gemini-3.1-pro-preview
52.3% Overall
55.4% Capability
45.3% Rel+Safety
45.50% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-11 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 48.6% | 25 | 0 |
| Compliance | 26.6% | 25 | 11 |
| Domain comprehension | 73.3% | 25 | 0 |
| Geospatial reasoning | 33.9% | 25 | 0 |
| ISR reasoning | 29.6% | 25 | 0 |
| Legal and ethical reasoning | 75.8% | 25 | 0 |
| Planning support | 71.3% | 25 | 0 |
| Robustness and calibration | 38.3% | 25 | 1 |
| Safety | 40.4% | 25 | 3 |
| Source evaluation | 58.1% | 25 | 1 |
| Sovereign deployability | 56.0% | 25 | 0 |
| Structured reporting | 73.2% | 25 | 0 |
#19
nvidia-build:deepseek-ai/deepseek-v4-flash@default
51.7% Overall
55.5% Capability
42.0% Rel+Safety
48.30% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-21 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 45.0% | 25 | 0 |
| Compliance | 30.6% | 25 | 8 |
| Domain comprehension | 67.9% | 25 | 0 |
| Geospatial reasoning | 42.0% | 25 | 0 |
| ISR reasoning | 33.4% | 25 | 0 |
| Legal and ethical reasoning | 71.5% | 25 | 0 |
| Planning support | 71.2% | 25 | 0 |
| Robustness and calibration | 34.8% | 25 | 2 |
| Safety | 31.2% | 25 | 4 |
| Source evaluation | 55.5% | 25 | 1 |
| Sovereign deployability | 61.6% | 25 | 0 |
| Structured reporting | 73.1% | 25 | 0 |
#20
gemini-cli:gemini-3.1-flash-lite
50.8% Overall
53.8% Capability
41.7% Rel+Safety
48.97% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-06-10 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 46.9% | 25 | 0 |
| Compliance | 19.9% | 25 | 12 |
| Domain comprehension | 70.5% | 25 | 0 |
| Geospatial reasoning | 30.3% | 25 | 1 |
| ISR reasoning | 30.1% | 25 | 0 |
| Legal and ethical reasoning | 71.1% | 25 | 0 |
| Planning support | 68.5% | 25 | 0 |
| Robustness and calibration | 36.8% | 25 | 2 |
| Safety | 38.8% | 25 | 1 |
| Source evaluation | 61.0% | 25 | 0 |
| Sovereign deployability | 62.9% | 25 | 0 |
| Structured reporting | 69.1% | 25 | 0 |
#21
nvidia-build:qwen/qwen3.5-397b-a17b@default
48.1% Overall
50.0% Capability
47.0% Rel+Safety
36.30% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 32.1% | 25 | 0 |
| Compliance | 34.5% | 25 | 1 |
| Domain comprehension | 61.9% | 25 | 0 |
| Geospatial reasoning | 28.8% | 25 | 0 |
| ISR reasoning | 29.6% | 25 | 0 |
| Legal and ethical reasoning | 68.7% | 25 | 0 |
| Planning support | 72.0% | 25 | 0 |
| Robustness and calibration | 41.7% | 25 | 1 |
| Safety | 43.2% | 25 | 0 |
| Source evaluation | 52.7% | 25 | 0 |
| Sovereign deployability | 37.6% | 25 | 0 |
| Structured reporting | 73.3% | 25 | 0 |
#22
nvidia-build:qwen/qwen3-next-80b-a3b-instruct@default
47.6% Overall
49.5% Capability
42.8% Rel+Safety
43.64% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-23 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 41.5% | 25 | 1 |
| Compliance | 39.3% | 25 | 3 |
| Domain comprehension | 55.5% | 25 | 0 |
| Geospatial reasoning | 26.1% | 25 | 0 |
| ISR reasoning | 30.5% | 25 | 0 |
| Legal and ethical reasoning | 67.3% | 25 | 0 |
| Planning support | 70.8% | 25 | 0 |
| Robustness and calibration | 30.0% | 25 | 2 |
| Safety | 34.7% | 25 | 1 |
| Source evaluation | 50.0% | 25 | 1 |
| Sovereign deployability | 52.3% | 25 | 0 |
| Structured reporting | 71.9% | 25 | 0 |
#23
nvidia-build:mistralai/mistral-medium-3.5-128b@default
46.6% Overall
49.4% Capability
38.9% Rel+Safety
45.64% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-24 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 38.0% | 25 | 0 |
| Compliance | 21.1% | 25 | 14 |
| Domain comprehension | 58.0% | 25 | 0 |
| Geospatial reasoning | 29.1% | 25 | 0 |
| ISR reasoning | 34.7% | 25 | 0 |
| Legal and ethical reasoning | 64.3% | 25 | 0 |
| Planning support | 55.4% | 25 | 0 |
| Robustness and calibration | 40.9% | 25 | 1 |
| Safety | 29.5% | 25 | 2 |
| Source evaluation | 61.6% | 25 | 0 |
| Sovereign deployability | 56.3% | 25 | 0 |
| Structured reporting | 69.0% | 25 | 0 |
#24
nvidia-build:openai/gpt-oss-20b@default
46.4% Overall
51.9% Capability
34.5% Rel+Safety
44.84% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-23 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 39.9% | 25 | 0 |
| Compliance | 28.8% | 25 | 4 |
| Domain comprehension | 63.1% | 25 | 0 |
| Geospatial reasoning | 37.2% | 25 | 0 |
| ISR reasoning | 33.1% | 25 | 0 |
| Legal and ethical reasoning | 58.1% | 25 | 0 |
| Planning support | 67.8% | 25 | 0 |
| Robustness and calibration | 28.2% | 25 | 7 |
| Safety | 22.7% | 25 | 2 |
| Source evaluation | 49.5% | 25 | 1 |
| Sovereign deployability | 54.7% | 25 | 0 |
| Structured reporting | 72.7% | 25 | 0 |
#25
nvidia-build:google/gemma-4-31b-it@default
46.4% Overall
48.4% Capability
40.5% Rel+Safety
44.30% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 40.6% | 25 | 0 |
| Compliance | 18.3% | 25 | 1 |
| Domain comprehension | 64.9% | 25 | 0 |
| Geospatial reasoning | 26.0% | 25 | 0 |
| ISR reasoning | 17.2% | 25 | 0 |
| Legal and ethical reasoning | 70.9% | 25 | 0 |
| Planning support | 62.8% | 25 | 0 |
| Robustness and calibration | 36.4% | 25 | 1 |
| Safety | 36.2% | 25 | 2 |
| Source evaluation | 57.2% | 25 | 0 |
| Sovereign deployability | 53.6% | 25 | 0 |
| Structured reporting | 70.1% | 25 | 0 |
#26
nvidia-build:nvidia/nemotron-3-nano-omni-30b-a3b-reasoning@default
46.3% Overall
50.0% Capability
36.5% Rel+Safety
45.10% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 27.0% | 25 | 1 |
| Compliance | 21.1% | 25 | 10 |
| Domain comprehension | 60.6% | 25 | 0 |
| Geospatial reasoning | 37.3% | 25 | 0 |
| ISR reasoning | 32.3% | 25 | 0 |
| Legal and ethical reasoning | 68.3% | 25 | 0 |
| Planning support | 65.6% | 25 | 0 |
| Robustness and calibration | 24.5% | 25 | 7 |
| Safety | 32.1% | 25 | 2 |
| Source evaluation | 56.7% | 25 | 0 |
| Sovereign deployability | 55.2% | 25 | 0 |
| Structured reporting | 70.5% | 25 | 0 |
#27
nvidia-build:mistralai/mistral-nemotron@default
46.0% Overall
48.2% Capability
38.2% Rel+Safety
47.36% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 48.1% | 25 | 0 |
| Compliance | 21.8% | 25 | 10 |
| Domain comprehension | 53.3% | 25 | 0 |
| Geospatial reasoning | 27.3% | 25 | 0 |
| ISR reasoning | 35.4% | 25 | 0 |
| Legal and ethical reasoning | 58.9% | 25 | 0 |
| Planning support | 55.3% | 25 | 0 |
| Robustness and calibration | 37.7% | 25 | 0 |
| Safety | 34.5% | 25 | 2 |
| Source evaluation | 54.5% | 25 | 0 |
| Sovereign deployability | 59.7% | 25 | 0 |
| Structured reporting | 63.4% | 25 | 0 |
#28
nvidia-build:nvidia/llama-3.3-nemotron-super-49b-v1.5@default
46.0% Overall
48.9% Capability
37.2% Rel+Safety
46.57% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-21 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 40.8% | 25 | 0 |
| Compliance | 17.2% | 25 | 12 |
| Domain comprehension | 56.0% | 25 | 0 |
| Geospatial reasoning | 31.7% | 25 | 0 |
| ISR reasoning | 30.7% | 25 | 0 |
| Legal and ethical reasoning | 61.2% | 25 | 0 |
| Planning support | 58.8% | 25 | 0 |
| Robustness and calibration | 36.9% | 25 | 2 |
| Safety | 33.5% | 25 | 2 |
| Source evaluation | 58.3% | 25 | 1 |
| Sovereign deployability | 58.1% | 25 | 0 |
| Structured reporting | 65.9% | 25 | 0 |
#29
nvidia-build:nvidia/nemotron-3-super-120b-a12b@default
45.6% Overall
49.6% Capability
34.8% Rel+Safety
46.84% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-21 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 42.2% | 25 | 0 |
| Compliance | 22.4% | 25 | 9 |
| Domain comprehension | 66.0% | 25 | 0 |
| Geospatial reasoning | 26.9% | 25 | 0 |
| ISR reasoning | 28.1% | 25 | 0 |
| Legal and ethical reasoning | 61.1% | 25 | 0 |
| Planning support | 64.1% | 25 | 0 |
| Robustness and calibration | 23.9% | 25 | 4 |
| Safety | 32.0% | 25 | 1 |
| Source evaluation | 45.4% | 25 | 0 |
| Sovereign deployability | 58.7% | 25 | 0 |
| Structured reporting | 74.3% | 25 | 0 |
#30
nvidia-build:mistralai/mistral-large-3-675b-instruct-2512@default
43.1% Overall
46.2% Capability
37.7% Rel+Safety
37.77% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 44.6% | 25 | 0 |
| Compliance | 22.7% | 25 | 10 |
| Domain comprehension | 70.8% | 25 | 0 |
| Geospatial reasoning | 26.2% | 25 | 0 |
| ISR reasoning | 31.0% | 25 | 1 |
| Legal and ethical reasoning | 63.9% | 25 | 0 |
| Planning support | 52.6% | 25 | 0 |
| Robustness and calibration | 31.7% | 25 | 1 |
| Safety | 32.6% | 25 | 2 |
| Source evaluation | 52.3% | 25 | 0 |
| Sovereign deployability | 40.5% | 25 | 0 |
| Structured reporting | 45.8% | 25 | 0 |
#31
nvidia-build:meta/llama-3.3-70b-instruct@default
42.9% Overall
48.0% Capability
39.0% Rel+Safety
30.30% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-23 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 43.9% | 25 | 0 |
| Compliance | 29.7% | 25 | 0 |
| Domain comprehension | 52.7% | 25 | 0 |
| Geospatial reasoning | 26.9% | 25 | 0 |
| ISR reasoning | 21.9% | 25 | 0 |
| Legal and ethical reasoning | 60.6% | 25 | 0 |
| Planning support | 74.0% | 25 | 0 |
| Robustness and calibration | 38.3% | 25 | 1 |
| Safety | 27.6% | 25 | 2 |
| Source evaluation | 48.9% | 25 | 0 |
| Sovereign deployability | 25.6% | 25 | 0 |
| Structured reporting | 67.3% | 25 | 0 |
#32
nvidia-build:nvidia/nemotron-3-nano-30b-a3b@default
42.9% Overall
45.1% Capability
35.7% Rel+Safety
43.77% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-21 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 27.0% | 25 | 0 |
| Compliance | 21.4% | 25 | 6 |
| Domain comprehension | 54.4% | 25 | 0 |
| Geospatial reasoning | 27.7% | 25 | 0 |
| ISR reasoning | 21.9% | 25 | 0 |
| Legal and ethical reasoning | 64.8% | 25 | 0 |
| Planning support | 60.6% | 25 | 0 |
| Robustness and calibration | 26.0% | 25 | 7 |
| Safety | 30.7% | 25 | 1 |
| Source evaluation | 54.8% | 25 | 2 |
| Sovereign deployability | 52.5% | 25 | 0 |
| Structured reporting | 69.3% | 25 | 0 |
#33
nvidia-build:mistralai/mistral-small-4-119b-2603@default
42.8% Overall
46.9% Capability
34.8% Rel+Safety
39.10% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 32.4% | 25 | 0 |
| Compliance | 16.7% | 25 | 16 |
| Domain comprehension | 51.1% | 25 | 0 |
| Geospatial reasoning | 26.4% | 25 | 0 |
| ISR reasoning | 33.4% | 25 | 0 |
| Legal and ethical reasoning | 66.0% | 25 | 0 |
| Planning support | 64.3% | 25 | 0 |
| Robustness and calibration | 26.9% | 25 | 6 |
| Safety | 29.8% | 25 | 3 |
| Source evaluation | 52.6% | 25 | 0 |
| Sovereign deployability | 43.2% | 25 | 0 |
| Structured reporting | 68.4% | 25 | 0 |
#34
nvidia-build:google/diffusiongemma-26b-a4b-it@default
41.7% Overall
43.7% Capability
38.5% Rel+Safety
37.64% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 26.7% | 25 | 1 |
| Compliance | 25.6% | 25 | 8 |
| Domain comprehension | 57.4% | 25 | 0 |
| Geospatial reasoning | 28.6% | 25 | 0 |
| ISR reasoning | 21.6% | 25 | 0 |
| Legal and ethical reasoning | 59.7% | 25 | 0 |
| Planning support | 45.9% | 25 | 0 |
| Robustness and calibration | 31.3% | 25 | 2 |
| Safety | 37.3% | 25 | 2 |
| Source evaluation | 56.5% | 25 | 0 |
| Sovereign deployability | 40.3% | 25 | 0 |
| Structured reporting | 69.3% | 25 | 0 |
#35
nvidia-build:nvidia/ising-calibration-1.5-31b@default
41.6% Overall
46.6% Capability
34.4% Rel+Safety
35.24% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 49.1% | 25 | 0 |
| Compliance | 15.9% | 25 | 14 |
| Domain comprehension | 46.7% | 25 | 0 |
| Geospatial reasoning | 27.4% | 25 | 1 |
| ISR reasoning | 25.7% | 25 | 0 |
| Legal and ethical reasoning | 62.6% | 25 | 2 |
| Planning support | 65.3% | 25 | 0 |
| Robustness and calibration | 21.0% | 25 | 3 |
| Safety | 38.1% | 25 | 3 |
| Source evaluation | 46.5% | 25 | 1 |
| Sovereign deployability | 35.5% | 25 | 0 |
| Structured reporting | 65.7% | 25 | 0 |
#36
nvidia-build:nvidia/nvidia-nemotron-nano-9b-v2@default
41.2% Overall
44.6% Capability
32.5% Rel+Safety
41.90% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 31.5% | 25 | 0 |
| Compliance | 15.3% | 25 | 13 |
| Domain comprehension | 44.4% | 25 | 1 |
| Geospatial reasoning | 30.0% | 25 | 0 |
| ISR reasoning | 30.6% | 25 | 0 |
| Legal and ethical reasoning | 63.4% | 25 | 0 |
| Planning support | 62.2% | 25 | 0 |
| Robustness and calibration | 18.6% | 25 | 10 |
| Safety | 32.7% | 25 | 0 |
| Source evaluation | 47.1% | 25 | 1 |
| Sovereign deployability | 48.8% | 25 | 0 |
| Structured reporting | 66.1% | 25 | 0 |
#37
nvidia-build:openai/gpt-oss-120b@high
40.9% Overall
45.1% Capability
28.8% Rel+Safety
47.10% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-23 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 39.7% | 25 | 0 |
| Compliance | 19.7% | 25 | 8 |
| Domain comprehension | 48.6% | 25 | 0 |
| Geospatial reasoning | 41.9% | 25 | 0 |
| ISR reasoning | 31.5% | 25 | 0 |
| Legal and ethical reasoning | 42.8% | 25 | 0 |
| Planning support | 42.3% | 25 | 0 |
| Robustness and calibration | 27.5% | 25 | 0 |
| Safety | 25.4% | 25 | 2 |
| Source evaluation | 49.7% | 25 | 1 |
| Sovereign deployability | 59.2% | 25 | 0 |
| Structured reporting | 62.0% | 25 | 0 |
#38
nvidia-build:poolside/laguna-xs-2.1@default
39.9% Overall
44.1% Capability
33.9% Rel+Safety
35.50% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-08-06 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 44.8% | 25 | 0 |
| Compliance | 16.6% | 25 | 13 |
| Domain comprehension | 35.4% | 25 | 0 |
| Geospatial reasoning | 26.2% | 25 | 0 |
| ISR reasoning | 25.4% | 25 | 0 |
| Legal and ethical reasoning | 61.8% | 25 | 0 |
| Planning support | 65.3% | 25 | 0 |
| Robustness and calibration | 23.4% | 25 | 7 |
| Safety | 33.6% | 25 | 0 |
| Source evaluation | 44.0% | 25 | 0 |
| Sovereign deployability | 36.0% | 25 | 0 |
| Structured reporting | 67.2% | 25 | 0 |
#39
nvidia-build:meta/llama-4-maverick-17b-128e-instruct@default
39.5% Overall
45.1% Capability
36.9% Rel+Safety
25.24% Deploy
hosted-api (declared) · hosted
300/300 · scored 2026-07-22 · scorer unknown
| Lane | Score | Tasks | Safety |
| Communication | 31.2% | 25 | 0 |
| Compliance | 29.5% | 25 | 4 |
| Domain comprehension | 58.5% | 25 | 0 |
| Geospatial reasoning | 26.2% | 25 | 0 |
| ISR reasoning | 25.7% | 25 | 0 |
| Legal and ethical reasoning | 61.7% | 25 | 0 |
| Planning support | 51.0% | 25 | 0 |
| Robustness and calibration | 28.5% | 25 | 3 |
| Safety | 28.0% | 25 | 1 |
| Source evaluation | 51.3% | 25 | 1 |
| Sovereign deployability | 15.5% | 25 | 0 |
| Structured reporting | 71.5% | 25 | 0 |
Caveats: scores come from a deterministic rubric grader (LLM-judge agreement provisional, kappa 0.156 measured 2026-06-27 — not a validation). Deployment profiles are operator-declared, not verified; edge-candidate is a model-name heuristic. Adjacent ranks are often not statistically separated — compare bands, not rank numbers.