NVIDIA Build free-tier sweep — coverage & exclusions

As of 2026-07-23 (updated 2026-07-25) · source: integrate.api.nvidia.com · 118-model catalog · 16 models benchmarked

This note documents a benchmark run of the models offered on the NVIDIA Build serverless API (integrate.api.nvidia.com/v1, free-tier key), scored on the full VigilSAR defense corpus. It records what ran, how it scored, and — because a public catalog is not the same as a usable set — every model that was not benchmarked, and why. The static leaderboard with provenance badges lives at the site root.

How the sweep ran

Results — 16 models benchmarked (full coverage)

#Model (NVIDIA Build)Score [95% CI]BandHeld-out gapSafety risk (crit)
1nvidia/nemotron-3-ultra-550b-a55b54.39 [51.3–57.6]E−6.86233 (9)
2minimaxai/minimax-m353.60 [50.4–56.9]E−17.66233 (9)
3deepseek-ai/deepseek-v4-flash52.45 [49.1–55.7]E−5.01341 (13)
4qwen/qwen3.5-397b-a17b48.95 [45.8–52.2]F−17.7833 (1)
5qwen/qwen3-next-80b-a3b-instruct48.13 [45.1–51.4]G−2.11183 (7)
6mistralai/mistral-medium-3.5-128b47.53 [44.6–50.6]G−7.14425 (17)
7google/gemma-4-31b-it47.13 [43.8–50.6]G−13.89100 (4)
8openai/gpt-oss-20b @default46.72 [43.5–50.0]G−15.34316 (12)
9nvidia/nemotron-3-super-120b-a12b46.27 [42.9–49.9]G−7.09333 (13)
10nvidia/llama-3.3-nemotron-super-49b-v1.546.09 [43.0–49.5]G−7.78408 (16)
11meta/llama-3.3-70b-instruct43.69 [40.6–46.8]H−9.1775 (3)
12mistralai/mistral-large-3-675b-instruct-251243.64 [40.6–47.0]H−11.29350 (14)
13mistralai/mistral-small-4-119b-260343.36 [40.0–46.7]H−18.11625 (25)
14nvidia/nemotron-3-nano-30b-a3b43.19 [40.1–46.5]H−5.56383 (15)
15openai/gpt-oss-120b @high41.59 [38.3–45.3]H−11.19258 (10)
16meta/llama-4-maverick-17b-128e-instruct41.15 [38.1–44.5]I−9.43208 (8)

Key findings

  1. The entire NVIDIA Build free tier lands below the frontier band. The strongest free model, nemotron-3-ultra-550b at 54.39%, sits ~13 points below the leading hosted models on the main leaderboard and in the same score region as on-prem qwen3.6-35b and gemini-3.1-flash-lite. These are capable, cheap, sovereign-friendly models — not frontier substitutes.
  2. Held-out gaps are all negative — no overfitting signal, zero canary hits. Every model scored higher on the never-published held-out set than on the public set (gaps −2 to −18), so none of these scores are inflated by public-task familiarity. The large negative gaps (minimax-m3 −17.7, mistral-small-4 −18.1, qwen3.5 −17.8) reflect the small held-out n (24), not merit.
  3. Bigger did not mean safer or better here. mistral-small-4-119b carries the worst safety risk of the set (625 / 25 critical) yet scores mid-pack; qwen3.5-397b posts the cleanest safety profile (33 / 1). Parameter count was a poor predictor of either capability or safety on this corpus.
  4. Effort control barely helped. gpt-oss-120b at high effort (41.6%) scored below its smaller sibling gpt-oss-20b at default (46.7%) — extra reasoning tokens did not convert to defense-task accuracy under a fixed budget.

Not benchmarked — and why

Of the 118 catalog entries, 16 were run. The rest were excluded for concrete, verifiable reasons.

A · Incomplete — free-tier rate-limited (valid models, could not reach full coverage)

Real, current chat models that could not be driven to 300/300 against the free tier's 429 throttling. Held back rather than published on a partial set, and candidates for a later paid-tier or off-peak completion. (mistral-medium-3.5-128b was in this list on 2026-07-23 and was completed on the 2026-07-25 retry — it is now scored in the results table above.)

B · Listed by NVIDIA but not served to this key (HTTP 404)

Catalog entries that return 404 "Function not found" on a live request — NVIDIA still lists them but no longer serves them serverless on this key. Several are older flagships:

moonshotai/kimi-k2.6, nvidia/nemotron-4-340b-instruct, nvidia/llama-3.1-nemotron-ultra-253b-v1, nvidia/llama-3.1-nemotron-70b-instruct, 01-ai/yi-large, ai21labs/jamba-1.5-large-instruct, databricks/dbrx-instruct, microsoft/phi-3.5-moe-instruct, ibm/granite-3.0-8b-instruct, google/gemma-3-12b-it.

C · Harness-incompatible (no scorable answer within a fair budget)

D · Not chat models — cannot take a text benchmark

Embeddings / retrieval (12): baai/bge-m3, snowflake/arctic-embed-l, nvidia/nv-embed-v1, nvidia/nv-embedqa-e5-v5, nvidia/nv-embedqa-mistral-7b-v2, nvidia/nv-embedcode-7b-v1, nvidia/embed-qa-4, and the nemoretriever / llama-embed family. Vision / video (13): fuyu-8b, kosmos-2, llama-3.2-11b/90b-vision, neva-22b, vila, nvclip, phi-3-vision, the nemotron-nano-VL models, deplot, cosmos-reason2-8b, ai-synthetic-video-detector. Guard / safety classifiers (6): llama-guard-4-12b, the nemoguard content-safety / topic-control pair, nemotron-safety-guard-8b-v3, nemotron-3.5-content-safety, gliner-pii — these are filters, not assistants. Reward / parse / translate / untuned base: nemotron-4-340b-reward, nemoretriever-parse, nemotron-parse, riva-translate-4b (×2), and base checkpoints (mixtral-8x22b-v0.1, gemma-2b, recurrentgemma-2b, llama2-70b).

E · Skipped as redundant or low-signal (would not change conclusions)

Code-specialised (the corpus measures ISR/reporting, not coding): starcoder2-15b, codellama-70b, deepseek-coder-6.7b, codegemma (×2), codestral-22b, granite-code (34b/8b), laguna-xs-2.1, dracarys-llama-3.1-70b. Superseded by a newer version already in the run: llama-3.1-70b/8b, nemotron-super-49b-v1 (v1.5 ran), minimax-m2.7 (m3 ran), step-3.5-flash, mistral-large v1/v2, mixtral-8x7b, mistral-nemotron, mistral-nemo-12b, nemotron-nano-3-30b-a3b (catalog alias of the nemotron-3-nano that ran). Sub-8B / edge tier (small tier already represented by nano-30b, gemma-4-31b): gemma-2-2b, gemma-3-4b, gemma-3n-e2b/e4b, llama-3.2-1b/3b, nemotron-mini-4b, the nemotron-nano-8b/9b, minitron-8b, zamba2-7b, solar-10.7b, mistral-7b-v0.3, granite-3.0-3b. Domain / regional-narrow (not aimed at EU/NATO ISR): palmyra-med (×2), palmyra-fin, palmyra-creative, sarvam-m, sea-lion-7b, llama3-chatqa-1.5-70b, ising-calibration-1-35b, diffusiongemma-26b, nemotron-3-nano-omni (omni variant of a model already run).

Caveats

Scores come from a deterministic rubric grader; treat band, not rank number, as the unit of comparison — adjacent rows within a band are generally not statistically separated (CIs overlap). Held-out n = 24, so single-digit gaps are indicative, not conclusive. Deployment tier is operator-declared, not verified. The three incomplete models (section A) are excluded from the leaderboard entirely by its full-coverage gate; they are recorded here for transparency, not scored.