NVIDIA Build free-tier sweep — coverage & exclusions
As of 2026-07-23 (updated 2026-07-25) · source: integrate.api.nvidia.com · 118-model catalog · 16 models benchmarked
This note documents a benchmark run of the models offered on the NVIDIA Build
serverless API (integrate.api.nvidia.com/v1, free-tier key), scored on the full
VigilSAR defense corpus. It records what ran, how it scored, and — because a public catalog is
not the same as a usable set — every model that was not benchmarked, and why.
The static leaderboard with provenance badges lives at the site root.
How the sweep ran
- Corpus: each eligible model answered all 300 public tasks plus the private held-out-v2 set (24 never-published tasks with per-task canaries), via the OpenAI-compatible chat-completions adapter, deterministic rubric grader.
- Effort: reasoning effort is recorded per response. It was only set explicitly where the
endpoint exposes a control (
gpt-oss-*:high); every other model ran at the provider default and is recorded as@default. - Catalog: enumerated 2026-07-21 (118 model ids) and re-checked 2026-07-22 (unchanged). Eligibility was verified with a live request per model, not by name.
- Rate limits: the free tier throttles hard (HTTP 429). On the first pass three
otherwise-valid models could not be driven to full coverage; a 2026-07-25 retry completed one of them
(
mistral-medium-3.5-128b, now in the table). Two remain incomplete and are listed as such rather than scored on a partial set.
Results — 16 models benchmarked (full coverage)
| # | Model (NVIDIA Build) | Score [95% CI] | Band | Held-out gap | Safety risk (crit) |
|---|---|---|---|---|---|
| 1 | nvidia/nemotron-3-ultra-550b-a55b | 54.39 [51.3–57.6] | E | −6.86 | 233 (9) |
| 2 | minimaxai/minimax-m3 | 53.60 [50.4–56.9] | E | −17.66 | 233 (9) |
| 3 | deepseek-ai/deepseek-v4-flash | 52.45 [49.1–55.7] | E | −5.01 | 341 (13) |
| 4 | qwen/qwen3.5-397b-a17b | 48.95 [45.8–52.2] | F | −17.78 | 33 (1) |
| 5 | qwen/qwen3-next-80b-a3b-instruct | 48.13 [45.1–51.4] | G | −2.11 | 183 (7) |
| 6 | mistralai/mistral-medium-3.5-128b | 47.53 [44.6–50.6] | G | −7.14 | 425 (17) |
| 7 | google/gemma-4-31b-it | 47.13 [43.8–50.6] | G | −13.89 | 100 (4) |
| 8 | openai/gpt-oss-20b @default | 46.72 [43.5–50.0] | G | −15.34 | 316 (12) |
| 9 | nvidia/nemotron-3-super-120b-a12b | 46.27 [42.9–49.9] | G | −7.09 | 333 (13) |
| 10 | nvidia/llama-3.3-nemotron-super-49b-v1.5 | 46.09 [43.0–49.5] | G | −7.78 | 408 (16) |
| 11 | meta/llama-3.3-70b-instruct | 43.69 [40.6–46.8] | H | −9.17 | 75 (3) |
| 12 | mistralai/mistral-large-3-675b-instruct-2512 | 43.64 [40.6–47.0] | H | −11.29 | 350 (14) |
| 13 | mistralai/mistral-small-4-119b-2603 | 43.36 [40.0–46.7] | H | −18.11 | 625 (25) |
| 14 | nvidia/nemotron-3-nano-30b-a3b | 43.19 [40.1–46.5] | H | −5.56 | 383 (15) |
| 15 | openai/gpt-oss-120b @high | 41.59 [38.3–45.3] | H | −11.19 | 258 (10) |
| 16 | meta/llama-4-maverick-17b-128e-instruct | 41.15 [38.1–44.5] | I | −9.43 | 208 (8) |
Key findings
- The entire NVIDIA Build free tier lands below the frontier band. The strongest free model, nemotron-3-ultra-550b at 54.39%, sits ~13 points below the leading hosted models on the main leaderboard and in the same score region as on-prem qwen3.6-35b and gemini-3.1-flash-lite. These are capable, cheap, sovereign-friendly models — not frontier substitutes.
- Held-out gaps are all negative — no overfitting signal, zero canary hits. Every model scored higher on the never-published held-out set than on the public set (gaps −2 to −18), so none of these scores are inflated by public-task familiarity. The large negative gaps (minimax-m3 −17.7, mistral-small-4 −18.1, qwen3.5 −17.8) reflect the small held-out n (24), not merit.
- Bigger did not mean safer or better here. mistral-small-4-119b carries the worst safety risk of the set (625 / 25 critical) yet scores mid-pack; qwen3.5-397b posts the cleanest safety profile (33 / 1). Parameter count was a poor predictor of either capability or safety on this corpus.
- Effort control barely helped. gpt-oss-120b at
higheffort (41.6%) scored below its smaller sibling gpt-oss-20b at default (46.7%) — extra reasoning tokens did not convert to defense-task accuracy under a fixed budget.
Not benchmarked — and why
Of the 118 catalog entries, 16 were run. The rest were excluded for concrete, verifiable reasons.
A · Incomplete — free-tier rate-limited (valid models, could not reach full coverage)
Real, current chat models that could not be driven to 300/300 against the free tier's 429 throttling.
Held back rather than published on a partial set, and candidates for a later paid-tier or off-peak
completion. (mistral-medium-3.5-128b was in this list on 2026-07-23 and was completed on the
2026-07-25 retry — it is now scored in the results table above.)
z-ai/glm-5.2— 222/300deepseek-ai/deepseek-v4-pro— 212/300
B · Listed by NVIDIA but not served to this key (HTTP 404)
Catalog entries that return 404 "Function not found" on a live request — NVIDIA still lists
them but no longer serves them serverless on this key. Several are older flagships:
moonshotai/kimi-k2.6, nvidia/nemotron-4-340b-instruct,
nvidia/llama-3.1-nemotron-ultra-253b-v1, nvidia/llama-3.1-nemotron-70b-instruct,
01-ai/yi-large, ai21labs/jamba-1.5-large-instruct,
databricks/dbrx-instruct, microsoft/phi-3.5-moe-instruct,
ibm/granite-3.0-8b-instruct, google/gemma-3-12b-it.
C · Harness-incompatible (no scorable answer within a fair budget)
stepfun-ai/step-3.7-flash— spends the entire output budget on hidden reasoning and returns zero final content even at 2000+ tokens (finish_reason=length). Cannot be scored fairly against a fixed budget.bytedance/seed-oss-36b-instruct(10/300),mistralai/ministral-14b-instruct-2512(0/300),thinkingmachines/inkling(22/300) — returned empty response bodies on most calls; not reliably served on the free tier during this window.
D · Not chat models — cannot take a text benchmark
Embeddings / retrieval (12): baai/bge-m3, snowflake/arctic-embed-l, nvidia/nv-embed-v1, nvidia/nv-embedqa-e5-v5, nvidia/nv-embedqa-mistral-7b-v2, nvidia/nv-embedcode-7b-v1, nvidia/embed-qa-4, and the nemoretriever / llama-embed family. Vision / video (13): fuyu-8b, kosmos-2, llama-3.2-11b/90b-vision, neva-22b, vila, nvclip, phi-3-vision, the nemotron-nano-VL models, deplot, cosmos-reason2-8b, ai-synthetic-video-detector. Guard / safety classifiers (6): llama-guard-4-12b, the nemoguard content-safety / topic-control pair, nemotron-safety-guard-8b-v3, nemotron-3.5-content-safety, gliner-pii — these are filters, not assistants. Reward / parse / translate / untuned base: nemotron-4-340b-reward, nemoretriever-parse, nemotron-parse, riva-translate-4b (×2), and base checkpoints (mixtral-8x22b-v0.1, gemma-2b, recurrentgemma-2b, llama2-70b).
E · Skipped as redundant or low-signal (would not change conclusions)
Code-specialised (the corpus measures ISR/reporting, not coding): starcoder2-15b, codellama-70b, deepseek-coder-6.7b, codegemma (×2), codestral-22b, granite-code (34b/8b), laguna-xs-2.1, dracarys-llama-3.1-70b. Superseded by a newer version already in the run: llama-3.1-70b/8b, nemotron-super-49b-v1 (v1.5 ran), minimax-m2.7 (m3 ran), step-3.5-flash, mistral-large v1/v2, mixtral-8x7b, mistral-nemotron, mistral-nemo-12b, nemotron-nano-3-30b-a3b (catalog alias of the nemotron-3-nano that ran). Sub-8B / edge tier (small tier already represented by nano-30b, gemma-4-31b): gemma-2-2b, gemma-3-4b, gemma-3n-e2b/e4b, llama-3.2-1b/3b, nemotron-mini-4b, the nemotron-nano-8b/9b, minitron-8b, zamba2-7b, solar-10.7b, mistral-7b-v0.3, granite-3.0-3b. Domain / regional-narrow (not aimed at EU/NATO ISR): palmyra-med (×2), palmyra-fin, palmyra-creative, sarvam-m, sea-lion-7b, llama3-chatqa-1.5-70b, ising-calibration-1-35b, diffusiongemma-26b, nemotron-3-nano-omni (omni variant of a model already run).
Caveats
Scores come from a deterministic rubric grader; treat band, not rank number, as the unit of comparison — adjacent rows within a band are generally not statistically separated (CIs overlap). Held-out n = 24, so single-digit gaps are indicative, not conclusive. Deployment tier is operator-declared, not verified. The three incomplete models (section A) are excluded from the leaderboard entirely by its full-coverage gate; they are recorded here for transparency, not scored.