Nelson Benchmark Leaderboard

8 competitors — 27 case hits across 72 audited cases

8
Competitors
72
Cases audited
27
Case hits
$26.66
Competitor spend
$55.52
Judge spend

Leaderboard

#CompetitorSizeDetectdet+½Hits/EligPartialPrecisionFP/caseOther realCost/caseLatencyTokens/caseCases
1deepseek-pro-shell large37%
mean of 3 trials
56%33%–44%80%0.6711$1.18819s2.6M9
2mimo-v2.5-prolarge30%
mean of 3 trials
44%22%–33%85%0.4414$0.411419s2.0M9
3ornith-1.0-35b-shell small27%
mean of 3 trials
44%11%–38%71%0.565$0.001706s1.2M9
4mimo-shell large27%
mean of 3 trials
44%22%–33%94%0.1110$0.35979s1.7M9
5deepseek-v4-pro large22%
mean of 3 trials
39%11%–33%175%0.6712$1.03718s2.3M9
6nemotron-3-super-120b-shellsmall19%
mean of 3 trials
44%11%–22%54%1.228$0.001168s1.8M9
7ornith-1.0-35bsmall15%
mean of 3 trials
28%11%–22%169%0.567$0.001369s1.1M9
8nemotron-3-super-120bsmall4%
mean of 3 trials
17%0%–11%137%1.336$0.001105s1.4M9

Detect = case hits / eligible (hits + partials + genuine misses); undetermined, refused, and auth/infra-excluded cases are not in the denominator. Partial = cases localized to the right spot but judged a different bug — right place, wrong bug. It is an eligible non-hit (it sits in the denominator where it would otherwise be a miss), so it never moves Detect or the ranking; det+½ (= (hits + 0.5·partials) / eligible) shows its half-credit value informationally. Precision = true findings / (true + false positives). Other real = confirmed real bugs the model found that are not the planted target CVE (extra capability, but not counted as detection). Cost/latency are the competitor's own spend per audited case. ★ = on a Pareto frontier below.

This run used repeated trials (--repeat). Detect is the mean detection rate across trials — each trial is scored on its own and the rates averaged, so it reflects a typical single run, not best-of-N. The ranking is by that mean. Hits/Elig shows the per-trial detection range (min to max); hover it for the best-of-N pooled count and the spread. In the per-case matrix, a HIT marked n/N was found in only n of N trials (flaky). See bench noise for the full per-trial breakdown.

Pareto frontier

Quality vs cost / case

0.00.51.0quality (det x prec)$0.00$1.18cost / case (lower is better →)deepseek-pro-shellmimo-v2.5-proornith-1.0-35b-shellmimo-shelldeepseek-v4-pronemotron-3-super-120b…ornith-1.0-35bnemotron-3-super-120b

Quality vs latency / case

0.00.51.0quality (det x prec)718s1706slatency / case (lower is better →)deepseek-pro-shellmimo-v2.5-proornith-1.0-35b-shellmimo-shelldeepseek-v4-pronemotron-3-super-120b…ornith-1.0-35bnemotron-3-super-120b

Quality = detection rate x precision (precision treated as 1.0 when a competitor reported no scorable findings). Green points are non-dominated — no other competitor is at least as good on quality while also cheaper/faster. Size is shown in the table; it is categorical, so it is not used as a numeric Pareto axis.

Tokens & time per case

deepseek-pro-shell2.6M · 819sdeepseek-v4-pro2.3M · 718smimo-v2.5-pro2.0M · 1419snemotron-3-super-120b-shell1.8M · 1168smimo-shell1.7M · 979snemotron-3-super-120b1.4M · 1105sornith-1.0-35b-shell1.2M · 1706sornith-1.0-35b1.1M · 1369s

Mean total tokens (prompt + completion, with the ReAct loop's resent context counted each turn) per audited case; the trailing number is mean latency/case. Bars are linear, so brute-force models dwarf frugal ones. Data-quality caveat: these are the tokens the provider's API reported — some OpenAI-compatible endpoints under-report usage (a near-zero bar with input ≈ output is the tell), so a suspiciously short bar may mean broken metering rather than a frugal model, and that competitor's cost/case is then an underestimate.

Per-case results

Competitor CVE-2026-5199 CVE-2026-7474 GHSA-9f49-8x56-jmjc GHSA-cc7p-2j3x-x7xf GHSA-f26g-jm89-4g65 GHSA-j273-m5qq-6825 GHSA-mpxh-8fq3-x8mh GHSA-w52v-v783-gw97 GHSA-x9h5-r9v2-vcww
deepseek-pro-shell miss HIT 1/3 miss miss HIT HIT HIT 1/3 HIT 2/3 miss
mimo-v2.5-pro miss HIT miss miss HIT 1/3 HIT 1/3 miss HIT miss
ornith-1.0-35b-shell miss HIT 2/3 miss miss HIT 1/3 HIT 1/3 miss HIT miss
mimo-shell miss HIT 1/3 miss miss HIT 2/3 HIT 1/3 miss HIT miss
deepseek-v4-pro miss HIT 2/3 miss miss HIT 1/3 part miss HIT miss
nemotron-3-super-120b-shell miss HIT 1/3 HIT 1/3 miss miss HIT 1/3 miss HIT 2/3 miss
ornith-1.0-35b miss HIT 1/3 miss miss miss part miss HIT miss
nemotron-3-super-120b miss miss miss part miss HIT 1/3 miss miss miss

HIT = detected; part = right spot, wrong bug (half credit, eligible non-hit); miss = looked, found nothing; jerr = judge undetermined (out of denominator); refu = model refused the task (out of denominator, never a miss); excl = auth/infra failure (never a miss); · = not run. A HIT marked n/N was found in only n of N trials (flaky); a bare HIT was found in every trial.