8 competitors — 27 case hits across 72 audited cases
| # | Competitor | Size | Detect | det+½ | Hits/Elig | Partial | Precision | FP/case | Other real | Cost/case | Latency | Tokens/case | Cases |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | deepseek-pro-shell ★ | large | 37% mean of 3 trials | 56% | 33%–44% | — | 80% | 0.67 | 11 | $1.18 | 819s | 2.6M | 9 |
| 2 | mimo-v2.5-pro | large | 30% mean of 3 trials | 44% | 22%–33% | — | 85% | 0.44 | 14 | $0.41 | 1419s | 2.0M | 9 |
| 3 | ornith-1.0-35b-shell ★ | small | 27% mean of 3 trials | 44% | 11%–38% | — | 71% | 0.56 | 5 | $0.00 | 1706s | 1.2M | 9 |
| 4 | mimo-shell ★ | large | 27% mean of 3 trials | 44% | 22%–33% | — | 94% | 0.11 | 10 | $0.35 | 979s | 1.7M | 9 |
| 5 | deepseek-v4-pro ★ | large | 22% mean of 3 trials | 39% | 11%–33% | 1 | 75% | 0.67 | 12 | $1.03 | 718s | 2.3M | 9 |
| 6 | nemotron-3-super-120b-shell | small | 19% mean of 3 trials | 44% | 11%–22% | — | 54% | 1.22 | 8 | $0.00 | 1168s | 1.8M | 9 |
| 7 | ornith-1.0-35b | small | 15% mean of 3 trials | 28% | 11%–22% | 1 | 69% | 0.56 | 7 | $0.00 | 1369s | 1.1M | 9 |
| 8 | nemotron-3-super-120b | small | 4% mean of 3 trials | 17% | 0%–11% | 1 | 37% | 1.33 | 6 | $0.00 | 1105s | 1.4M | 9 |
Detect = case hits / eligible (hits + partials + genuine misses); undetermined, refused, and auth/infra-excluded cases are not in the denominator. Partial = cases localized to the right spot but judged a different bug — right place, wrong bug. It is an eligible non-hit (it sits in the denominator where it would otherwise be a miss), so it never moves Detect or the ranking; det+½ (= (hits + 0.5·partials) / eligible) shows its half-credit value informationally. Precision = true findings / (true + false positives). Other real = confirmed real bugs the model found that are not the planted target CVE (extra capability, but not counted as detection). Cost/latency are the competitor's own spend per audited case. ★ = on a Pareto frontier below.
This run used repeated trials (--repeat). Detect is the mean detection rate across trials — each trial is scored on its own and the rates averaged, so it reflects a typical single run, not best-of-N. The ranking is by that mean. Hits/Elig shows the per-trial detection range (min to max); hover it for the best-of-N pooled count and the spread. In the per-case matrix, a HIT marked n/N was found in only n of N trials (flaky). See bench noise for the full per-trial breakdown.
Quality = detection rate x precision (precision treated as 1.0 when a competitor reported no scorable findings). Green points are non-dominated — no other competitor is at least as good on quality while also cheaper/faster. Size is shown in the table; it is categorical, so it is not used as a numeric Pareto axis.
Mean total tokens (prompt + completion, with the ReAct loop's resent context counted each turn) per audited case; the trailing number is mean latency/case. Bars are linear, so brute-force models dwarf frugal ones. Data-quality caveat: these are the tokens the provider's API reported — some OpenAI-compatible endpoints under-report usage (a near-zero bar with input ≈ output is the tell), so a suspiciously short bar may mean broken metering rather than a frugal model, and that competitor's cost/case is then an underestimate.
| Competitor | CVE-2026-5199 | CVE-2026-7474 | GHSA-9f49-8x56-jmjc | GHSA-cc7p-2j3x-x7xf | GHSA-f26g-jm89-4g65 | GHSA-j273-m5qq-6825 | GHSA-mpxh-8fq3-x8mh | GHSA-w52v-v783-gw97 | GHSA-x9h5-r9v2-vcww |
|---|---|---|---|---|---|---|---|---|---|
| deepseek-pro-shell | miss | HIT 1/3 | miss | miss | HIT | HIT | HIT 1/3 | HIT 2/3 | miss |
| mimo-v2.5-pro | miss | HIT | miss | miss | HIT 1/3 | HIT 1/3 | miss | HIT | miss |
| ornith-1.0-35b-shell | miss | HIT 2/3 | miss | miss | HIT 1/3 | HIT 1/3 | miss | HIT | miss |
| mimo-shell | miss | HIT 1/3 | miss | miss | HIT 2/3 | HIT 1/3 | miss | HIT | miss |
| deepseek-v4-pro | miss | HIT 2/3 | miss | miss | HIT 1/3 | part | miss | HIT | miss |
| nemotron-3-super-120b-shell | miss | HIT 1/3 | HIT 1/3 | miss | miss | HIT 1/3 | miss | HIT 2/3 | miss |
| ornith-1.0-35b | miss | HIT 1/3 | miss | miss | miss | part | miss | HIT | miss |
| nemotron-3-super-120b | miss | miss | miss | part | miss | HIT 1/3 | miss | miss | miss |
HIT = detected; part = right spot, wrong bug (half credit, eligible non-hit); miss = looked, found nothing; jerr = judge undetermined (out of denominator); refu = model refused the task (out of denominator, never a miss); excl = auth/infra failure (never a miss); · = not run. A HIT marked n/N was found in only n of N trials (flaky); a bare HIT was found in every trial.