AI Security
Blog →
AI Security

LLMScanBench Q3 2026: Choosing the Right Approach for AI Vulnerability Discovery | Harness Blog

Compare LLM vulnerability scanners with AI SAST on precision, recall, speed, cost, and consistency to choose the right AI security approach.

TL;DR

  • Not surprisingly, LLM scanners have high precision on this benchmark.   
  • LLM scanners fall short on recall. 
  • GPT 5.5 performed better than Opus 4.8 in precision and recall in most cases. 
  • LLM scanners are inherently probabilistic. The same scan on the exact same unchanged code produced different findings across repeated runs, used different tokens, and took different times to complete the scans. 
  • Because of that, LLM scanners also took 1.6-37x longer to scan and cost 6-39x more than Harness AI SAST. 
  • Using LLMs to validate and reason the results of traditional scanners yields comparable precision, and more importantly deterministic findings, cost and time-to-scan. 

Harness recommendation for customers 

  • LLM scanners face unpredictability, latency, and cost challenges that make them a harder fit for scanning regular PRs or code commits today.
  • We recommend using LLM scanners for pen testing and periodic deep scans of critical applications. 
  • For continuous scanning as part of code pipelines, we recommend using LLMs to validate and reason on the results of traditional scanners and to identify novel vulnerabilities.   

Research purpose 

LLM-based scanners are getting better at finding vulnerabilities in code, reasoning through code the way a human reviewer would. We wanted to test if security teams can rely on LLM scanning for scanning regular PRs or code commits. That question matters most for a typical enterprise security team, which doesn’t have in-house expertise to build and tune a custom LLM harness for its environment and have limited token budgets.

This benchmark measures unattended scanning specifically - with no human in the loop to set the scope, ask the right questions, and check the answers. That’s a deliberate scope that excludes the kind of work a security engineer would normally do: triaging a specific finding, tracing a suspicious data flow, explaining why a pattern is dangerous in your framework, or pressure-testing a fix before merge.  

We tested GPT-5.5, Claude Opus 4.8, and Harness AI SAST - judging all options across accuracy, speed, cost, and consistency. We also used the same harness (Open AI Codex Security’s security prompts) across every model tested, to minimize the impact of the harness on the test results. 

Finally, our benchmarking is based on three well-known vulnerable test applications - OWASP WebGoat, OWASP Juice Shop, and OWASP Benchmark for Java. We realize that LLMs are trained on these applications, which impacts their precision and recall. We still chose them because security teams will decide whether or not to use LLMs for continuous scanning not just on their detection accuracy but also their predictability, cost and latency. Using these applications also allows third parties to reproduce our results. 

The challenge with measuring precision

Vulnerability scanners are typically evaluated on two numbers: precision (of the scan findings, how many were actual vulnerabilities?) and recall (of all the vulnerabilities present, how many got found?). Traditional SAST fares poorly in the real world, with high false positive (FP) rates a common complaint. In the benchmark world, LLM-based scanners provide a sharp contrast. Across our testing, they demonstrated precision values from 85% to 100%. But numbers this clean, on applications this well-documented, are worth scrutinizing before taking it at face value.

Measuring LLM scanner precision on WebGoat, Juice Shop, and Benchmark for Java

OpenAI’s GPT-5.5 stands out with a perfect 100% precision on WebGoat and BenchmarkJava, but a slightly lower 86.4% precision on Juice Shop. Anthropic’s Claude Opus 4.8 follows closely on most tests. If you stopped here, you could conclude that LLM-based scanning has solved the precision challenge in vulnerability scanning. It hasn't necessarily, and the reasons why are worth walking through.

Comcast, a Project Glasswing partner, ran Mythos against real-world code and reported 44% false positives (<56% precision), a wide gap from what we saw in our tests. Two different mechanisms could explain that gap:

  1. Familiarity: a model that's encountered this code, or write-ups describing its vulnerabilities, during training doesn't need to reason carefully to score well; recognizing the application may be enough.
  2. Information leakage: OWASP Benchmark for Java ships its own answer key inside the repository being scanned. A tool instructed to review every file in scope could read that file and know exactly which test cases are vulnerable. WebGoat and Juice Shop have softer versions of the same issue: WebGoat’s source is organized around named, labeled security lessons, with some files containing comments explaining how to exploit the vulnerability they demonstrate, and Juice Shop ships extensive public documentation describing its intentionally vulnerable behavior.

WebGoat, Juice Shop, and BenchmarkJava are well-known, publicly documented vulnerable applications either way. But even between WebGoat and Juice Shop, GPT-5.5’s precision differs by 13 points, suggesting neither familiarity nor leakage is the whole story.

A different challenge with recall

Recall asks the question that precision left out: of every real vulnerability in the application, how many did the scanner actually find? On WebGoat, both GPT-5.5 and Opus 4.8 fell well short. GPT-5.5 reached 26.5% recall (62 TP / 172 FN), while Opus 4.8 reached 16.2% (38 TP / 196 FN). Both models missed more than two-thirds of the vulnerabilities present. Juice Shop tells a similar story - both models posted under 25% recall. And yet, on Benchmark for Java, GPT-5.5 reached 100% recall, while Opus 4.8 managed just 1.1% (16 TP / 1399 FN).

Measuring LLM scanner recall on WebGoat, Juice Shop, and Benchmark for Java

The gap is worse at the severity levels that matter most. On WebGoat, GPT-5.5 caught only 38.5% of critical vulnerabilities and Opus 4.8 caught 33.3%. Juice Shop reverses that order - GPT-5.5 drops to 12.5%, while Opus 4.8 jumps to 37.5%. On Benchmark for Java, GPT-5.5 found all 126 critical vulnerabilities, while Opus 4.8 found zero. Precision, even accounting for its own distortions, gave no hint that any of this was happening underneath.

Comparing true positives and false negatives side-by-side on WebGoat

Part of this may come down to how these models handle uncertainty rather than what they know. A model that only reports what it's confident about will look precise, at the cost of recall. That's a different mechanism from the familiarity effect in the previous section, where a model recognizes a well-documented app and scores well without deep reasoning. The two are hard to tell apart from the outside, but call for different fixes. Familiarity is solvable with harder, less-public test cases, but conservatism would follow these models into unfamiliar code.

Some of that conservatism may be instrumentation rather than incidental. We orchestrated both models with Codex Security’s workflow, which is explicitly designed to validate candidate vulnerabilities before anything is reported, a deliberate step to keep false positives low. A real vulnerability that a model can’t confirm with enough confidence doesn’t make it into the final report, even if the model found it. If that is what’s happening, the low recall seen here may say more about Codex Security’s high bar for validation than it does about any model’s ability to reason through a vulnerability itself.

Same code, different stories

Every scan so far has been a single attempt. But LLMs are fundamentally probabilistic tools. We know that if you run an LLM scanner multiple times on the exact same code, it won’t return the same answer twice. 

To put data behind that statement, we ran Opus 4.8 three times each on three applications, and got three different outcomes - not just three different numbers, but three different shapes of result.

  • WebGoat: real improvement. Recall climbed with each attempt - 16.2% on the first scan, 19.7% on the second, 35.5% on the third. The third scan found more than double what the first one did, on identical code.
  • Juice Shop: up, then back down. Recall jumped from 16.0% on the first scan to 30.0% on the second, then fell to 18.8% on the third, giving back almost all of that gain. Two identical re-runs, two opposite directions.
  • Benchmark for Java: consistently poor. Recall stayed within a point across all three scans. Unlike WebGoat or Juice Shop, more attempts didn't move the needle in either direction; the tool simply missed nearly everything, every time. It’s worth calling out that Opus 4.8 found zero critical vulnerabilities across all three scans.
Recall with Opus 4.8 shifted across multiple scans

There's no way to tell from a single scan which of these scenarios you're in. A team running this once and getting WebGoat's first-scan result (16.2%) would have no signal that a second attempt might nearly double it. A team running it once on Juice Shop and landing on the second scan's 30.0% might reasonably conclude the tool is improving - only to see that gain evaporate on a third try. Only on Benchmark for Java is a single scan actually representative of what you'd get from trying again. 

For this test, we only benchmarked Opus 4.8. 

Getting there with a deterministic scanner

Adding an AI reasoning layer to a deterministic scanning engine can move its precision into the same range LLM-only approaches posted on this benchmark. On WebGoat, tightening Harness AI SAST's confidence filter moves both numbers in a predictable direction:

Filter level Precision Recall
All findings 72.9% 54.8%
Baseline (no AI reasoning) 83.0% 46.5%
Exclude contextually safe 83.5% 44.0%
Confirmed risk only 85.1% 40.3%

At its most filtered setting, Harness's precision (85.1%) lands in the same range the LLM-only approaches posted on this benchmark, without requiring the model to have seen the application before, and without the corresponding drop in recall those approaches showed. Each step here trades a few points of recall for a few points of precision, on the same underlying set of findings. The AI reasoning layer gives the Harness deterministic scanner a tunable path into LLM-competitive precision territory, which a raw LLM pass doesn't have any equivalent lever for. Speed and cost are where the two approaches diverge more clearly.

Slow, for no predictable reason

Every LLM-based scanner in this benchmark took longer to run than AI SAST. On WebGoat, Harness finished in 87 seconds; the fastest LLM approach took over 30 minutes, and the slowest took nearly an hour. That gap doesn’t hold everywhere: on Benchmark for Java, the relative gap narrowed to 1.6x with Opus 4.8, while reaching almost 8x for GPT-5.5.

GPT-5.5 was the slowest of the two LLM approaches on Benchmark for Java, by a wide margin - nearly three hours, compared to 43-53 minutes on WebGoat and Juice Shop. That's the one case in this benchmark where scan time scales with codebase size: Benchmark for Java is the largest application tested, and GPT-5.5's scan time grew to match. Opus 4.8 shows no such relationship. Its slowest run (Juice Shop, 58 minutes) wasn't its largest codebase, and its fastest (WebGoat, 30 minutes) wasn't its smallest.

Every LLM-based scanner had dramatically longer scan times

It's tempting to read this as "more reasoning takes more time," and that may be part of what's happening, but the data doesn't show a clean, predictable relationship between codebase size and scan time for either tool, apart from GPT-5.5's jump on Benchmark for Java. A team choosing an LLM-based scanner can't assume scan time will scale with the size of what they're scanning, in either direction. What is predictable is Harness's scan time by comparison: fast on every application tested, and dramatically faster than any LLM approach regardless of which one, or how it scales on its own.

Longer Scan Times Mean Higher Costs

Scanner WebGoat Juice Shop Benchmark for Java
GPT-5.5 $27.20 $21.79 $133.58
Opus 4.8 $29.54 $60.66 $47.17
AI SAST $3.40* $3.40* $3.40*

Token cost tracks scan time closely for both LLM approaches. The longer a scan runs, the more it costs, regardless of which tool. GPT-5.5's cost and scan time rank in the same order on every application: cheapest and fastest on Juice Shop ($21.79), most expensive and slowest on Benchmark for Java ($133.58). Opus 4.8 shows the same relationship: cheapest and fastest on WebGoat ($29.54), most expensive and slowest on Juice Shop ($60.66).

*None of these costs apply to AI SAST. There's no per-scan token bill to run up and no cost that climbs if a scan needs to be repeated, but rather a predictable monthly fee - which the table above amortizes over an average number of customer scans per month. 

The fix for both

The last two sections leave an obvious question. If LLM-based scanning is this slow, this expensive, and this unpredictable, is there any way to actually run it in a real pipeline? That's a separate question from the research above, but we wanted to explore it anyway. We changed what a model is asked to do on each run: instead of re-reasoning over an entire repository from scratch every time, an incremental scan loads the prior scan's findings, re-validates each one against the current code, and limits new discovery to what actually changed.

We benchmarked this directly on three different test applications - dvpwa, govwa, and vulpy - running GPT-5.5 (258K context window) both as a full scan and as an incremental scan covering the same commit range. 

On dvpwa, a full scan cost $26.20 and took 50 minutes; the incremental scan covering the same range cost $4.51 and took 16 minutes - and found more, not fewer, real issues (20 versus 16), with every finding from the full scan retained. govwa showed the same shape: 83% cheaper, 72% faster, with every finding retained. Vulpy tells a different story: cost savings were smaller (58%), and incremental scanning missed 9 of the 13 rule categories the full scan caught. Its commit range included an unusually large change, over 44,000 lines added across just 3 commits, which likely pushed the 'only look at what changed' assumption past where it holds.

Because incremental scans get cheaper and faster as volume increases (the fixed cost of the initial full scan gets spread across more runs), the savings compound. At 100 scans, that's the difference between $2,620 and roughly $473.

Two of three repositories show incremental scanning retaining full coverage at a fraction of the cost; the third shows that assumption can break down when the "incremental" change isn't actually small. 

This addresses the same instability "Same Code, Different Stories" documented, from a different angle: not by making the model consistent, but by giving it a narrower question to answer each time, and less room to be inconsistent about - provided that narrower question is actually narrow.

Methodology

Scanners and versions

  • GPT-5.5, via OpenAI Codex Security's security skills
  • Claude Opus 4.8, via OpenAI Codex Security's security skills
  • Harness AI SAST - unversioned SaaS, scans performed Aug 18-20, 2026

Test applications and commits (precision/recall/speed/cost benchmark)

  • OWASP WebGoat - commit c596b729866792d0e5b73330948713d5852b5ec6
  • OWASP Juice Shop - commit 1618a611b173b4bf114028e6e02549950606e29d
  • OWASP Benchmark for Java - commit 0b6b371ad4ae444dd16498f5a69edef8317b3e46

Test applications and commits (incremental scan orchestration benchmark)

  • dvpwa - base 47d7ea5 to HEAD fe2ee16 (24 commits in range, +379/-69 lines)
  • govwa - base dbf4fe5 to HEAD 4058f79 (25 commits in range, +220/-171 lines)
  • vulpy - base 567663f to HEAD 5249cc8 (3 commits in range, +44,620/-78 lines)
  • Model: GPT-5.5, 258K token context window, via OpenAI Codex Security's deep-security-scan skill
  • Cost formula (GPT-5.5 as tested): fresh input tokens x $5/M + cached input tokens x $0.50/M + output tokens x $30/M (GPT-5.5 short-context pricing)
  • Cost formula (Opus 4.8 for reference): fresh input tokens x $5/M + cached input tokens x $0.50/M + output tokens x $25/M

Ground truth

  • For OWASP Benchmark for Java, an official expected-results file provides labeled test cases with known CWEs - this is used as-is
  • For OWASP WebGoat and OWASP Juice Shop, ground truth is built iteratively: it starts as an empty or partial list and gets enriched as scanners' true-positive findings are confirmed and added
  • Some additional entries found by human and LLM-aided analysis were added manually
  • Each ground-truth entry records a vulnerability ID, CWE-based bug class, description, file locations, and CVSS 3.1 severity

Matching findings to ground truth

  • For OWASP Benchmark for Java: matching is deterministic - test case IDs are extracted from findings and checked against expected CWEs, with some per-category CWE flexibility
  • For OWASP WebGoat and OWASP Juice Shop: an LLM matches scanner findings to ground-truth entries in batches, using pre-filtered candidates based on file overlap and semantic category similarity (CWE mappings and keyword hints)
  • The LLM is instructed to match on the underlying vulnerability mechanism, not just co-location in the same file or line
  • When multiple findings from a single scanner match the same ground-truth entry, the LLM ranks them. The best match is kept, other findings are either reassigned to related unmatched ground-truth entries or new entries are created as needed

Classification of unmatched findings

  • Findings that don't match any existing ground-truth entry are classified by the LLM as true positive (new ground-truth entry created), false positive, or non-security
  • Scanners are processed sequentially so earlier scanners' confirmed true positives enrich the ground truth before later scanners are evaluated
  • After the full pipeline, unmatched ground-truth entries are audited and classified as duplicates of matched entries, invalid, or genuine missed vulnerabilities (false negatives)

Metrics

  • Non-security findings are excluded (neither counted as true positive nor false positive)
  • Precision = TP / (TP + FP) - what fraction of reported findings are real
  • Recall = TP / (TP + FN) - what fraction of known vulnerabilities were found
  • For OWASP Benchmark for Java, additional standard metrics: TPR, FPR, and Score (TPR − FPR), broken down by vulnerability category
  • Severity is assigned via CVSS 3.1 - either by scanner consensus or LLM judgment when scanners disagree
  • Scan time and token cost were measured directly from each scanner's own logs and session data
← Previous:
Next: →‍

Related Resources

Get Started

Get Started with Harness AI

Try the full platform free. No module restrictions, no credit card.

Gabriel Acevedo
Senior Director, Software Engineering
Hands-on experience in cybersecurity, R&D, and software engineering and product develop- ment focused on cloud first solutions. Experience maintaining, optimizing, enhancing, and rearchitecting existing systems, including mission critical services. Played a critical role in the creation of technologies, products, and solutions from the ground up.
gabriel-acevedo
Gabriel Acevedo
https://www.linkedin.com/in/gabrielacevedo/
Malte Kraus
Principal Software Engineer, Security
Malte Kraus is a Principal Software Engineer for Security at Harness, where he focuses on the vulnerability detection capabilities of Harness SAST and SCA. He brings 15+ years of software security experience and has published research on program analysis.
malte-kraus
Malte Kraus
https://www.linkedin.com/in/malte-kraus-706744135/
Renny Shen
Senior Director, Product Marketing
Renny Shen is a 25-year veteran of the technology industry and has deep experience both building and marketing a broad range of products and solutions.
renny-shen
Renny Shen
https://www.linkedin.com/in/renny-shen