Rebuild AI API classification with Jev to cut costs, reduce latency, improve calibration, and make production decisions more efficient without sacrificing accuracy.

TL;DR
Replacing the judgment step in a production classification pipeline with a purpose-built decision model changes the economics of the problem entirely.
- The cost gap is not incremental — it is structural. Moving from a generative LLM to Jev for the yes/no decision reduced our per-endpoint classification cost by roughly 1,200× against our original pipeline and 150× against a single-shot LLM. The saving exists because Jev makes a decision without generating text; generation was never the bottleneck, it was the cost.
- Calibrated probability turns confidence into a control surface. Every LLM approach we evaluated reports a self-assessed confidence level that clusters and correlates poorly with accuracy. Jev returns a probability. In our eval, every error Jev made fell in the 0.3–0.8 uncertainty band — just 3.4% of endpoints. At both tails of the distribution, Jev made zero mistakes. That is not something you can do with a self-reported label.
- A stable output contract makes the switch reversible. We replaced the entire decision engine behind four fields without touching a single downstream consumer. The architecture that made this safe is the transferable lesson: if you can place a stable contract in front of a decision, what sits behind it can be changed, improved, and rolled back without coordination cost.
An earlier post on the Harness blog introduced Jev and the idea of treating a model's judgment as a typed, auditable primitive rather than something buried inside a generative call. That post covers what Jev is — the Choice / Score / Noul output types — and why a calibrated probability is fundamentally different from a confidence value an LLM reports about itself. We recommend reading it first; we will not repeat that material here.
Here, we put that idea to the test on a production system. One of the more expensive pipelines in our AppSec platform identifies which of a customer's API endpoints communicate with an LLM, and it prompted a specific question: the majority of what this pipeline pays for is not text generation but a repeated yes-or-no judgment, so what happens if we make that judgment the inexpensive part of the system? We evaluated two approaches to answer it — one of which uses Jev — and this post reports what we learned.
Why AI API classification matters
Nearly every organization we work with is shipping AI-powered features, and very few can produce an accurate list of which endpoints actually send data to a model. That gap is the core security problem.
An endpoint that forwards user-supplied text into an LLM carries a different threat model than a conventional CRUD endpoint. It introduces a prompt-injection surface, it is a potential path for sensitive data to leave the organization, and it may route requests to a third-party model vendor that the security team has never reviewed. None of that can be governed, monitored, or threat-modeled until the endpoint is discovered and its behavior understood. The pipeline's job is therefore inventory: for every API observed in a customer's traffic, determine whether it is an AI API, whether it returns AI-generated content, and — for those that qualify — extract the models, vendors, and the precise locations of prompts and responses within each request.
The classifier produces a small, fixed record:
is_ai_api: bool
returns_ai_generated_content: bool
confidence: "high" | "medium" | "low"
reason: string # human-readable justificationThis record is the central constraint on everything that follows. Dashboards, the downstream enrichment stage, and every product surface that surfaces these findings all depend on these four fields in exactly this shape. We are free to change how the verdict is produced. We are not free to change what the verdict looks like.
The current pipeline: V1
Before discussing how we optimized the pipeline, it is worth establishing what it actually does, because the cost problem only makes sense against that backdrop.
The input is API traffic. Our platform observes each customer's requests as spans — structured records of individual request/response exchanges, including URLs, headers, bodies, query parameters, and status codes. A single endpoint may be represented by thousands of spans over a scan window. The pipeline's task is to reduce that traffic to a per-endpoint inventory record: is this an AI API, does it return AI-generated content, and if so, which models and vendors are involved and where in the request the prompts and responses live.
V1 produces that record through the stages shown below.

The pipeline's first stage aggregates raw spans into per-API statistics and filters the population down to viable candidates — endpoints with enough traffic, and the kind of request content that could indicate AI use, to be worth classifying. For each candidate it selects a small, representative sample of spans and trims them, so that a model sees a compact, uniform view rather than the full raw traffic. V1 then normalizes that sample through two LLM passes before classifying: the first summarizes every individual span into a nine-field description, and the second aggregates those into a single per-API summary. Only then does the classification call read the summary and emit the four fields. Finally, for endpoints classified as AI APIs, an enrichment stage extracts the models, vendors, and prompt and response locations, and writes the completed record to the inventory table that the rest of the product reads from.
The classification decision is the subject of this post. Enrichment sits downstream of it and is shared by every approach we evaluated, so it was never the target of the change.
Why V1 is expensive
The economics are the problem, and they are concentrated in one place. The two summarization passes account for approximately 89% of V1's total cost, and they run before the pipeline reaches a single decision. We pay, in volume, to pre-process traffic for endpoints that are frequently unambiguous — and we do so for every candidate endpoint, regardless of whether it turns out to be relevant. By our internal estimate this amounts to roughly $0.18 per endpoint, or approximately $176 per thousand endpoints. This was a good start that delivered immediate value, but it was not cost-scalable across our fleet of enterprise customers, where endpoint counts far exceed a thousand and scans run on a recurring cadence.
The salient point is that almost all of this spend goes toward judgment — the is-it-or-isn't-it determination — which is precisely the kind of narrow, high-volume decision for which a general-purpose generative model is an inefficient instrument.
Two paths to optimization
Both approaches we evaluated begin from the same step: remove the summarization passes and classify directly from trimmed raw spans (top-3 spans, truncated leaves, capped bodies). They differ in how the verdict itself is produced.

Option A: a single LLM call
The direct approach. Collapse summarization and classification into one LLM call per endpoint — supply the trimmed spans, request the same four-field JSON, and leave extended thinking enabled so the model can reason through ambiguous cases.
This is the lower-risk option. Call volume falls from tens of thousands to roughly one thousand, it relies on the same model family already in use, and because it is a generative model it produces the reason string natively. The trade-off is where the cost migrates. With summarization removed, the classification call is now dominated by thinking-token spend, and the thinking budget was inherited from V1's more demanding task. In effect, the model deliberates over endpoints that a simple rule could resolve immediately. The confidence field also remains a self-assessment — the same self-reported high/medium/low that, as the earlier post demonstrated, tends to cluster and correlates poorly with actual accuracy.
Option B: Jev for the decision, an LLM for the explanation
The second approach divides the work according to what each tool does well. Jev renders the decision; an LLM writes the explanation. The field mapping is nearly one-to-one.

is_ai_api and returns_ai_generated_content are yes/no determinations, so each becomes a Noul question and returns a calibrated probability that we threshold in our own code. confidence is no longer the model's self-assessment; it derives from Jev's probability distribution — a measurable quantity we can gate on. This is a genuine improvement over both LLM approaches rather than a lateral move: a field that today means "the model expressed high certainty" comes to mean something quantifiable.
That leaves reason, which is where the approach requires a deliberate design decision.
The constraint Jev introduces
Jev produces typed judgments, not free-form text. Our output contract, however, includes a human-readable reason field that surfaces directly in the product for every flagged endpoint, and which Jev cannot generate on its own. Rather than push the model beyond what it is designed to do, we treated this as a division of labor.
Our solution retains an LLM, but only for narration, and only where the cost is justified. Approximately 4.5% of endpoints — the AI-API prevalence we observed in this corpus — are actually AI APIs, and those already trigger a downstream LLM enrichment call. The reason for a positive can be generated as part of that existing call at negligible marginal cost. Negatives receive a templated reason derived from Jev's own signals: no prompt-like fields, no recognized AI vendor host, low is_ai_api probability. An alternative policy would generate prose only for low-confidence cases, targeting the endpoints most likely to warrant human review. In either design, the expensive generative step moves off the critical path of judgment and onto the narrow path of explanation.
The principle we return to is this: the goal is not to remove the LLM, but to reassign it from making every decision to explaining the decisions that merit explanation.
The comparison
Because all three approaches emit the same four fields, they can be compared directly. The eval set comprises 500 stratified endpoints drawn from Zendesk production traffic (tenant 5a3f7192, cross-date pool, most-recent run per API). The positive class — AI APIs — represents 4.49% of the full classified population (43/957 unique endpoints). Labels were seeded from the current single-shot pipeline and hand-verified by the team, with particular attention to the 20 mixed-signal and ambiguous rows. Jev was evaluated on the same frozen trimmed-span inputs the single-shot pipeline reads from.
Cost
Where does the 1,200× come from?
- V1 makes one LLM call per span, not per endpoint. With ~5 spans per API, 1,000 endpoints produce ~5,000 span-summary calls, ~1,000 aggregation calls, and ~1,000 classification calls — roughly 7,000 total. Each call processes raw request and response bodies with thinking tokens enabled. Cost:
~$0.176 per endpoint. - Jev bills input tokens only — no output, no thinking budget. At
$0.042per million tokens and a measured average of 3,445 input tokens per call, that comes to~$0.00015 per endpoint. - Ratio:
0.176 ÷ 0.00015 ≈ 1,200×. V1 runs a text-generation pipeline before reaching a decision; Jev skips generation entirely. Fewer tokens billed, and billed cheaper.
Single-shot LLM sits between them: one call per endpoint instead of 7,000, but still with extended thinking enabled. At roughly $0.022 per endpoint it is ~8× cheaper than V1 and ~150× more expensive than Jev.
Latency
V1 latency is not directly comparable and is omitted from this table. Unlike the other two approaches — which make a single LLM call per endpoint — V1 issues a batch of calls per endpoint across multiple passes. A per-decision latency figure would not be meaningful in the same way.
Accuracy
V1 is our established baseline — it has been running in production for multiple enterprise customers and its classification results are trusted. We treated V1 as the accuracy benchmark for this comparison: any replacement needs to match what V1 already does well.
Measured on is_ai_api and returns_ai_generated_content against the 500-endpoint human-verified eval set (43 positives):
Both optimized approaches match V1's accuracy tier. The cost and latency advantages come without a meaningful accuracy trade-off.
Confusion matrix, is_ai_api:

Calibration of Jev's is_ai_api probability:

The chart confirms that the probability means what it claims. At both ends of the scale — below 0.3 and above 0.8 — Jev's classifications are consistently correct. The uncertain middle band (0.3–0.8) is where all the errors live, and that band covers just 17 endpoints, 3.4% of the total. This is what makes the probability a useful operational signal rather than a label.
What Jev's probability is really saying. When Jev evaluates an endpoint it returns a number between 0 and 1 — not a categorical yes or no, and not a self-reported confidence label. A score of 0.04 means Jev is highly confident the endpoint is not an AI API. A score of 0.96 means it is highly confident it is. Scores in the middle — say, 0.5 — mean Jev genuinely cannot tell.
The key finding is that Jev's expressed uncertainty is predictive of where its errors actually occur. Every one of the seven errors (4 false positives, 3 false negatives) came from endpoints that scored between 0.3 and 0.8 — the region where Jev itself was unsure. That band covers only 17 endpoints across the 500-endpoint eval — 3.4% of the total.
Outside that band, Jev was perfect. The 448 endpoints it scored below 0.3 (highly confident not AI API) are all true negatives — zero errors. The 25 it scored above 0.9 (highly confident is AI API) are all true positives — zero errors. The model's uncertainty and its actual error rate are correlated, which is what a well-calibrated probability is supposed to do. An LLM's self-reported high/medium/low cannot make this claim: it is a label the model attaches to its own output, not a probability derived from the distribution of outcomes.
The trade-offs at a glance
Our assessment
We approached this experiment with a fair amount of skepticism — replacing a general-purpose model with a specialized decision model is the kind of change that looks attractive in a cost model and can quietly erode accuracy in production. The numbers surprised us. Here is what we take away.
- The cost gap comes from eliminating generation, not from a better prompt. V1 spends
~$176per thousand endpoints running a text-generation pipeline before reaching a single decision. Jev makes the decision without generating text at all: at$0.042per million input tokens and 3,445 avg tokens per call, that works out to roughly$0.14per thousand endpoints. Even single-shot LLM, already ~8× cheaper than V1, is still ~150× more expensive than Jev.
- Accuracy held. Against human-verified labels, Jev makes 7 errors across 500 endpoints. All three approaches operate at a similar accuracy tier (~98–99%). For an AppSec use case the false negatives matter most — a missed AI API is a governance blind spot — and three missed endpoints in 500 is a rate we find acceptable, especially given what the calibration tells us (see below).
- Jev's probability surface tells you exactly where to look. Every one of those 7 errors came from the 17 endpoints (3.4%) that scored in the 0.3–0.8 uncertainty band. Outside that band Jev made zero mistakes — 448 confident negatives and 25 confident positives, all correct. A self-reported
confidence: highlabel cannot make this claim.
- The uncertainty band is where to spend the LLM budget. We are deploying Jev with a confidence-gated fallback: endpoints where Jev's probability falls between 0.2 and 0.8 are routed to a single-shot LLM call. In this corpus that is ~17 extra calls per thousand — the LLM is used precisely where it earns its cost, and not a call wider.
- Single-shot alone is a valid, simpler option. Drop the two summarization passes, point the classification call at trimmed spans, nothing else changes. The ~8× cost reduction is real and the implementation risk is low. What you do not get is a probability you can threshold — the model's
confidencefield stays a self-reported label.
- The contract pattern is the transferable lesson. We replaced the entire decision engine behind four stable fields without touching a single downstream consumer. Dashboards, the enrichment stage, every product surface reading the inventory table — none of them changed. That was not accidental. For any team running an expensive LLM pipeline that is fundamentally doing judgment: place a stable contract in front of the decision, and what sits behind it can be changed, improved, and rolled back without coordination cost.
For background on Jev's decision primitive and its role in agent governance, see the earlier post on the Harness blog.

