Can a cheap decision model take triage off your team's plate? What happened when Jev finally answered

My last post ended with a model I could not reach. I had set out to test TypeSafe’s Jev, a decision model that answers a narrow question with a probability attached, on issue triage and code review. For days, almost every request came back with “high demand”. So I measured the options I could run, published the method, and waited.

On 2 October the requests started going through. In a little over an hour, Jev answered all 3,700 decisions from the first study, and the next morning another 1,700 for a follow-up test. The bill for all of it came to 23 cents.

Cheap is the easy part. What I wanted to know is whether its confidence can be trusted, because that is what decides how much work you can hand over.

Where we left off

In the first post I sorted decisions into three lanes: no human, review later, and human first. A model earns the no-human lane only if its confidence separates the calls it gets right from the ones it gets wrong, at an error rate a team accepts. I used one wrong call in twenty (5%). Last time, triage came close and code review did not: every lane that opened missed its 5% target, landing at 6 to 8% on new data.

Decision models promise exactly the missing piece: a probability you can act on, at a fraction of a language model’s price. Jev is listed at about four cents per million tokens of input. On my data that worked out at about a twentieth of what a small language model cost per decision.

Availability is part of the evaluation

On my best day of trying, seven requests got through. The refusals came from the gateway before either of the two hosts that run Jev was even tried, which suggests the limit was the capacity allotted to the model rather than one busy server. Then, on 2 October, every request was answered, typically in about 0.6 seconds, by both hosts.

That is not a criticism of a young product that many people want to try at once. But if you plan to put one in your delivery pipeline, days without answers is a risk you have to design for: a fallback, a timeout, and a way to keep work moving when the model says no.

How I tested it

Same questions, same data, same rules. Nothing changed from the first post, and everything was fixed before any Jev result existed. The ground truth is again Terraform’s own 2026 history, 203 issues its maintainers labelled and 400 merged pull requests its reviewers either commented on or approved, with two public research datasets as a cross-check: 1,500 issues and 1,000 code changes. Jev was compared with Claude Haiku 4.5, the local open classifier, and, on a 1,000-item sample of the public data, Claude Opus 4.7.

The test I promised. Last time I said I would check whether narrow, well-designed questions close the gap on code review. One independent test had lifted Jev from 63% to 95% on a different task by splitting one broad question into five narrow ones. So instead of asking “will a reviewer comment?”, I asked six narrow checks about each change, from “could it introduce a bug?” to “is it a trivial change?”, and let a small model, trained only on the calibration examples, learn how much each check matters. I wrote the method down before running it, including the rule for a “safe to skip review” lane I had spotted in the data.

What I found

On Terraform’s own history:

  Jev Claude Haiku 4.5 Local classifier
Issues sorted correctly 92% 92% 80%
Issues it could decide with no human 77%, at 4.6% error 82%, at 6% error 51%, at 7% error
Pull requests judged correctly 49% (86% with six checks) 65% 24%
Cost per 1,000 issue decisions $0.04 $0.92 $0

For comparison, a rule that says “no reviewer will comment” on every pull request would be right 86% of the time.

On your own history, the promise held. Jev sorted Terraform’s issues exactly as well as Haiku, at about a twentieth of the cost. Its no-human lane was set for 5% error on one set of issues and landed at 4.6% on new ones. Across both posts and four models, it is the only lane that kept its promise. Most of its answers came with full certainty, and on this repository that certainty was earned.

On noisy labels, the certainty broke. On the public issues, Jev and Haiku both sorted 72% correctly and agreed on 92% of their answers, including calling about half of the maintainer-labelled questions bugs. Jev was fully certain about more than half the issues and wrong on one in six of those, so no lane opened. Certainty is not a property of the model alone; it belongs to the model and your labels together. It is the strongest argument I have found for measuring on your own history.

Code review moved closer to the expensive model, and is still not ready. On the public code changes, Jev judged 60% correctly against Haiku’s 56%, and the six combined checks reached 64%. On the shared sample, the checks scored 66% against Opus’s 65%, at about a hundredth of the cost. But no model, with or without the checks, could be trusted with a single review decision on its own.

The six checks fixed the over-flagging, not the judgement. Asked once, Jev said “a reviewer will comment” on 62% of Terraform’s pull requests, when reviewers engaged with about one in seven. Combined, the checks brought that down to 2%, which is where the 86% comes from: it is close to the rule that never comments. How well Jev ranked changes from least to most likely to draw a comment barely moved, and with only 17 commented pull requests to learn from, there was not much more to learn. One detail is worth keeping: on both datasets, the check that best predicted a comment was “missing error, validation or edge-case handling”.

The finding I liked best did not survive. After the first run, Jev’s “no comment” answers on Terraform looked clean: only 4% of them needed a comment, against 14% overall. It looked like a lane for taking uneventful changes off the review queue. Tested with a rule written down in advance, it found no workable threshold on the public changes, and on Terraform it skipped 63% of pull requests with 10% of those needing a comment, double the target. A pattern spotted after seeing the data is a hypothesis, not a result.

What I’d do as a CTO

Pilot a decision model on triage where your labels are clean. At a twentieth of the cost and with a lane that held, triage is the obvious place to start. Measure it on your own labelled history before you trust it.

Treat certainty as a claim to test. “Fully certain” was right on one repository and wrong one time in six on another. A model’s confidence is only as good as the match between your question and how your team labels things.

Plan for the day it says no. Young models can be unreachable for days. Keep a fallback model, a timeout and a queue, so work keeps moving.

Keep people on review, and keep measuring. Nothing earned a no-human lane for code review, at any price. Cheaper models may still help order the queue, but they do not replace the reviewer yet.

Write the rule down before you look. My favourite finding fell over the moment I tested it properly. Deciding the threshold and the target in advance is cheap, and it is what separates a result from a hunch.

Try it yourself

The kit runs Jev with one flag, and the six-check version with a second short script. For a repository the size of Terraform, Jev costs a few cents and finishes in about a quarter of an hour; you need an API key for Vercel’s AI Gateway. The limits from last time still apply: your results are only as good as your team’s labelling and review habits.

What’s next

So far I have tested models. Next I want to look at the other half of the system: the harness around the model, and how a team can tell whether a change to its setup actually helps.


Appendix: method, data and limits

The study’s rules, questions, data and metrics were fixed before any results were seen, and nothing changed for Jev. Jev’s answers were collected between 26 September and 3 October 2026, almost all of them on 2 and 3 October, through Vercel’s AI Gateway and served by both TypeSafe and DigitalOcean; all 5,403 decisions were answered, for $0.23 in total. The six-check test was written down on 3 October before any call: six yes/no checks per change, combined by a logistic regression fitted on the calibration split with cross-validation, then scored on the evaluation split exactly like every other model. Opus 4.7 ran on a 1,000-item stratified sample only. On Terraform, a pull request counts as “drew a comment” if any reviewer other than the author left a comment review or requested changes; diffs over 20,000 characters were excluded. The local classifier reads at most 512 tokens.

Data: the NLBSE 2024 issue report classification set; Microsoft’s CodeReviewer data (Li et al. 2022, CC BY 4.0); hashicorp/terraform issues and pull requests created in 2026, fetched through the GitHub API and reported in aggregate only. Models: Jev by TypeSafe, Claude Haiku 4.5 and Claude Opus 4.7, all through Vercel’s AI Gateway; DeBERTa-v3-large zero-shot v2.0 by Moritz Laurer (MIT licence).

This is personal research, done in my own time and on public data only. The views are my own and not my employer’s.

I used Claude to help build the test harness and draft this post

Sources