Is it the model or your labels? What it takes to hand issue triage to AI

Coding agents have made writing code cheaper, and the queue of things to sort has grown with it. More teams now let a model do the first pass of issue triage: what kind of issue this is, which team owns it, how urgent it is. Several of the busiest open-source projects I looked at already do.

Before you hand that work over, you measure whether a model can do it, and you measure against your labels. Those labels are not just answers. They carry your definitions, your organisation chart and, more and more often, the output of an earlier model.

So I wanted to separate two things that usually get blurred: the part of triage a better model can fix, and the part no model can fix because it was never written down.

Why labels

In my previous two posts, the labels kept deciding more than the model did. On a public benchmark, every model called about half of the maintainer-labelled questions bugs, because reporters had filed them through the bug template. That is not new: a study of more than 7,000 issue reports from five open-source projects found a third of those filed as bugs were something else.

I set out to test the labels head on, on a project that looks like a typical SaaS codebase, and ran into the problem straight away. The first three projects I checked, Metabase, Sentry and a large observability project, all use a bot to set some of their labels when an issue arrives. Of about twenty others I screened, none combined the volume I needed, triage on GitHub and labels set by people alone.

So I chose Metabase, an open-source analytics product with a hosted version, for four reasons. It looks like the codebases most of us run: a web product with a backend, a frontend and customers. Its community opened about 2,000 issues in nine months. It triages on GitHub, with labels for type, owning team and priority. And its bot marks every issue it touches, and people remove that mark when they review, which leaves a trail of which labels a person actually confirmed. That trail let me change the question to fit: I scored the models only against labels a person had confirmed.

How I tested it

Nine months of one project’s history. Metabase opened about 2,090 issues between January and September 2026. I kept the 1,526 that had been triaged and read each one’s label history: who added every label, and when. The bot had touched 85% of them. When a person reviews the bot’s work, they remove its marker and correct what is wrong; that happened on 256 issues.

Only labels a person confirmed. A label counts if a person set it, or if a person reviewed the issue and left it in place. Everything else, a bot’s or a template’s label that nobody looked at again, is left out. That left 289 issue types, 389 teams and 334 priorities. It is not a typical slice: people step in mostly to correct the bot, and the confirmed types are nine in ten bugs. I also scored everything against all labels as a cross-check.

Two ways of asking. Each model answered three triage decisions, type, team and priority, using Metabase’s own labels and their descriptions as the options. For issue type it also answered the generic question from my first post (bug, feature or question), so I could see what the team’s own labels add.

Three models, and two ways to bring in what the labels lack. I compared TypeSafe’s Jev, a decision model, with Claude Haiku 4.5 and Claude Opus 5.5. After the first results I added two follow-ups, written down before any call. For teams, Jev only names the product area an issue concerns, and a simple table learned from the training issues maps areas to the teams that own them. For priority, Jev answers nine yes-or-no checks taken from Metabase’s own priority definitions. Everything cost $12.58, of which Opus was $7.57.

What I found

On labels a person confirmed:

  Always the most common label Jev Claude Haiku 4.5 Claude Opus 5.5 Jev, learning from history
Issue type, own labels 90% 94% 94% 95%  
Team 32% 24% 18% 30% 63%
Team, learned January to April, tested May to September 33% 20% 17% 25% 46%
Priority 44% 49% 45% 61%  
Cost per 1,000 decisions $0 $0.05 $1 $6 $0.06

Your own labels did not help, and sometimes hurt. On issue type, the generic three-way question did as well as Metabase’s eight labels or better; scored against all labels, they cost Jev and Haiku about four points each. Most of the extra mistakes went to a label with no description, picked for bugs and feature requests alike. More options with thin definitions give a model more ways to be wrong. Type is also the easy decision here, for an unflattering reason: the label mostly reflects the reporter’s choice of template, and the models read the same framing the reporter chose.

Asked cold, no model routed issues better than “send it to the biggest team”. Metabase has sixteen team labels, most without a description, and names that say little about what each team owns. The strongest model picked the right team three times in ten, slightly worse than always choosing the largest team. One model kept choosing teams whose labels say, in their own description, that they are no longer in use. That is not a criticism of Metabase’s labels, which are written for people who already know the teams; it is exactly the point. In one trial answer Opus picked up a clue I had not expected, another team’s ticket number mentioned in the issue, and followed it to the wrong team. It is a single case from a twelve-issue trial, so read it as an illustration rather than a finding: some of the organisation leaks into the text, but not reliably.

History did what model size could not. Asking only which product area an issue concerns, then learning from 116 confirmed examples which areas each team owns, got the team right 63% of the time, at about a hundredth of Opus’s cost. Much of that comes from the history rather than the model: built from the area labels already on each issue, the same table is ten points worse on confirmed labels and just as good on all labels. The honest test is the future. Learned on January to April and tested on May to September, it fell to 46%, still well ahead of Opus at 25%. Team ownership moved during the year: one team label was still being applied in the spring and has since been retired in favour of another, and four are now marked as no longer in use. Ownership drifts, so a table learned from history needs relearning.

A stronger model helped where the definitions were written down. Metabase’s priority labels each come with a clear definition, and there Opus was clearly best: 61% exactly right, almost never more than one level out. The smaller models put most issues at P1, the second-highest level: Haiku two in three, where maintainers chose it for one in three. Opus did so for one in four. Splitting the definitions into nine checks made Jev less alarmist, but its accuracy gain did not hold on later issues. Still, no model was good enough to set a priority without a person.

Whose labels you grade against changes the verdict. Scored against the labels the bot set and nobody reviewed, Jev looked twice as good at routing as against labels a person confirmed: 47% against 24%. Part of that gap is that people step in on harder issues. Either way, a team that grades a new model against its existing labels is partly grading it against its old one. When a person did review Metabase’s bot, they kept its team half the time and its priority three times in four.

What I’d tell a fellow technical leader

Write down what each team owns. Routing failed because the knowledge was missing, not because the model was weak. A short, current description per team, in words a newcomer would understand, is the cheapest improvement available, for people and models alike.

Prune and define your labels. Remove labels you no longer use, merge the ones that overlap, and give every remaining one a definition. Each vague option is another way to be wrong, and a model will happily use the label you retired.

Record who set each label. If a bot labels your issues, mark what it touched and let people mark what they checked. Then measure any new model only against labels a person confirmed; otherwise you are mostly measuring agreement with the last model.

Learn routing from your history, and relearn it. A few hundred issues with a confirmed owner beat the strongest model I could run, at a fraction of the cost. Teams change, so refresh the mapping when they do and check it on recent issues, not old ones.

Pay for a bigger model where judgement is the gap. Priority was the one decision where the strongest model clearly earned its price, because the definitions were there for it to apply. Keep a person on priority for now, but that is where a stronger model buys the most.

Try it yourself

The kit now includes the taxonomy test. A short config file names your label groups and, if you use one, your bot’s marker; one script fetches your issues and their label history, another runs the models, and a third scores them. Every GitHub answer is kept locally, so you can re-run the analysis without asking GitHub again. For about 1,500 issues, Haiku costs about $5 and Jev a few cents. You need a read-only GitHub token and a key for Vercel’s AI Gateway. The usual limit applies with more force here: the results are only as good as the labels you can trust.


Appendix: method, data and limits

The study was written down on 3 October 2026 before any model ran, and every change after that was recorded before the affected runs: Metabase was briefly set aside when its bot came to light and taken up again with the confirmed-labels rule; a reference comparison with the bot was corrected after I found it measured my rule rather than the bot; and Opus 5.5, which often writes a sentence before its answer, was given room to finish after it lost half the answers in a twelve-item trial. Jev and Haiku ran on 3 October and Opus and the follow-ups on 4 October, through Vercel’s AI Gateway; every one of about 14,000 decisions was answered, and four of Haiku’s answers could not be read and count as wrong.

About 30% of issues were used to set thresholds and fit the follow-ups, the rest to score them. A label is “confirmed” if a person set it, or if a person removed the bot’s marker and the label stayed; I read that marker’s removal as a review from the label histories, not from Metabase’s documentation. The time split and the area-label comparison were added after the results and are exploratory. Of the predictions written down in advance: the team’s own labels did not beat the generic question; no model could route teams without a person checking; the two-step routing beat one question by a wide margin; Opus beat Haiku on team and priority; the priority checks made Jev less alarmist but did not clearly improve accuracy.

Limits: one project; the confirmed labels lean towards bugs and towards the bot’s mistakes, so the figures describe reviewed issues rather than all issues; team and priority may depend on information outside the issue text, such as who reported it or the roadmap.

Data: metabase/metabase issues created in 2026, fetched through the GitHub API and reported in aggregate only; for Sentry and the observability project, only the label histories of a handful of issues, to see who applied their labels. No issue authors or reviewers are named. Models: Jev by TypeSafe, Claude Haiku 4.5 and Claude Opus 5.5, through Vercel’s AI Gateway.

I am not affiliated with Metabase or Sentry, and neither has reviewed or endorsed this post.

This is personal research, done in my own time and on public data only. The views are my own and not my employer’s.

I used Claude to help build the test harness and draft this post.

Sources