Skip to content
Back to blog
EvaluationA quality gate shown openly, failure included

Why 0.828 remains visible beside a 0.850 target

In resident ticket triage, the first category-routing run missed its threshold. We kept the result, documented the errors, and published the later passing run beside it.

The product screen keeps the full measurement history: the first below-target result, its errors, and the later run after changes to the system.

5 min read

The screen of our ticket triage tower shows the quality-gate history above the case queue. The first category-routing run scored 0.828 on the combined precision-and-recall measure (macro-F1), below the 0.850 target. The duplicate-detection result of 1.00 and a later routing run that scored 1.000 also remain visible.

What exactly failed

The first recorded run covered 24 category-routing cases. Four went wrong: two roof tickets landed in plumbing, one heating ticket in plumbing, one safety ticket in construction. Every error is described in the evidence case by case, with the ticket's text and the wrong classification.

We did not recompute the score or change the test set after seeing the result. We diagnosed the errors, refined the taxonomy, reduced the influence of the category entered by the resident, and added narrow rules for unambiguous signals. A later 24-ticket evaluation reached a macro-F1 of 1.000. The two runs used different test sets, so the later result establishes that its own gate passed but cannot demonstrate a production improvement. Both results remain visible.

Why the failed result remains visible

Prepared demonstrations often show only successful results. Keeping a failed quality gate in the product record gives buyers information about how the team responds when a measurement misses its target.

Keeping the quality gate and run history on the product screen also makes measurement part of daily operation. A failed result remains in the record after the next run.

Architecture around an imperfect model

The first red routing tile did not make the system useless, because the architecture assumes the classifier's fallibility. The resident's original text is never overwritten and always stands next to the result. A category correction by the dispatcher is part of the flow, with both values preserved. Safety tickets go to the top of the queue regardless of classification, and the contractor decision always belongs to a human.

Duplicate detection, catching the same failure reported in different words, found all 18 prepared duplicate pairs and merged none of the other 18. The full six-call measurement cost about PLN 0.04, or less than PLN 0.01 per ticket. That is where much of the operational value sits, because an uncaught repeat can mature into an incident over weeks.

How to interpret macro-F1 in an operational queue

Macro-F1 measures every category separately and gives each equal weight. That helps prevent a rare safety category from disappearing under a large volume of routine plumbing tickets. A single score still cannot tell you which errors carry the greatest operational risk.

Pair it with at least four views:

  • a confusion matrix showing which categories are mistaken for one another,
  • performance for safety-critical categories on their own,
  • the share of tickets held for dispatcher review,
  • the time required to inspect and correct a proposed route.

A strong aggregate score can still hide an unacceptable process when the few missed cases involve gas, fire, or building access. The release threshold should follow the cost of each failure class, rather than the desire for one green metric.

Prevent a test result from becoming a rehearsal score

The evaluation after a change should contain fresh tickets that were not used to refine the taxonomy or rules. Otherwise the run measures the project team's familiarity with the examples rather than performance on the next intake.

The test also needs the language and channel mix of the real queue. Residents report the same failure by phone, email, and portal, with abbreviations, spelling mistakes, and building-specific names. A clean synthetic set omits much of that difficulty.

Duplicate detection deserves a separate control set. Incorrectly merging two incidents can hide a new failure beneath an existing work order. Include similar descriptions from different locations and near-duplicate reports from the same building at different times.

What to measure in a pilot

Set these items before automated routing is enabled:

  1. the categories, their owners, and the precedence rules,
  2. cases that rise to the top of the queue regardless of model output,
  3. separate release thresholds for routine and safety-related reports,
  4. the maximum wait for dispatcher review,
  5. a manually reviewed sample after launch,
  6. the condition that disables automated routing when quality degrades.

Alongside macro-F1, track time from receipt to first decision, correction volume, merged tickets later split by a dispatcher, and the age of the oldest open case. Those measures show whether classification improves the shift or merely moves work into a correction screen.

What to ask a vendor

If a vendor shows you only green metrics, ask which uncertain or failed result changed a release decision. The answer shows whether thresholds and measurement are part of engineering work or merely presentation.

Also request the definition of the post-change set, the confusion matrix, category versioning rules, and the procedure for rolling automated routing back. A single score is insufficient for a production decision.

The full measurement history, including the first red result and the later run, is on the case page. A free process scan is a good place to discuss quality thresholds before implementation begins.

Free process scan

Start with a free process scan.

  • A 30-minute call with the engineer who would lead the work.
  • A review of the processes that cost you the most time and money.
  • A written summary of what to automate first and the likely cost range.

The scan is free and creates no obligation. If automation is unlikely to pay off, the written recommendation will say so.

€0

30 minutes · written takeaway within 2 business days

Book a free process scan (30 min)

Times are shown in your own time zone. We work with clients across time zones.

Describe the process in the form