Resolution mix over time
Automation potential
Tickets by module
Confidence distribution
Estimated cost by model
Estimated cost by module
Estimated cost per day
Tickets per day
Tokens (all measured tickets)
Spend anomalies — each day against its own baseline
The centre line is not zero dollars, it is the trailing baseline — what a day like this one usually costs. So an ordinary day has no height, and the only thing with height is the part that is unusual. Bars above the line are overspend, below is underspend; the shaded corridor is ordinary noise, and a bar leaving it is named underneath. The dark line is the 7-day drift: one loud day is weather, a fortnight of quietly hot ones is money.
How a day is judged unusual — and the four questions this chart deliberately refuses to answer
Two measures, because “we spent more” has two unrelated causes. Daily spend is the budget question: dollars that day, which move mostly with how many tickets arrived. Cost per run is the pricing question with volume divided out — it is what moves when a model swaps, a prompt grows, or the cache stops hitting, and it stays flat on a merely busy day. A spike on the first and not the second means a big day; a spike on the second is a price event, and that is the one worth chasing.
Why per run and not per ticket. Cost is booked to the day the run happened, so a ticket re-triaged a week later spends on the later day. Runs are therefore the only denominator that lands on the same day as its own numerator. Dividing by the tickets that arrived that day would make every heavy-reprocessing day look like a price rise.
The baseline is trailing, robust, and excludes the day itself. It is the median of the comparable days behind this one — median rather than mean, so a single $9 day does not drag up the line that is meant to catch the next one; and strictly excluding today, because a baseline containing today would be pulled toward it and a large enough spike could hide inside the baseline it created. Choose the window with the buttons: 7d is thin enough that one odd day visibly moves what the following days are judged against, 28d takes a fortnight to even notice a step change, and 14d sits between them. All three are whole weeks, so a window never holds 2.14 Saturdays.
A Saturday is compared with Saturdays. Weekend ticket volume is a fraction of a weekday’s — the monthly projection splits the two for the same reason — so one blended baseline would put every weekend far below the line, every week, and the chart would spend its whole vertical range describing the calendar. Each day is measured against the same-day-type days behind it instead. That is why the corridor visibly pinches at weekends: the noise floor really is lower there.
The corridor is ±2σ of that measure’s own noise, where σ is a MAD-based estimate — robust, so the outliers it exists to find cannot inflate it and hide themselves inside it. It is pooled across the whole visible range rather than estimated inside each window, because a window holds only a handful of comparable days and a spread computed from four numbers is not a spread. It is pooled separately for weekdays and weekend days, since one shared figure would either drown every weekend anomaly or call every ordinary weekday unusual.
Four questions it refuses, on purpose. A day with no baseline yet is drawn as a gap rather than a zero — the left edge of any range stays blank until enough comparable days sit behind it. Where there is too little history to estimate noise at all, the corridor disappears and nothing is flagged, rather than everything. Today is drawn but never flagged: an hour into a day, spend is always anomalously low. And a day with fewer than three runs can never be a pricing anomaly — one ticket that happened to carry a large attachment is not a price change, though its dollars are real and do still appear under Daily spend.
The baseline is the last fortnight, not all time. Every figure here compares a day against a trailing 14-day window of comparable days — weekdays against weekdays. So a day can be flagged as high-volume while looking unremarkable on the ticket chart above, because that chart shows months and this comparison shows the fortnight immediately behind the day. Both are true; they are answering different questions. The run counts are stated in the cause column so the comparison can be checked rather than taken on trust.
What drove it: volume or rate. Spend is runs × cost per run, so an unusual day is exactly one of two things — we did more work, or each piece of work cost more. Those have opposite fixes, so the column splits the deviation between them. The split is exact: each factor is held at the midpoint of the day and its baseline, which leaves no interaction term to drop or fudge on the days where both moved, which are the interesting ones.
The sentence underneath names the sub-driver. If volume led, the question is whether the extra runs were new tickets or repeats — runs per ticket rising means retries, replays or update re-triggers rather than more work arriving. If rate led, the candidates are bigger prompts (tokens per run), a wordier model (output per run), or a fallen cache hit rate — a cache write prices at 12.5× a read, so that one raises cost with no extra work done at all, and it is invisible in every other number here. When nothing moved enough to name, it says so instead of picking the largest number regardless.
Model mix is a pointer, not an audit. It compares each model’s cost that day against its own trailing median. Per-model cost is booked to the triage day rather than the run day, so those figures explain the day’s model mix; they will not sum to the deviation. A model swap in particular shows up here as one model up and another down, which is a substitution and not a cause — identically priced tiers net to nothing. Read the volume/rate column for the cause.
Why this replaced the cumulative-spend curve. A cumulative line only ever goes up, so every day looks like the last, and a day that cost triple reads as a slightly steeper segment nobody notices. The total it showed is on the Triage spend card at the top of this page, stated once and exactly — which is all a running total was ever worth here.
Daily cost by model
Cost per ticket (distribution)
Cache-hit rate over time
Tokens over time
Learning-loop cost per day
Spend by source
| Source | Runs | Avg time | Avg turns | Avg / run | Total cost |
|---|
By model
| Model | Runs | Input tok | Output tok | Cost |
|---|
Recent runs i
| When (UTC) | Source | Model | In | Out | Cost | Duration | Run |
|---|
Prediction vs. reality over time
Of the tickets where we can tell what the humans really decided, how often did the bot's recommendation match? Each point is a trailing window ending that day — bold = 28 days, faint = 7 — and the shaded band is the 28-day 95% confidence interval, so band width is sample size. ▲ marks the day a learning-loop prompt change went live; ◆ marks a day one of these signals changed meaning — the classifier moving and the ruler moving look identical on a line chart, and they are not the same news.
How these numbers are measured — why module is plotted twice, and why this reads lower than the weekly report
Module (team routing) — did the ticket reach the right team? The classifier never picks a team directly; it picks a primary_module (a fine-grained, partly technical axis like accounting or oem_integration). We derive from that which operational team the classification implies, and compare it against the swimlane_* tag a human ended up putting on the ticket. So this is a team vs. team comparison, not a module-string match — a ticket can land on a slightly different module and still be counted correct, provided it would have gone to the same team.
Issue type — did we call the kind of ticket right? The bot's issue_type (bug / enhancement / how-to / configuration …) against the class the human's tags imply. A simpler, more direct comparison, which is part of why it usually scores higher.
Agent approval — the share of 👍 among tickets a support agent actually rated. This is the only series where a human judged our recommendation rather than us inferring agreement from tags they happened to apply, which makes it the one number here that the circularity problem cannot touch. Two caveats, both real: rating is voluntary, so it is biased toward the notably good and the notably bad rather than being a fair sample; and very few tickets get rated, so its band is by far the widest on the chart. Treat a movement in it as a prompt to go read the 👎 notes below, not as a measurement. If you want this series to become trustworthy, the fix is more ratings, not more maths.
What counts toward a window. Only tickets carrying usable ground truth. A ticket is skipped when a human never applied a tag we can read, so we genuinely cannot say who was right. Dividing by all tickets instead would mean a stretch where outcome capture ran poorly looked like a stretch where triage performed poorly — and improving capture would show up as a regression.
Why module and issue type are each plotted twice. If a swimlane tag was already on the ticket when triage ran, the classifier was fed that tag as input. It agreeing with the tag afterwards proves nothing — it is an echo, not a verdict. Module — clean (solid) throws those out; Module — raw (dashed) keeps them. The catch is that the exclusion can only ever fire on an agreement — a disagreement is never an echo — so it strips wins and keeps every loss. Clean is therefore the number to trust after a merge, but read alone it looks like the bot's hit rate when it is a rate over a deliberately adversarial subsample; raw is much closer to how often the ticket actually reached the right team. The vertical gap between the two lines is the share of the signal that is circular, which is worth watching in its own right: as it widens, less and less of your traffic is independently measurable. Measured on 2026-09-08 that gap was 71.0% raw against 17.7% clean — 240 agreements excluded against 1 disagreement, so 91% of the raw agreements were echoes. Neither end of that bracket is "the accuracy"; the honest reading is the range, and which end you quote depends on the question.
Issue type is now plotted twice for the same reason. It has carried the identical exclusion since 2026-08-24 and had no raw twin until this was added, so when the guard began landing the only line on the chart slid downward with nothing beside it to say why. The pair makes that legible the way it already was for module.
Group — clean (dotted, same blue) is a third read on the same question as module: the team implied by the bot's module against the team implied by the Zendesk group the ticket sat in. Worth having because its echo path is much weaker — metadata.module maps one-to-one onto the answer space and the prompt makes the model declare whether it overrode it, whereas a group is an org-unit name a model would have to infer a module from. It is not echo-free, though: the group is handed to the classifier too, so it carries the same exclusion. Expect it to be thin and gappy — the support queue spans every team, so it yields no verdict at all rather than a fabricated one — and expect a wide band for a while. It only accrues history if it is drawn.
So it will not match the weekly ai-eval issue. That report's headline module disagreement rate is computed over every ticket with any readable tag, circular ones included, so it reads much better. Both are arithmetically correct; they answer different questions. This chart answers "when the human's decision was independent of ours, how often were we right?" — the harder and more useful question.
Reading the chart. Every point aggregates the days behind it, so the newest point is always current — there is no "still filling" week to discount. Two windows are drawn: 28 days bold, because that is the shortest window where module's interval narrows to roughly ±8pp and becomes worth acting on, and 7 days faint, which responds fast but is noise-dominated at this volume. Both are whole numbers of weeks on purpose — weekend ticket volume is a fraction of a weekday's, so a 30-day window would contain 4.3 weeks and quietly over-weight whichever weekdays fell in it five times. A rate is left as a gap only when nothing was measurable; thin-but-real data shows up as a wide band instead, so "we measured a little and know little" no longer looks like "we measured nothing". Hover any day for the counts, the interval, and the merges that landed on it.
Where the series starts. It begins on the first day anything was measurable, not the first day a ticket was triaged. Outcome capture began months after triage did — roughly 500 tickets from May carry no outcome at all — so an unfiltered view would otherwise open on two months of blank axis and read as a broken chart. Only that leading stretch is trimmed: a gap in the middle is left drawn, because that one means capture or triage genuinely stopped and is worth seeing. The coverage line above the chart always states the true total either way.
Adjacent points are not independent. Two neighbouring 28-day points share 27 of their 28 days. The line is a smoothed level, not a sequence of fresh measurements, so a gentle slope is not evidence of a trend on its own — that is what the band is for. If the band after a merge still overlaps the band before it, the honest reading is "we cannot tell yet".
The module lines stop before 2026-07-20. That day (PR #75) the module signal was redefined: a raw module-string match became a team-vs-team comparison, and the circularity exclusion was introduced at the same time. Outcomes captured before then are in the old frame and are not comparable, so any window reaching back past that date omits both module lines rather than averaging two definitions into a fake step change. They reappear on their own as the older rows age out of the window. Green spans the full range because the issue-type comparison has not changed since 2026-06-16 — but that is not the same as the plotted line never changing meaning: see the next note.
The clean issue-type line steps down on 2026-08-24, and that is the ruler moving. Its circularity exclusion shipped that day; before it, every issue-type agreement counted, circular or not. Because the exclusion can only fire on an agreement, the clean line falls as it lands while the raw line does not move at all — over the 28 days to 2026-09-08 the clean rate went 69.3% → 61.7% while its denominator fell 189 → 120, losing agreements at nearly five times the rate of disagreements. It is marked ◆ rather than gated: the pre-2026-08-24 rows are still valid readings of a less guarded signal, so blanking a month of history to avoid a visible step would cost more than labelling the step. ◆ means the ruler changed; ▲ means the classifier changed. Do not read one as the other.
Same measure as the loop's own verdicts. The per-ticket judgement here is shared code with the post-merge verification that labels an improvement improved / regressed on the agentic-changelog page (PRODUCTION_SIGNALS), so the chart and those verdicts cannot drift apart.
What the module signal was made of
Of every ticket whose outcome we captured, how much yielded a usable module signal — and where the rest went. Trailing 7 days, stepped daily. The warm band is the share a code fix could recover. When the blue band shrinks, every rate on the chart above it gets noisier, so this is the early warning for that.
Did each change actually help?
One row per merged prompt change — agreement before it landed against after. Read the bars, not the dots: where the two 95% intervals overlap, the change has not been shown to do anything, however far apart the two markers look. The trend chart above cannot answer this, because a 28-day window smears several merges together.
Does the bot know when it's wrong?
Its stated confidence against what actually happened. On the dashed diagonal, confidence is earned — a point below the line is the bot claiming more certainty than it delivered, which is the direction that matters. Bars are 95% intervals; a band that crosses the diagonal is not evidence of anything. Ignores the date filter above: calibration is a property of the prompt, and five bands over a fortnight would be all noise.
Time to resolve
The outcome this whole system exists to move. Trailing median with the interquartile band — a median because one ticket parked for three weeks drags a mean somewhere no real ticket lives. Read this as evidence, not proof: the bot recommends but does not act, so a fall here is equally consistent with a quieter ticket mix or a staffing change.
Module agreement (bot vs. human tag)
Resolved (of tickets with outcomes)
Support feedback (👍 / 👎) i
| Step | Model | Runs | % tickets | Avg / run | Est. cost |
|---|