The Queue Looked Fine on Monday
Want to talk about this essay? Email me: amitkvint@gmail.comcopied· or message me on LinkedIn
For a few years at WPML, the support numbers arrived once a week, in a report.
The report was fine. It had ticket volume, average wait time, resolution time, satisfaction. It was the kind of thing you could put in front of the executive team without anyone asking questions.
The problem was the gap between when the report was produced and when we were reading it. By the time a number was on the page, it was describing last week. And we were using it to make decisions about this week.
Who covers the weekend. Whether to move a second-tier supporter down to first tier for a few days. Whether the backlog that just appeared is a blip or the start of something. Those are Tuesday decisions. We were making them with data that stopped on Sunday.
Monday: healthy. Wednesday: underwater.
The moment it stopped being a theoretical problem was fairly ordinary, because it kept happening. Every few WordPress releases, something broke. The one I remember is a release, I think in 2023, that broke the CSS of the WPML language switcher on some sites. Nothing dramatic from the outside, but every one of those sites belonged to someone who could see it.
The Monday report showed a healthy queue. Wait times normal, backlog small. By Wednesday, we were underwater, and nobody had noticed on Tuesday because there was nothing to notice with. Supporters could see their own tickets. Team leads could see their tier. Nobody could see the whole thing, right now.
We got through it the way you usually do: people worked late, someone skipped a day off, a couple of angry customers got personal apologies. And then the next weekly report came out and confirmed, in retrospect, what we had already lived through.
That is the part I keep coming back to. The report wasn't wrong. It was just late. And a late number is a different thing from a wrong number, because everyone trusts it.
The pitch
I was on the executive team, so I took it to the CEO and my peers the way we took everything else: as a proposal with a cost attached.
Six slides. The problem, which was that we were making capacity decisions on week-old data. The evidence, which was the Monday-to-Wednesday story. What I wanted to build and what you'd see on it in real time. The options, including doing nothing. The cost: six weeks of developer time plus two weeks of QA. And the ask.
I think including "do nothing" as a real option mattered more than anything else on those slides. It made the cost of the current situation explicit instead of assumed. Every late-noticed spike had a price in overtime, in resolution time, in the occasional refund. Once that was on a slide next to eight weeks of engineering, the decision more or less made itself.
What we built
A small Python application sitting on top of our support data, refreshing continuously instead of weekly. It wasn't the analytics layer I wrote about in Your Ticket Queue Is a Defect Log. That one ranked recurring issues so we could fix them at the source. This one was about the queue right now.
Otto Wald, one of my lead supporters and a dear colleague, built it. I tested and specified changes. I wasn't writing the code, and I don't think I should have been. What I was doing was deciding what needed to be on the screen, and next to what.
The spec was organised by time.
First, the long view. New reports, chats and tickets counted together, per month for the last three years. For the last month: the average per weekday, the average per hour, the volume per forum, the resolution ratio, resolution times for tickets, chats and both, and how happy users were with support.
Then last week. Tickets per day, chats started, total reports, chats that turned into tickets, positive feedback, resolution ratio, resolution times, first reply time.
Then today. The same kind of numbers again, plus how many converted chats were sitting in the unassigned queue. And next to most of them, a plus or a minus against last week's average.
And then one row per supporter, live. Working or not, hours left in the shift, on chat or not, assigned tickets, replies on tickets today, ongoing chats, tickets waiting on the user, negative feedback today, first reply time today, reports resolved today.
That last part only team leads could see, and it was for moving work around, never for reviews. A live row of numbers about a person turns into a leaderboard the moment it shows up in an evaluation.
So the weekly numbers didn't go away. They moved. The report used to be the news, a week late. On the screen, last week's average became the thing today was measured against, and the monthly averages per weekday and per hour told you what a normal Tuesday afternoon was supposed to look like.
That is the part the old report could never do. It could tell you last week was busy. It couldn't tell you, on a Tuesday, that this Tuesday was already heading somewhere worse.
What changed
I used it every day.
After it went live, a big problem in the queue couldn't hide until the next report anymore. If a few people weren't available that day, or tickets or chats suddenly started piling up, it showed up on the screen as it happened. I could shift focus straight away, ask for help, open the tickets and see exactly what was going on.
The Monday-to-Wednesday surprise didn't happen again.
I don't remember the numbers from that period, so I'm not going to put any here.
The other change was in how fast things reached me. Before, the whole picture existed once a week, in a document, for the people who read documents. After, the team leads were looking at the same live picture, and when something was off, they told me much sooner.
That is the version of "data-driven" I actually believe in. Not more numbers. The same numbers, closer to the moment where someone has to decide something, and sitting next to what normal looks like.
The part that isn't about Python
Tools change. We built this in Python because that was what Otto and I were both comfortable with; today I'd probably have a rough version running on sample data within a day or two with an AI coding tool, and hand Otto something already working instead of a spec.
The goal doesn't change. If you're running a support team and your numbers arrive after the week they describe, you don't have a measurement problem. You have a timing problem, and every decision in between is being made by feel.
Make today's numbers live, and put last week's average next to each one. You don't have to throw the weekly report away. It just stops being the news and becomes the baseline.
And if you're pitching it upward, put "do nothing" on the slide. It's usually the most expensive option, and the only one nobody has priced :)