There's a word researchers started using last year that we can't stop thinking about, because it names something we'd been watching happen without quite having a label for it. The word is workslop, and it describes AI output that looks finished and professional on the surface but turns out to be hollow underneath, the document that reads well until you need it to mean something, the summary that misses the one fact that mattered, the email that sounds exactly right and somehow says absolutely nothing. A team out of Stanford and BetterUp studied it and found that around 41 percent of workers had received some in just the past month, and that each instance ate close to two hours of cleanup.
We pay close attention to this because our whole field, AI customer service, is one of the easiest places in the world for workslop to take root, and we'd be kidding ourselves to pretend otherwise. A customer asks a question, an AI generates a reply that is fluent and confident and structured beautifully, and at a glance it looks like great service. But fluent and correct are not the same thing, and a reply that's polished and wrong is arguably worse than no reply at all, because it sends the customer off in the wrong direction feeling reassured, which is the most expensive kind of mistake you can make with someone who trusted you.
The researchers had a phrase that stuck with several of us, the description of workslop as something cheap to produce and expensive to detect, and that asymmetry is the heart of why it's so dangerous in support specifically. The surface quality is high enough to sail past a first glance, which means it gets sent, and the failure only shows up later, on contact with a real customer making a real decision based on it. By then the damage is done in a way that's hard to trace, the customer who got bad information and acted on it, the complaint that arrives a week later, the trust that quietly erodes without anyone being able to point to the moment it broke.
What makes this more than an abstract worry is what we know about how customers already feel. Surveys keep finding that only around 44 percent of consumers trust AI to handle their service needs, while a much larger share of service professionals assume customers trust it, a perception gap that should make all of us nervous. People are already primed to be skeptical, already half-expecting the bot to be confidently useless, and every piece of workslop they receive confirms that suspicion and hardens it. You don't get many chances to prove you're different, and a single empty, plausible-sounding reply can spend that chance in one go.
The instinctive fix, and we understand the appeal, is to slow down and have a human check everything, but that quietly defeats the entire point of using AI in the first place, and it doesn't scale past a certain volume anyway. The better answer, as far as we've been able to figure out, is upstream, in making sure the AI is actually anchored to real information rather than improvising. Workslop is fundamentally what happens when a system generates language without being tied to truth, when it's optimizing for sounding right instead of being right, and the way you fight it is by refusing to let it answer from nothing.
This is where the technical choices behind a tool stop being trivia and start mattering to the customer in a very concrete way. A system that pulls its answers from your actual policies, your real prices, your genuine hours, and refuses to wander beyond them is structurally much harder to turn into a workslop machine than one that just generates fluent text and hopes. We're not claiming any approach is foolproof, because nothing is, and we've seen grounded systems still stumble. But the difference between an AI that has to cite your real information and one that's free to invent something that sounds plausible is enormous, and it's mostly invisible until a customer gets burned.
There's an honesty point buried in here that our field has been slow to face, and we're going to be plain about where we've stood, because we didn't arrive at this view recently. A lot of AI customer service got sold on speed and volume, on how many tickets it could deflect and how fast, and those metrics quietly reward workslop, because a fast confident wrong answer counts as a deflection too. The customer went away, the ticket closed, the dashboard looked great, and nobody measured whether the person actually got helped or just got dismissed efficiently. We built around accuracy from the first line of code and deployed it in enterprise production as early as May 2024, before "deflection" became the number everyone wanted to brag about, and we've said since then that it is one of the most misleading figures in this business. Whether the ticket closed was never the question. Whether the customer left satisfied always was, and we put a money-back guarantee behind that position.
So the number we trust is CSAT, asked of the customer once the conversation is done, because it is the only measure that hands the verdict to the person on the other side instead of to the dashboard. Workslop does not survive that question. A polished reply can fool someone skimming it, but it cannot fool the customer who is still stuck, and when you ask them plainly whether they got what they needed, they will tell you. We care about it enough that the surveying is built into the platform rather than left as someone else's reporting problem, because a company confident in its accuracy should want to be graded by the people it serves. Re-contact rate is a useful companion, the share of people who come back within a few days because they weren't really helped, and it catches the same failure from the other side. But CSAT is the benchmark, and we would rather be judged on it than on any number designed to flatter us.
Looking forward, our guess is that workslop is going to become the dividing line in AI customer service, separating the tools that genuinely help from the ones that just generate convincing noise, and customers are going to get better and better at telling the difference. The novelty of a fluent bot is wearing off fast, and what's left when the novelty goes is the plain question of whether the thing actually solved your problem. The businesses leaning on polish are going to find that polish stops impressing anyone, while the ones obsessed with being correct, even at the cost of sounding a little less slick, are going to quietly earn the trust that's getting scarcer.
We hold this view with humility about the difficulty, because keeping AI honest is genuinely hard and we work at it every day. But if there's one thing we'd want a business owner to take away, it's to stop being impressed by how good the replies sound and start checking whether they're actually true and actually helpful, because your customers are already doing exactly that. Workslop thrives on the gap between sounding right and being right, and the only real defense is to care more about the second one than the first, even when the first is the one that demos well.



