AI Silent Failures & Conversational Loss Patterns Killing Your Product Adoption

Which failures to look for, why your dashboards miss them, and what to fix — grounded in Stanford research on 100,000 real AI conversations

B
Brittany Hobbs · · 21 min read
AI infrastructure sinking and failing 0 image demonstrated the silent failure of an image generation app
Overview
  • 79% of AI failures are invisible — they never fire an alert, never dent CSAT, and never show up in the telemetry your engineering team watches.
  • The damage lives in the behavioral layer, not the model layer. Even as newer models cut overall failure rates from ~42% to under 10%, the failures that remain stay overwhelmingly invisible and the mix of failure types barely moves.
  • Seven archetypes account for 99% of these silent failures. This guide gives you the signature of each and the design fix.
  • Your most skilled users see the most failures — and recover from them. That paradox is the key to instrumenting your product.
  • The single highest-leverage move is not a new eval. It is reading real transcripts with people who hold different job titles.

Which failures to look for, why your dashboards miss them, and what to fix — grounded in Stanford research on 100,000 real AI conversations.


Most of the ways your AI product fails will never reach your dashboard. In a study of 100,000 real conversations between people and live AI systems, Christopher Potts and Moritz Sudhof of the Stanford NLP group found that 79% of AI failures are invisible: the AI got something wrong, and the user gave no sign of it — no correction, no complaint, no dip in any metric you review on Monday. The session closed clean, and on paper it looked like a success.

The loud failures are the ones that make headlines. Deloitte Australia refunded a government department after a $440,000 report shipped with fabricated citations. Latham & Watkins filed a court declaration with hallucinated citation details in a copyright case.

Here you'll learn the seven failure archetypes, the signature of each, why your current tooling can't see them, and the fix for each. At PH1 Research we do this work with product teams, and the pattern holds: the failures that damage a product most are the ones no one is measuring. Save this, and pull it up the next time you audit a conversational experience.


What Are AI Silent Failures & LLM Loss Patterns

Every conversational AI product you have ever shipped is failing constantly, in conversations happening right now. Not the way a server outage fails. Nothing crashes. Nothing throws an error. These are failures only the user in the room can catch, and most of them never say a word.

Here is the shape of it. You ask an AI for a great sushi restaurant nearby, and it recommends a well-reviewed dim sum spot instead. Close enough to sound helpful, wrong enough to miss what you asked for. Or you request a report, and get one back that reads as thorough and confident, but is quietly missing a requirement you gave it up front. Nothing broke. The output looks fine. It just is not what you asked for, and unless you catch it yourself, no one will.

Different teams have arrived at this same problem from different directions, which is part of why it stays invisible inside product organisations: everyone has given it a different name. On our episode with Moritz Sudhof, CEO of Bigspin, he calls them silent failures, the term this guide uses throughout. On our episode with the UX research team at Microsoft AI, they call the same phenomenon loss patterns. Different vocabulary, same underlying problem: a conversation that looks successful on paper while quietly failing the person having it.


Why four in five failures never reach your dashboard

These failures stay hidden because they happen at a different layer than the one your tools watch. Telemetry tracks the technical layer: latency, uptime, completion, session length. The failures live in the behavioral layer — the decisions the model makes about how to act while it answers, and how the user reacts to them.

Model improvements have made real progress on the technical layer. When Potts and Sudhof re-ran old user prompts through today's frontier models, overall failure rates fell from around 42% to under 10%. But the failures that survive are still overwhelmingly invisible, the mix of failure types barely changes, and some behavioral problems got worse: the newer models over-deliver and pad their answers more than their predecessors did.

The clearest behavioral failure is what Moritz calls generating rather than clarifying. A user sends an ambiguous request. A person would often slow down and ask which requirement matters more. The model, trained to be helpful, treats helpful as producing a long, thorough answer immediately, so it guesses and moves on. Whether the guess is right has little to do with how smart the model is. It is a choice about how to behave, and that choice is what your dashboard cannot see.

So everything technical stays green. Latency is fine, and a 20-turn session reads as engagement even when it is a user fighting the model, though the dashboard cannot tell the difference. Only about 3% of users leave explicit feedback. You are running a conversational product on instruments built for a different machine, and the rest of this guide is about seeing what they miss.


The terms everyone building AI products should know

A handful of terms get used interchangeably for problems that are not the same. Anyone building or evaluating a conversational product should be able to tell them apart, because each one points you at a different fix.

  • Invisible failures — the term used here, from Potts and Sudhof — and the CX industry's silent failures name the same thing: the AI got something wrong and the user gave no signal. The seven patterns in this guide are their failure archetypes.
  • Failure modes is the broadest label, but the agent-reliability literature has mostly claimed it: system-level taxonomies of tool-call errors, reasoning drift, and multi-agent coordination breakdowns. That literature is legitimate and mostly technical; it answers "why did the system break," not "why did the user leave without what they needed."
  • Loss patterns, or productivity loss, is Microsoft Research's frame, from Ironies of Generative AI. It measures something different: the time and value a user loses even when the AI technically succeeds — the shift from producing work to reviewing it, the interruptions, the way automation makes easy tasks easier and hard tasks harder. A conversation can pass every check in this guide and still be a productivity loss in that sense.
  • Breakdowns and repair is the older dialogue-systems vocabulary — the study of conversations that derail and the moves users make to get them back on track.

The difference that matters for a product team is the vantage point. Failure modes and loss patterns are measured from the system or the org side. Invisible failures are measured from inside the conversation, where the user is. Most teams are missing that view, and it is the one this guide takes.


The field guide: seven silent failures

The Stanford study sorted invisible failures into eight archetypes. Seven of them account for more than 99% of what happens; the eighth, the mystery failure, is a sub-1% catch-all for the rare cases that fit none of the others (it appeared on just 10 transcripts, which the authors read as a sign the other seven are close to comprehensive). Those seven are what you should be able to recognize on sight. For each one below, I have given you the definition, the transcript signature to search for, the reason your current tooling stays quiet, and the design move that addresses it.

A note before you use these: the archetypes are not independent. A confidence trap can trigger a death spiral, which can end in a walkaway. Treat them as a diagnostic vocabulary, not seven separate bugs.

1. The Confidence Trap

What it is. The AI asserts something wrong, or not quite what the user asked for, and delivers it with such polish that the user is unlikely to catch it. It shows up in about a third of all failure transcripts (32%), and the researchers single it out as the most insidious archetype precisely because the interaction looks so much like a success.

The signature. Precise numbers where precision isn't earned — you see "2.68%" and it reads as rigorous because it looks specific. Clean walls of formatting and authoritative-looking citations do the same work. What gives it away is not that the answer is wrong but that it is dressed up to be believed. As Moritz put it, the more specific and helpful it looks, the more fake it often is.

Why telemetry misses it. The interaction completes successfully. The user, especially a newer user reaching to AI for something outside their own expertise, has no way to verify the claim and no instinct to try. The Deloitte and Latham & Watkins cases were confidence traps that happened to get caught by an outside expert. Most never do.

The fix. Stop putting the burden of catching this on the user. Instrument your responses for confidence signatures — specific figures, definitive claims, citations — and tune the model to surface its own caveats: what it could not find, where information may be stale, what it is inferring rather than knowing. In a knowledge-base or search product, honesty about uncertainty is the feature, not a bug that slows people down.

2. The Silent Mismatch

What it is. The AI answers a slightly different question than the one the user asked, without ever flagging the gap. It quietly reinterprets the request into something it can satisfy, and the response is plausible enough that neither side notices the swap. It is the second most common archetype in the study, present in more than half of all failure transcripts (53%).

The signature. My favorite example from the data: a user asks for 100 vegetables starting with the letter G. Around number 25, the model runs out of real ones, so it starts putting the word "green" in front of things — green hot dog, green oxtail, green pizza. It technically produced 100 items. It answered a question the user never asked. Watch for responses that satisfy the shape of the request (right length, right format) while missing what the user asked for.

Why telemetry misses it. Every automatic check passes — it returned 100 items, and the response was well-formed. The mismatch lives in the difference between the literal output and the real requirement, and that difference is often ambiguous even to a careful human reviewer.

The fix. You cannot catch this with single-turn evals that check format or keyword presence. You need a system that reads the whole interaction and asks what the user wanted versus what they got. And you need to design the model to name its own compromises out loud: "I could only find 24 vegetables starting with G — do you want me to broaden the criteria?" A surfaced mismatch is a recoverable one.

3. The Drift

What it is. Over a long conversation, the AI loses track of what it is supposed to be doing. The goal shifts, and the model stops holding the thread. Sometimes the user leads the drift on purpose, which is one of the pleasures of working with AI. Often the model drifts on its own, because long conversations erode the grounding that humans hold intuitively. Roughly 4% of failure transcripts carry it, and because it needs several turns to develop, it concentrates in longer conversations.

The signature. The example that stuck with me: a user was translating finance slides from Chinese to English. Around turn 15, they asked a side question about a career choice. That turned the AI into a life coach — but it still thought it was a translator, so every subsequent answer came with career advice and a translation of that advice. Watch for responses that carry vestigial instructions from earlier in the conversation, or that answer a question the user has clearly moved on from.

Why telemetry misses it. A long, active session reads as engagement, and the model responds fluently to every turn without erroring. The conversation simply stops arriving anywhere, and "did not reach a resolution" is not a field in your logs.

The fix. Drift is a context and memory problem that gets worse as conversations lengthen. Design for it: periodic re-grounding ("just to confirm, we're now focused on X"), clear affordances for the user to reset scope, and monitoring that flags when a conversation's topic has moved but the model's framing hasn't caught up.

4. The Death Spiral

What it is. The user has a clear goal, keeps trying to get the AI to deliver it, and keeps failing. Unlike drift, the user knows what they want. They just cannot get there, and they try the same door over and over. It is rarer than the headline archetypes — around 2% of failure transcripts — and brutal for the user who lands in one.

The signature. You see it constantly in coding tasks: "it's still not working," "it still has the bug," "that didn't fix it." Repeated requests with escalating frustration and shrinking politeness. The user rephrases, re-explains, and burns tokens without progress. Expert users escape by changing strategy. Most users just keep pushing until they give up.

Why telemetry misses it. High turn count and long session time look like your best-case engagement metrics. A death spiral and a delighted power session produce nearly identical numbers. Without reading the content, you cannot separate the user who got 20 things done from the user who failed at one thing 20 times.

The fix. Detect the pattern — repeated near-identical requests, rising frustration markers, no forward movement — and intervene. Offer a different path, escalate to a human, or have the model itself step back and propose a new strategy instead of retrying the failed one. The goal is to shorten the distance between "this isn't working" and a real change in approach.

5. The Contradiction Unravel

What it is. The AI confidently contradicts itself across turns. It states one thing early, something incompatible later, and never acknowledges the reversal. It is one of the rarest archetypes on its own — 0.3% of failure transcripts — but the study found it travels closely with the confidence trap, and the pairing is what marks the most damaging cases: a confident answer and its later reversal both sliding past the user unnoticed.

The signature. An answer on turn 12 that quietly conflicts with the answer on turn 4. Numbers that shift between turns without explanation. Recommendations that reverse. The user often doesn't scroll back to check, so the contradiction sits there unresolved while the conversation moves on.

Why telemetry misses it. Each individual turn is internally coherent and passes any single-response check. The failure only exists across the arc of the conversation, which is the view most eval systems never take.

The fix. This is the strongest argument in the whole guide for monitoring at the interaction level rather than the turn level. You need a system that reads the full conversation and checks for internal consistency across turns — because no single-response eval will ever catch a contradiction that spans ten of them.

6. The Walkaway

What it is. The user gives up. They abandon the task, or quietly shrink what they were trying to do because they have realized they can't get the whole thing, and they leave with less than they came for — no error, no complaint. It is the single most common archetype in the entire study, appearing in 85% of failure transcripts, because so many failures simply end with the user walking off.

The signature. A conversation that ends mid-task. A user who started asking for a complete deliverable and settled for a fragment. Scope that contracts over the session — from "write the full report" to "just give me the outline" — not because the user changed their mind but because they gave up on more.

Why telemetry misses it. A completed session and an abandoned one can look identical in the logs. Worse, a walkaway with reduced scope can register as a success, because the user did get the smaller thing they downgraded to. You never learn what they wanted and didn't get.

The fix. Instrument for scope contraction and abandonment, not just completion. When a user reduces what they're asking for, treat it as a signal worth investigating rather than a resolved ticket. And design earlier interventions so users hit an off-ramp — a clarifying question, a human handoff — before they reach the point of quietly leaving.

7. The Partial Recovery

What it is. Partial recovery is the hopeful archetype, and it rewards study more than any of the others. The user hits a failure — a contradiction, a drift, a bad answer — and instead of spiraling or walking away, they steer the AI toward something that mostly works. It appears in about 5% of failure transcripts, and the follow-up fluency research shows it concentrated among the most skilled users.

The signature. A user who catches the model's mistake and redirects: "that's not right, you contradicted what you said earlier — let's go back to X." Added constraints. A reframed request after a failed one. The conversation bends back toward success instead of breaking.

Why it matters more than telemetry. Partial recovery is the one archetype you want to propagate rather than eliminate. These users are doing, in real time, the quality control your product should eventually do for everyone. Reading their transcripts shows you which interventions turn a failing interaction around.

The fix. Study your recoveries as closely as your failures. The moves your best users make to rescue a conversation — pushing back, adding constraints, asking for a confidence check — are a specification for the affordances and behaviors you should build into the product itself. Recovery is a design pattern hiding in your transcripts.


The seven silent failures at a glance

Save this table. It is the fastest version of everything above — the one to have open while you read a transcript.

Failure What it looks like The tell to search for The fix
Confidence Trap Wrong or off-target answer, delivered with polish Precise numbers, clean formatting, authoritative citations where certainty isn't earned Instrument for confidence signatures; tune the model to surface its own caveats
Silent Mismatch Answers a slightly different question than was asked Response fits the shape of the request but misses the intent ("green hot dog") Check whole-interaction intent, not format; have the model name its compromises
Drift Model loses the thread over a long conversation Vestigial instructions from earlier turns; answers to abandoned questions Periodic re-grounding; scope-reset affordances; topic-shift monitoring
Death Spiral User keeps trying the same thing and failing "Still not working," repeated near-identical requests, rising frustration Detect the loop; intervene with a new strategy, path, or human handoff
Contradiction Unravel Model confidently contradicts itself across turns A later answer conflicts with an earlier one, unacknowledged Interaction-level consistency checks, never single-turn
Walkaway User abandons or quietly shrinks the task Session ends mid-task; scope contracts from "full report" to "outline" Instrument for scope contraction and abandonment; build earlier off-ramps
Partial Recovery User catches a failure and steers back to success Pushback, added constraints, reframed requests that work Study these; turn the recovery moves into product affordances

The fluency paradox: your best users see the most failures

Moritz and Potts ran a follow-up study, A Paradox of AI Fluency, on 27,000 conversations, this time looking at the human side of the interaction. A small minority of users — around 2 to 3% in the data — behave completely differently from everyone else. The paper calls their stance augmentative: they iterate with the AI, refine their goals mid-conversation, push back, ask for confidence assessments, and critically assess what they get back. In the data, 93% of high-fluency interactions are augmentative, against under 1% for the least fluent users.

Now the part that breaks most dashboards: these expert users encounter more failures than novices, not fewer. 64% of their conversations contain at least one failure, against 24% for the lowest-fluency group. If you ranked users by how many problems they hit, your most sophisticated people would look like your unhappiest. What saves them is that 59% of their failures are visible, because they catch the model and steer it, against just 12% for novices. The same underlying failures get surfaced in one group and buried in the other. Experts also take on harder tasks and succeed at them more often.

The majority take the opposite, delegative stance. They treat the AI like a vending machine: put in a prompt, take out a result, leave. If they don't like the result, they say "try again" or walk away. A delegative user has no habit of verifying and no instinct to push back, so a confident, well-formatted, wrong answer gets accepted whole. The users with the highest share of invisible failures are the ones least equipped to make those failures visible.

The reframe I keep coming back to with teams at PH1 Research and AI Value Acceleration: every AI product already has a human doing quality control, and it is the user. In how they steer, push back, and abandon, they are telling you what the model does well and where it misses. Your expert users are teaching the model to succeed at tasks it would fail on its own. The open question is whether any of that signal survives past the end of the session, or evaporates every time.


Designing the user is half the job

As product builders, we are not only designing how the model behaves. We are also designing how the user behaves, and we tend to forget the second half. The fluency paper reaches the same conclusion: we design user behavior as much as model behavior, and encouraging deep engagement rather than friction-free experiences is what produces more success overall.

He proved it with an experiment I have not stopped thinking about. At BetterUp, his team ran a role-play coaching product — users rehearsing a difficult conversation, like giving a boss bad news. Same model, same prompt, same system, pixel-identical UX. The only thing they changed was the introduction. In one version, the AI was framed as a supportive, empathetic mentor. In the other, it was framed as a machine: non-judgmental, always available, never tired, with access to a lot of information.

The outcomes diverged sharply. Users in the machine-framed group reported being twice as prepared for their real conversation afterward. They were less worried, sometimes even excited, and their repeat engagement was meaningfully higher. Nothing about the model changed. What changed was how users thought about the thing they were talking to, which changed how honest and raw they were willing to be, which changed what the model could give back.

That loop is where the design happens: how the user thinks about the AI shapes what they say to it, which shapes what it can return, which shapes what they do next. You set the mental model with framing and shape behavior with affordances, and the right settings depend on your product. A coaching tool wants the model asking Socratic, clarifying questions instead of dumping a bulleted answer. A knowledge-base search tool wants the opposite, with fast answers and low friction, but paired with honesty about what the model couldn't find or isn't sure of. In both cases you are tuning the interaction to protect the lower-fluency user who would otherwise accept a confident answer and move on.


Why coding works and everything else is harder

There is a reason AI feels so much more reliable for engineers, and it is not that engineers are smarter users. Their work has a verification layer built in. The AI can run the code, run the tests, check whether the linter complains, and self-correct before a human ever sees the output. It can loop on a task for a long time and come back with something closer to complete, because it has a definition of "good" it can check itself against.

The failure data backs this up from the other direction. Software development shows the highest rate of visible failure of any domain in the study, because the users are experts who spot mistakes and talk back to the model. The domains that should worry you are the quiet ones — creative writing, general knowledge, everyday advice — where failures go uncorrected because the user can't see them.

Now think about writing an email, drafting a strategy memo, synthesizing research, or working through a design problem. There is no test suite for tone. There is no linter for "this is what the client needs." The model cannot self-verify, so it cannot safely run loose, and the burden of judging quality falls back on a human who often does not know their own full requirements until they see the work.

That last part is the trap in the current enthusiasm for delegation. The dream is to write one detailed spec, send the AI off for twelve hours, and review a finished deliverable. For anything requiring judgment, creativity, or expertise, that model fights how the work gets made. You don't know all your requirements up front. You discover them by reacting to a draft — the tone should be warmer, you forgot a constraint another stakeholder mentioned, the approach reads worse in practice than it did in the spec. The longer the AI runs unattended, the more thinking happens out of your view, and the more likely you are to review a polished output that quietly misses the mark and sends you back to the start.

The way through is lower-friction iteration, not less AI. Build products, and build your own habits, so you can hill-climb toward the deliverable in tight loops rather than gambling on one long unattended run. The reviewer layer comes in here: virtual reviewers that encode what "good" looks like for a given stakeholder, so judgment-heavy work gets the self-verification loop that code already has. That layer does not exist yet, but its raw material — the record of what your users want — is already sitting in your transcripts.


The one-transcript audit: what to do Monday

If you are a researcher, PM, or designer who suspects your AI product is underperforming and can't get your technical team to see it, this is the most useful thing I can hand you. It costs nothing and it is where every serious improvement Moritz has seen started.

Read real transcripts. Not summaries, not aggregate metrics, but actual sessions between real users and your AI, or your own dogfooding sessions. Reading transcripts is what separates teams who understand their product from teams who guess. Most of the seven failures above are invisible in a dashboard and obvious in a transcript once you know what to look for.

Read one transcript together, with mixed job titles in the room. Put a researcher, a PM, a designer, and an engineer around one conversation. Everyone will notice something different, and no single person has the whole picture: a trust break, a missed intent, a moment the user needed a clarifying question. "Good" for an AI product cannot be defined by technical metrics alone, and the people who understand the user and the business usually see the most actionable failures.

Translate what you see into behavioral requirements. Behavioral requirements are what give you leverage with engineering. Don't report "the AI feels off." Report the decision the model made: "On turn 6, the request was ambiguous and the model generated a full answer instead of asking which constraint mattered, which led to the user re-prompting four times." An engineer can turn that into a prompt change or an eval case. You are handing them a specification, in their language, grounded in real data.

Then close the loop. Your expert users are already revealing preferences, requirements, and recovery moves in every session. Mine your past transcripts for the patterns and encode them as prompt instructions, eval cases, and skill files that make the behavior durable instead of ephemeral. Intelligence is commoditized; anyone can make the API call. The proprietary asset is the feedback loop from your own users about how to make a generic model produce work that satisfies them. You build that moat one transcript at a time.

Evals are necessary and not sufficient. They catch only what you already knew to look for, usually on single turns in a walled set of test scenarios. The failures in this guide emerge across the arc of an interaction, and in the behavior of users who look nothing like the people who wrote the evals. Keep the evals, and add the transcripts.


The research behind this guide

Everything here traces to Moritz Sudhof and Christopher Potts's published work, and the data and code are open if you want to go deeper or run the analysis on your own transcripts.


What to do this week

  • Researchers: Pull five real transcripts. Tag each turn for the seven archetypes. Bring the three clearest examples to your next product review as behavioral requirements, not vibes.
  • PMs: Add two fields to how you evaluate conversations — scope contraction and unresolved endings. A completed session is not a successful one.
  • Designers: Audit your product's framing and first-run experience. What mental model are you setting before the user says a word? Run the BetterUp test: change only the introduction and watch what changes downstream.
  • Anyone: Book one hour, one transcript, four job titles, one room. It is the highest-return meeting you will run this quarter.

The uncomfortable truth in Moritz's research is that we have handed the definition of quality to the layer of the stack least equipped to hold it. Model benchmarks cannot tell you whether your users are getting what they came for. The people closest to those users can, and right now most of them have never read a transcript. The teams that win the next eighteen months will be the ones who fix that first — not by buying a smarter model, but by finally looking at what happens in the room.


Go deeper

This guide compresses two hour-long conversations into a field reference. If you want the full reasoning, the disagreements, and the details that did not fit here, listen to the full episodes: Moritz Sudhof on silent failures and the Microsoft AI UX research team on loss patterns.

How helpful was this article?

Have a story to share?

0 / 500

Latest Episodes

All episodes

Product Impact Newsletter

AI product strategy delivered weekly. Free.