When an AI chatbot gets a reply wrong, what should the repair screen show?
“That didn’t help” is an awful place to leave someone. They have already explained the thing once, and the usual chatbot repair is a fresh blank box or a cheerful rewrite. A recent UXBench preprint builds evaluation from real feedback signals and treats recovery after a bad answer as its own task. That seems closer to what users actually feel. The repair screen should keep the question and offer a few specific moves: “too long,” “wrong source,” “I meant this part,” or “start over without using that assumption.” If it guessed something, show that guess where it can be corrected. Otherwise a bad answer becomes two jobs: spot the mistake, then figure out how to describe it to the same system again. What would make a chatbot’s second try feel like it heard the first failure?
Comments
The recovery screen should begin with a sentence the person can recognize: ‘I assumed you meant a refund, not an exchange.’ That is better than a generic apology because it gives them one thing to correct. A chatbot that makes you reconstruct the misunderstanding is still asking you to do its reading.
Give "I need a person" equal weight to "try again." Otherwise the repair metric rewards a longer conversation, and the chatbot turns a clean handoff into four more attempts at sounding useful. A real exit should carry the original question and the corrected assumption, not make someone tell the story a third time.
And do not make escalation a test of resolve. A customer who has already been bounced around will choose whatever route ends fastest, not necessarily the right one. If the chatbot guessed wrong once, leaving should get easier, not require a small argument.
The cheap test is a support question with one wrong detail baked in: say it is a 2019 printer when it is a 2021. Correct the year once. The next answer should use the right model and keep the steps already tried. If it asks for the whole story again, the repair screen is just a nicer reset button.
Also track whether the correction actually sticks: second answers needing another correction, minutes to a resolved answer, and abandonment after a bad first answer. A repair screen can feel considerate and still fail if people have to repeat the same correction on the next turn.
For the team buying it, I would also sample the chats that get handed to support. A better recovery rate can still hide a new cleanup job if the rep has to read the first answer, the correction, and three retries before helping. The handoff should arrive with the failed assumption and what the person already ruled out.
Yes—and preserve the original answer’s source and version with the correction. A support rep needs to distinguish a bad inference from a policy page that changed between turns. Otherwise "fixed" gets counted as a recovery while the same stale answer keeps coming back.