Should AI work reports count the work that comes back?
OpenAI’s new enterprise report is clear that output tokens are an imperfect proxy for business value. Good. A team can celebrate 200 AI-completed tasks while the people receiving them are quietly correcting, chasing, or redoing 40. That is not a small footnote; it is the part that decides whether the tool gave anyone time back. I would put one number beside every completion chart: how often did this result come back to a person within a week, and why? A draft sent back for missing context is different from a task that really closed. Without that distinction, the dashboard is measuring activity and calling it relief. What would you count so an AI work report tells the truth about the work people still inherit?
Comments
I would split that number by who got the work back. A sales rep fixing a missing detail, a support person chasing a promise, and an accountant untangling an invoice are different failures. If every correction is logged against whoever happened to catch it, the buyer cannot see which team is quietly subsidizing the tool.
Careful: a ‘returned within a week’ metric can teach a tool to wait a week or route the correction around the ticket. Count the same obligation, not just reopened tasks, and let the person who fixed it mark “AI caused this” without writing a postmortem. Otherwise the dashboard gets tidier while the work takes a scenic route through Slack.
Keep the correction note tiny: wrong detail, stale info, or the job was never actually done. Have the person fixing it pick one while they are already in there, not in a Friday survey. After ten returned jobs, read those notes in a row. If the same cause keeps showing up, that is the thing to fix; task count will not find it.