When is an AI assistant’s work actually done?
OpenAI has published GDPval as a way to assess knowledge-work tasks. That is useful, but a task score has a boundary: it tells us something about the work selected for the test, not whether a real person’s obligation is complete. A draft is not done if approval is missing, a policy source is stale, or the next owner has to reopen the trail to discover what changed. If an AI assistant calls a task complete, I want the label to mean something concrete: delivered or still draft; who owns the next decision; which source and date support the claim; and what has not been checked. Otherwise a polished output can quietly become somebody else’s 5 p.m. cleanup.
Comments
My cheap test: give it one ordinary task—turn a Friday call into a follow-up email, then put the draft on the right thread for review. Walk away for an hour. If the person still has to hunt for it, copy it into the mail app, or ask whether it went out, the helper only finished the writing part. The saved hour is in getting the next person to the next click without a scavenger hunt.
Yes. And "put it on the right thread" still needs to be distinguishable from "sent." A system may have created the email, attached it to the conversation, and queued it for review. Those are different states with different next actions. A completion label that hides them does not remove the risk; it leaves the recipient guessing what happened.
Those states should not all accrue as "completed" in the vendor dashboard. Drafted, queued for review and sent have different value to the person paying for this, but collapsing them makes the product look busiest exactly where the human still has work. Count the follow-up as part of the task or stop calling it done.
I’d make the status change the surface, not leave a green check in a history panel. If the email is queued for review, open the card on the draft, the recipients, and three plain choices: send, edit, or leave unsent. Nobody should have to decode “complete” and then hunt for the thing that still needs them.
“Leave unsent” still has an owner problem. For a customer follow-up, it is a decision to stop, not a harmless holding state. Give that draft an owner and a date; otherwise the cleanest-looking status is how a promise disappears.
Exactly. A draft that was deliberately parked needs to look different from one nobody has opened yet. On the customer thread, I would show “Not sending — Maya decided Aug. 29” beside the last message, with a reminder date only if it is meant to come back. Otherwise a quiet card reads like a forgotten one.
Then the completion rate needs two columns: system complete and obligation complete. Sample a week of items the assistant marked done; count how many still needed a send, approval, clarification, or manual record change before the next person could move. If the second number lags, calling the first one complete is teaching everyone to distrust the label.