What should an AI model maker show before calling a model 'aligned'?
OpenAI called GPT-6 Astra its ‘most aligned model yet’ when it announced the model last week. That phrase is doing more work than the launch. A buyer cannot use it to decide whether the model will stay inside a task boundary, flag uncertainty, or make a bad situation worse. I would rather see a small, ugly scorecard: cases it refused correctly, cases it over-refused, what changed after outside testing, and where the company still tells people not to use it. If ‘aligned’ cannot survive those questions, it is a compliment, not product information. What would you need before that label changed a real decision?
Comments
“Aligned” is too far from the moment a normal team needs help. Put the limits beside the task: “Use this to draft the customer update; do not use it to decide whose account gets closed.” And when it stops, say why in ordinary words. People should not have to read a scorecard to learn the boundary after the damage is done.
I’d want the claim translated into a task-level error table, not a model-wide badge: among ordinary customer-update drafts, how often did it invent a fact, omit a required caveat, or refuse something allowed? Put that beside correction minutes and sample size. A lower incident rate matters if the team can spend less time rereading routine work without discovering problems later.