Prentis is training AI on the work people do between systems
Prentis describes itself as an AI research lab focused on computer-use models. Its models are meant to perceive screens and operate software across desktop, browser, and mobile interfaces. The company says the missing ingredient is not simply a smarter model. It is exposure to real work: the exceptions, handoffs, and recovery steps that rarely make it into a process document.
That is a credible target. An insurance claim does not live in one clean form. A person checks an email, compares a policy, opens an old account system, notices a name mismatch, asks a colleague, and records why the ordinary rule did not apply. Customs refunds, healthcare administration, and manufacturing paperwork have similar seams. Those seams are where computer-use AI agents either save time or create a second job for the reviewer.
TechCrunch says Prentis plans tailored agents for jobs such as insurance claims and customs-duty refund exceptions. It also reports that the company claims its 32-billion-parameter Hive-32B model beats larger rivals on WindowsAgentArena and ScreenSpot-v2 at roughly one-tenth the task cost. The publication could not independently verify those claims, and Prentis has not published enough detail for an outside reproduction.
The benchmark names still matter. WindowsAgentArena tests agents on more than 150 tasks inside a real Windows environment, including documents, spreadsheets, browsers, settings, file management, and other ordinary applications. It rewards completed tasks, not a convincing explanation of what the agent meant to do. That is closer to office work than a chatbot leaderboard. It is not the buyer's office.
Twenty percent of savings can align incentives—or distort them
Performance pricing fixes one common software problem: the vendor gets paid for licenses while the customer is left to discover whether anything improved. If the fee depends on a finished claim, recovered refund, or retired manual step, the vendor has a reason to stay through implementation instead of celebrating activation.
The trouble begins when the denominator moves. Suppose a claims team costs $1 million a year and an AI agent lets the company run the same volume with fewer overtime shifts. Is the saving the overtime removed, the projected hires avoided, or part of the whole team's payroll? If volume drops for unrelated reasons, does the agent get credit? If employees spend four hours a week checking its work, where does that cost appear?
A percentage fee also gives both sides reasons to favor flattering definitions. The vendor wants a large baseline and a visible reduction. The buyer may want to exclude internal review, integration work, or mistakes because admitting those costs weakens the business case it already sold upstairs. Nobody needs to lie. They only need to put inconvenient work in a different column.
Labor makes the accounting sharper. A company can count fewer paid hours as savings while a worker experiences a smaller paycheck, a denser queue, or more exception-heavy work. Those may still be rational changes. They are not the same outcome. A credible savings statement should say who kept the gain and who absorbed the change.
A benchmark win does not settle the invoice
Computer-use benchmarks answer a real question: can the model locate controls, move across applications, and finish a defined task? They do not settle what completion means in a live business process.
An agent can enter a claim correctly and still choose the wrong claim. It can recover a customs refund and consume so much employee checking time that the net saving vanishes. It can finish 90 routine cases while quietly passing the ugliest ten to people with less context and less time. A task score sees the click path. A business has to price the aftermath.
Before agreeing to a share-of-savings contract, run the agent on a fixed sample from your own week. Include clean cases, incomplete records, conflicting instructions, locked accounts, policy exceptions, and one system outage. Keep the old process running beside it long enough to measure both.
Count finished cases that stayed finished. Subtract integration hours, review time, corrections, repeat handling, support tickets, retraining, downtime, and work pushed to another team. Keep safety and customer harm outside the bargaining range; a wrong payment or missed medical exception should not become acceptable because the average case got cheaper.
Priya wants a net number. Ivy sees a useful forcing function.
Priya Rao would not reject performance pricing. She would reject a gross number. Her version starts with the same case mix, volume, and service level before and after. Then it subtracts setup, review, correction, reruns, reopened cases, and any work shifted to another person. If those costs are unavailable, the savings claim is not ready for a percentage sign.
Ivy Chen is more open to the structure because it can force a vendor to name the chore it plans to remove. A small team may prefer paying from a verified result over funding another long software rollout. Her condition is blunt: the contract should identify the old task, the owner who confirms it is gone, the quality floor, and the point where the pilot stops. Paying for savings is useful only if the buyer can say no when the work merely changes shape.
Both positions beat arguing about whether AI agents are generally good for office work. Some repeated computer chores are overdue for removal. The practical fight is over the baseline, the after-state, and the human hours that refuse to fit in a demo.
How to measure AI agent savings without fooling yourself
Freeze the baseline before the pilot. Use several ordinary weeks, not the worst month the team has ever had. Record volume, completion time, error rate, repeat work, overtime, service level, and the people involved. Write down seasonal or staffing changes that could move the result on their own.
Define one accepted outcome. For an insurance claim, that might mean the record was complete, the decision followed the current policy, the customer was notified, and the case did not reopen within 30 days. A click sequence is not an outcome. A draft waiting for a human is not a finished case.
Price the checking. Track reviewer minutes even when the review finds nothing. Track corrections, second passes, escalations, system-owner help, and the time employees spend reconstructing what the agent did. Hidden review is how automation moves from the workday into the evening without appearing on the invoice.
Separate money saved from capacity created. Avoiding a future hire, cutting overtime, serving more customers, and reducing employee strain are different gains. Name which one occurred. Do not convert all four into the same cash figure unless cash actually moved.
Finally, keep a short list of outcomes the fee cannot reward: fewer escalations caused by blocking valid cases, faster handling caused by skipping checks, or lower labor cost created by pushing unpaid cleanup onto salaried staff. Incentives work. That is why the definition matters before the agent opens the first spreadsheet.
Prentis may prove that specialized, cheaper models can automate a large share of office computer work. Its reported pricing model is already useful because it drags the value question out of the demo. The answer should survive contact with payroll, customer complaints, reviewer time, and the person still at the desk after five.