AI agents in business: What a correctly completed transaction really costs
What does an AI agent really cost? Why businesses need to account for rework, escalations and processing time, not just token prices.
Consider an AI agent in a business that checks order confirmations. It reads the document, compares products and payment terms with master data and prepares the next steps. The run looks cheap in the model provider's dashboard.
Then an order confirmation arrives with two cash discount tiers. There is no clear output format for this case. The transaction goes to an employee, who opens the document, compares the details and requests a correction. A short model call has turned into several minutes of specialist work.
What still costs money after the model call
Many AI cost calculations leave out this time. The cost report includes tokens, API calls and hosting, although the transaction is still open in the business team. More costs arise there: specialist time, approvals, questions, delays and, in the worst case, a wrong decision. A commercial assessment has to follow the transaction through to a checked result.
A cheap model can therefore cost more in day-to-day use. Each call may cost less but produce incomplete data more often. A more capable model can cost more per request and still be the better choice if it completes more transactions without correction. The process reveals which option is more economical.
What belongs in the performance record
I therefore want to see a performance record for every digital employee. It should show the number of cases processed, the completion rate, average processing time, rework, escalations and the full cost per transaction. Human approvals belong in it as well. They may be planned by design or point to a weakness in the process.
These figures help with decisions that model prices alone cannot answer. Is a stronger model worth it? Does a special case need a fixed escalation? Does routing create more work than it saves? Can the digital employee take on more cases without increasing the amount of oversight?
Two figures instead of one token price
The next evaluation of an AI agent should therefore put two figures side by side: how many transactions did it complete correctly, and what did each one cost up to that point? Token usage belongs underneath as a technical detail.