Why most AI agent ROI calculations fail?
The standard business case claims an agent will deflect 10,000 tickets per year at $25 per ticket, totaling $250,000 in savings. This number gets approved, then forgotten when the agent goes live because deflection rates are rarely measured properly. Tickets the agent appears to handle often resurface as escalated calls, second-touch interactions, or frustrated users who abandoned the system. Measure outcomes that survive scrutiny, not vanity metrics that look good in a slide.
A real example: one agent was reported as deflecting 70% of password reset requests when actual resolution rates, measured with follow-up tracking, were closer to 22%. The 70% number came from conversation completion rates, not actual resolution rates. The CFO noticed the gap.
What does a custom AI agent actually cost?
Kernel Flow builds production-grade AI agents that integrate directly into your existing databases and software. The cost has two components: platform licensing and delivery and tuning.
Platform costs depend on message volume and agent complexity. A simple FAQ agent consumes fewer message credits than one performing real-time data lookups or workflow automation. Plan for production message consumption to run 2-4 times higher than pilot phase consumption once real user volume hits the system. Generative has multiply per-conversation cost significantly.
Simple FAQ agent: Three to four topics with basic routing typically costs $25,000 to $50,000 AUD in delivery.
Service desk agent: Ticket creation and routing integration costs $60,000 to $130,000 AUD in delivery.
Multi-skill agent: Power Automate workflows and data integration runs $130,000 to $280,000 AUD in delivery.
Complex enterprise agent: Custom plugins and enterprise data sources range from $250,000 to $500,000+ AUD in delivery.
Ongoing tuning and improvement is typically one to three days per month for the first six months. Budget $40,000 to $80,000 AUD in year one for ongoing work, declining in year two as the agent matures. Agents that receive no tuning degrade over time as knowledge sources become stale and user behavior shifts.
Which metrics actually prove ROI to finance?
Containment rate is the percentage of conversations that resolved without human escalation. It's the headline metric, but only if measured properly. Conversation completion is not the same as user resolution.
Measure containment correctly by tracking conversation outcomes for 14 days after each user interaction. Check whether the user raised a related ticket, called support, or had a second conversation on the same topic. True containment is conversations that resolved and did not trigger a follow-up. This requires more effort than counting conversation endings but produces a number finance teams will believe.
Realistic containment in month one: Service desk agents typically achieve 30% to 55% proper containment in the first six months.
Containment after active tuning: With ongoing optimization, containment climbs to 55% to 70% by month twelve.
Escalation handle time reduction: Conversations that escalate to humans should resolve faster because the agent collected relevant information and routed correctly, creating measurable productivity gains.
How to build a business case that survives scrutiny?
Start with conservative containment assumptions and real delivery costs. Calculate savings based on properly measured deflection rates, not pilot-phase optimism. Factor in year-one tuning costs and platform consumption beyond the base allocation.
Include escalation time savings and quality improvements as secondary benefits. A agent that routes complex cases correctly reduces resolution time for support staff, creating real operational use without assuming unrealistic deflection numbers.
Track metrics continuously from launch through month twelve. Adjust tuning based on actual performance. Present quarterly results to stakeholders to demonstrate progress and justify ongoing investment.
