The ROI measurement problem: why most organizations can’t tell if their AI investment is working

The ROI measurement problem: why most organizations can’t tell if their AI investment is working

 

The question most organizations can’t answer 

 

Ask a technology leader whether their AI agent deployment is delivering ROI, and you’ll get one of two responses. The first is a confident yes, backed by a number that turns out to be activity metrics dressed up as outcomes: API calls per day, interactions handled, time saved on a specific task in a controlled environment. The second is a pause, followed by an honest admission that the measurement framework hasn’t been fully defined yet. 

Both responses point to the same underlying problem. Organizations are deploying AI agents at an accelerating pace — 79% are already using them in some form, according to recent research — but the infrastructure for measuring whether those deployments are actually generating business value hasn’t kept up with the deployment pace itself. 

This isn’t a minor operational gap. It’s the reason why 60% of organizations are still primarily investing in pilots despite growing budgets, and why only 25% of AI initiatives have delivered expected ROI since 2023. If you can’t measure it, you can’t improve it. And if you can’t demonstrate it, you can’t justify scaling it. 

 

 

Why the obvious metrics don’t work 

 

The first instinct when measuring AI agent performance is to reach for operational efficiency metrics. How many tasks did the agent handle? How much faster is the process? How many hours of human work did it replace? 

These are useful numbers. They’re also insufficient, and in some cases actively misleading. 

An AI agent that handles 10,000 customer interactions per month is not necessarily delivering value if the interactions it handles are low-stakes and the ones it escalates — the complex, high-value cases — still require the same human time as before. An agent that reduces average handle time by 30% is not necessarily delivering ROI if the quality of outcomes in that shorter time has degraded in ways that generate downstream costs: callbacks, escalations, errors that surface later in the process. 

The operational efficiency frame measures what the agent is doing. It doesn’t measure what the business is getting as a result. 

Gartner’s research on AI deployment costs makes this concrete: at 25% AI agent adoption, governance costs alone can increase over 34%. An organization measuring ROI purely on task automation efficiency can show positive numbers while simultaneously accumulating governance, oversight, and error-correction costs that more than offset the efficiency gains — and never see it in their dashboard. 

 

 

The metrics that actually matter 

 

Meaningful ROI measurement for AI agents requires connecting agent activity to business outcomes at three levels. 

The first is process-level impact. Not how many tasks the agent completed, but what happened to the process as a whole. Did time-to-resolution improve across the full workflow, including human-handled steps? Did error rates at the process level go down or up? Did the volume of exceptions and escalations change, and in which direction? These metrics require looking beyond the agent’s perimeter to the system it operates within. 

The second is financial impact, calculated honestly. This means accounting for the full cost of the deployment — not just development and licensing, but infrastructure at scale, governance overhead, monitoring, human review of edge cases, and the ongoing cost of maintaining and improving the model. Against that full cost, what is the measurable revenue impact, cost reduction, or risk reduction that can be attributed to the agent? The attribution question is genuinely hard, and organizations that skip it tend to overstate returns significantly. 

The third is strategic impact, which is harder to quantify but critical to track. Is the organization accumulating data, process knowledge, and deployment capability that compounds over time? Are the use cases the agent handles expanding in complexity and value, or has it plateaued at the same low-stakes tasks it started with? The organizations that generate durable value from AI agents are those where each deployment creates the infrastructure — technical and organizational — for the next one. 

 

 

The baseline problem 

 

One of the most common reasons AI ROI measurement fails is that the baseline was never properly defined before deployment. 

If you don’t know exactly how long the process took before, with what error rate, at what cost, and with what outcome quality, you have no reference point for measuring improvement. This sounds obvious. In practice, it’s frequently skipped — either because the pre-deployment state was never systematically measured, or because the urgency to deploy overrode the discipline to baseline first. 

McKinsey’s research on AI high performers identifies baseline definition as one of the clearest differentiators between organizations that can demonstrate ROI and those that can’t. High performers are nearly three times as likely to have defined specific, measurable success criteria before deployment — not after the pilot succeeds and the pressure to justify scaling begins. 

The practical implication: if you’re planning an AI agent deployment and you don’t have a clear, quantified picture of the current-state process you’re improving, that’s the first thing to fix. Not the model selection. Not the integration architecture. The baseline. 

 

 

What a measurement framework actually looks like 

 

A functional ROI measurement framework for AI agents has four components that need to be defined before go-live, not retrofitted after. 

Success metrics tied to business outcomes, not agent activity. Define what the business result looks like in measurable terms: reduction in processing time for a specific workflow, decrease in error rate at a specific process step, improvement in a customer satisfaction score for interactions handled by the agent. These need to be specific enough to be attributable. 

A pre-deployment baseline with the same metrics applied to the current-state process. This is the reference point everything else is measured against. 

A cost accounting model that includes the full cost of ownership: development, infrastructure, licensing, governance, monitoring, and human oversight. Without this, efficiency gains look better than they are. 

A review cadence that separates short-term efficiency metrics — which are visible immediately — from medium-term outcome metrics — which take three to six months to stabilize — and long-term strategic metrics, which require a year or more of production data to assess meaningfully. Conflating these timeframes produces either false positives early or premature conclusions that the deployment isn’t working. 

 

 

The compounding advantage 

 

There’s a reason the organizations that measure well tend to scale well. Rigorous measurement produces something beyond accountability: it produces a learning loop. Each deployment cycle generates data about what the agent does well, where it fails, what human oversight catches, and what business outcomes correlate with what agent behaviors. 

That learning loop is the actual competitive advantage of early AI agent deployment — not the efficiency of any single use case, but the accumulating organizational capability to deploy, measure, improve, and expand. The organizations that are building that capability now, with the measurement infrastructure to support it, are the ones that will have a compounding advantage as the technology matures. 

The ones measuring API calls and calling it ROI will eventually have to rebuild their framework from scratch. That’s an expensive correction to make after significant investment. 

If your organization is deploying AI agents and you’re not confident in your measurement framework, that’s worth fixing before the next deployment cycle — not after.

Dedicated squad, staffing, turnkey, or managed evolution: How to choose the right model for where you actually are

Dedicated squad, staffing, turnkey, or managed evolution: How to choose the right model for where you actually are

 

 

The decision most organizations get wrong before the work even starts 

 

When a US organization decides to bring in a nearshore team, there’s a conversation that happens early and often gets resolved too quickly: how do we structure the engagement? 

The most common answer is whatever the provider defaults to, or whatever worked on the last project. Both are reasonable starting points. Neither is a strategy.  

The engagement model shapes everything that follows: how fast the team can move, how much coordination overhead the client absorbs, who owns quality, and whether the relationship has room to grow as the work evolves. Getting it wrong doesn’t show up as a single visible failure. It shows up as friction that compounds quietly across every sprint until someone asks why the project is three months behind a realistic schedule. 

 There are four models worth understanding clearly. Not as a menu to pick from, but as tools with specific use cases and specific failure modes when applied to the wrong situation. 

 

 

Dedicated Squad: when context is your most valuable asset  

A dedicated squad is a complete delivery team embedded in your development cycle. Architects, engineers, QA, and often a tech lead operating as a functional extension of your organization. They build context over time. They understand the codebase, the stakeholders, the unwritten constraints, and the history of decisions that led to the current architecture.  

This model performs best on ongoing product development, platform modernization, or complex builds where accumulated context is genuinely irreplaceable. The cost of rebuilding that context after a team rotation is not theoretical… it is measurable in ramp-up weeks, in bugs that resurface because no one remembered why a constraint existed, and in architectural decisions that get revisited because the reasoning was never transferred.  

The failure mode for dedicated squads is applying them to work that is actually bounded. If the scope is well-defined, the timeline is fixed, and the deliverable is clear, a dedicated squad introduces more overhead than the situation requires. You’re paying for context-building on a project that doesn’t need it. 

 

 

Staffing: filling a specific gap without changing the structure 

 

Staffing is the most straightforward model: specific profiles integrated directly into your existing team to fill a capability gap or accelerate capacity for a defined period. A senior backend engineer with healthcare interoperability experience. A machine learning engineer to carry a specific initiative through its first production deployment. A QA lead to establish testing standards before handing them off internally. 

 Done well, staffing doesn’t create dependency. It fills a gap, transfers knowledge, and leaves the client’s team stronger than it found it. Done poorly, it creates a situation where the client is managing individuals rather than working with a partner. 

 The signal that staffing is the right model: you have a well-functioning internal team with a clear gap, and you need someone who integrates into your culture and process rather than running a parallel one. The signal that it’s the wrong model: the “gap” is actually a structural capacity problem that will resurface as soon as the staffed individual rotates off. 

 

 

Turnkey: when outcome accountability matters more than visibility into the process 

 

A turnkey engagement is defined by what gets delivered, not by how many people are working on it or what their hours look like. Scope, timeline, and quality criteria are agreed upfront. The provider owns the delivery. 

 This model works well for projects with genuinely clear boundaries: a defined integration, a specific product feature, a migration with a clear before and after state. The client gets outcome accountability without the overhead of managing a distributed process. The provider has the room to allocate resources, adjust internal processes, and make technical decisions without running every choice through a client approval cycle. 

 The failure mode is scope ambiguity. Turnkey on a well-defined project is efficient. Turnkey on a project where the requirements will evolve creates conflict at every boundary. When the scope shifts, the contract becomes the conversation instead of the work. That’s an expensive place to spend energy. 

 Before choosing turnkey, the honest question is: do we actually know what we want built, or do we know what problem we’re trying to solve? Those are different situations that require different models. 

 

 

Managed Evolution: the model organizations forget to plan for 

 

Managed evolution is ongoing maintenance, enhancement, and operational support for applications in production. It covers compliance updates, integration with evolving platforms, incremental feature development, and the continuous work of keeping a system healthy as the environment around it changes. 

It is also the model most organizations fail to think about during the initial build. The project gets scoped, the team gets selected, the delivery happens… and then the question of who maintains and evolves the system is treated as a separate decision, often made under pressure after the original team has moved on.  

The organizations that handle this best are the ones that design for it from the beginning. They choose an initial engagement model that has a natural path to managed evolution. The handoff cost in that scenario is near zero. The handoff cost when the original team is completely different from the maintenance team is measured in months. 

 

 

The model mismatch problem 

 

The reason engagement model decisions go wrong most often is not that the options are poorly understood. It’s that organizations choose based on cost or habit rather than based on where they are in their technology roadmap and what the work actually requires. 

A staffing model applied to a complex platform modernization leaves the client managing individuals without a coherent delivery process. A turnkey model applied to a discovery-heavy initiative creates scope conflict at every turn. A dedicated squad applied to a bounded, well-specified integration introduces overhead the project doesn’t need.  

The more useful question before signing anything is not which model is cheapest. It’s: what does this specific project require from the people working on it, and what does it require from us as a client? Projects that need context-building over time need a squad. Projects that need a specific skill for a defined period need staffing. Projects with clear deliverables and tight timelines need turnkey. Systems in production that need to keep evolving need a managed evolution model, and ideally a team that was there for the original build. 

 

 

Flexibility as a structural feature, not a sales pitch 

 

One marker of a mature nearshore partner is the ability to move between models as the engagement evolves. A POC that starts as a turnkey engagement and transitions into a dedicated squad when the scope expands. A staffing arrangement that grows into a squad when the client’s internal team is ready to absorb more coordination. A dedicated squad that shifts into managed evolution after the initial platform build stabilizes. 

That kind of flexibility requires a provider with enough process maturity and team depth to restructure without losing delivery quality. It also requires a client willing to revisit the engagement model as a strategic decision rather than a contractual default.

The Nearshore Advantage

The Nearshore Advantage

 

The US technology talent shortage is not a hiring problem. It is a structural gap that keeps widening as AI-driven demand accelerates faster than any domestic pipeline can absorb. The organizations that gained ground over the last five years did not simply hire faster. They built better team models.

This whitepaper covers what the data actually shows about nearshore IT teams in 2026: the real economics behind the rate card, the industries where the model delivers documented results, and the criteria that separate a strategic partner from a commodity provider.

 

In this report you will find:

Why the total cost of offshore development is not what appears on the spreadsheet.
How time zone alignment directly impacts sprint velocity, code quality, and project outcomes.
The documented impact on healthcare, insurance, and financial services organizations.
What mature AI integration in the SDLC looks like and how to tell if a provider actually has it.
Three questions to ask before evaluating any nearshore partner.

Read the full report here

 

 

The enterprise AI paradox: everyone’s in, almost no one’s ready

The enterprise AI paradox: everyone’s in, almost no one’s ready

 

 

“The gap between enterprise ambition and production-ready AI is wider than most organizations admit…  and it has nothing to do with the technology.” 

  

Ask any enterprise leader in 2026 whether AI agents are a priority, and the answer is almost universally yes. 92% of companies plan to increase their AI spending over the next three years. Boardroom conversations have shifted. 34% of chief executives now identify AI as their top strategic theme, replacing digital transformation after decades at the top of the agenda. 

And yet, the production numbers tell a different story. 

Only 1% of companies consider themselves mature in AI, meaning AI is fully integrated into their operations. Fewer than 10% of deployed AI use cases make it past the pilot stage. According to IDC, 88% of AI proof-of-concepts never reach production. 

This is the defining tension of enterprise AI in 2026: enormous ambition, modest execution. And understanding why that gap exists (and how to close it) is the most important question technology leaders should be asking right now. 

 

The pilot trap 

 

Most organizations aren’t failing to start with AI. They’re failing to finish. 

60% of organizations are still primarily investing in pilots, and since 2023 only 25% of AI initiatives have delivered expected ROI. The pattern is consistent across industries: a promising proof of concept, early enthusiasm, a working demo, and then a slow stall when it comes time to move into production. 

The reasons are rarely technical. 70% of organizations discover that their data infrastructure is fundamentally lacking only after launching ambitious AI initiatives. That’s typically six months in, after a successful pilot, when the foundational systems can’t handle production workloads. 

In other words, the technology works. The organization isn’t ready for it. 

  

What actually separates winners from the rest 

 

The research is consistent on what differentiates organizations that generate real value from AI versus those that accumulate a graveyard of pilots. 

AI high performers are nearly three times as likely as others to say their organizations have fundamentally redesigned individual workflows. They don’t layer AI onto existing processes, instead they redesign the process around what AI can do. That distinction sounds subtle. In practice, it’s the difference between a chatbot that answers FAQs and an agent that resolves customer issues end to end. 

McKinsey also reports that 65% of AI high performers have defined human-in-the-loop processes, compared to only 23% of other organizations. Governance isn’t a constraint on deployment speed. It’s what makes deployment sustainable. 

And leadership engagement matters more than most organizations expect: 33% of high performers have senior leaders actively driving AI adoption, compared to significantly fewer in the general pool. AI transformation doesn’t happen bottom-up. It requires executives who treat it as a strategic operating model change, not a technology project delegated to IT. 

  

The agentic shift changes the stakes

 

While most organizations are still wrestling with basic GenAI deployment, the frontier has already moved. Agentic AI is becoming the new baseline expectation. 

By the end of 2026, 40% of enterprise applications will include task-specific AI agents, according to Gartner. PwC’s research shows that 79% of organizations are already using AI agents to some degree, with 88% planning budget increases specifically for agentic capabilities. 66% report measurable productivity improvements, and 62% expect ROI exceeding 100%. 

But the same dynamics that stall basic AI deployment apply at the agentic level, amplified. By 2027, organizations that don’t prioritize high-quality, AI-ready data are expected to suffer around a 15% productivity loss when trying to scale agentic solutions. The foundation matters more as the systems become more autonomous. 

  

The governance problem nobody wants to talk about 

 

There’s an uncomfortable reality buried in the research that doesn’t get enough attention: at 25% AI agent adoption, application development costs could rise approximately 16% and governance costs could increase over 34%. 

Deploying AI agents without governance infrastructure doesn’t just create risk — it creates cost. Runaway infrastructure spend, agents behaving outside policy boundaries, decisions that can’t be audited or explained. These aren’t edge cases. They’re the most common reasons projects get canceled after significant investment. 

The organizations that win with agentic AI will be those that treat it as an operating model and change program, not just a technology rollout. That means governance, observability, and clear business outcomes defined before a single line of code is written, not retrofitted after the pilot succeeds. 

  

The window is open, but it won’t stay that way 

 

Organizations that establish agent capabilities early accumulate data, experience, and process advantages that compound over time, creating competitive moats that become increasingly difficult for competitors to replicate. 

Having an agile product delivery organization with well-defined delivery processes is one of the factors most strongly correlated with achieving real value from AI. The organizations that are winning aren’t necessarily the ones with the biggest AI budgets. They’re the ones that combine technical capability with delivery discipline: short cycles, measurable checkpoints, and organizational maturity to move from pilot to production without losing momentum. 

The gap between ambition and execution in enterprise AI isn’t a technology problem. It’s a delivery problem. And in 2026, that distinction matters more than ever. 

  

At Huenei, we help companies bridge that gap. From strategy to production-ready AI, with the agile delivery process and governance model to make it stick. 

 

Want to see how we approach it? Let’s talk!

From Pilot to Production: The Real State of AI Agents in 2026

From Pilot to Production: The Real State of AI Agents in 2026

 

The organizations pulling ahead are not evaluating whether to deploy AI agents. They already have them running in production and they are measuring how many processes still lack one.

This whitepaper covers where the market actually stands, which architectures hold up in real deployments, where the ROI is most documented by industry, and why most projects never make it to production.

 

In this report you will find:

  • Why 2026 is the year AI agents moved from experiment to enterprise infrastructure
  • The measurable impact across healthcare, insurance, and financial services
  • The four architectures that dominate in production and when to use each one
  • A maturity model to assess exactly where your organization stands today
  • How we build and operate agents in production — including our own