
Why AI training ROI feels impossible to measure
You bought the AI tools. You ran the training. A few weeks later leadership asks the obvious question: so what did we actually get for it? And the room goes quiet.
You are not behind. This is the most common conversation in Indian L&D right now. Companies have spent real money putting AI in front of their teams, and almost none can put a number on whether it worked. Three things quietly break the measurement.
- Usage is invisible. You can see licenses bought and logins, not whether people use AI well, rarely, or as expensive autocomplete. A license is an input, not an outcome.
- Trained is not fluent. A workshop attended tells you someone was in the room, not that they changed how they work the following Monday.
- Confidence hides the gap. Heavy users assume they use AI well. The two are not the same, and self-report surveys make it worse, not better.
The jagged frontier: why your best people fail silently
A well-known field experiment from BCG and Harvard Business School (Dell'Acqua et al., Organization Science) found that on tasks inside AI's capability, consultants using it were markedly faster and produced higher-quality work. On tasks just outside that capability, AI made them worse, and they often could not tell.
25% faster, 40% better AI's lift on tasks inside its frontier, and a quiet decline on tasks outside it (BCG and Harvard)
The researchers called this the jagged frontier. AI is brilliant on one side of an invisible line and unreliable on the other. This is why headcount using a tool tells you nothing. Your most confident users can be the ones quietly shipping worse work, because they cannot see which side of the line a task falls on. That is what training has to fix, and what you have to measure.
The number your CFO actually wants
Finance does not want 142 people completed AI training. It wants a figure it can defend in a budget review. That number has a shape.
AI-exposable hours per week, times the cost of those hours, times a realistic productivity lift, adjusted for how ready your team actually is to adopt.
This is a Phillips-style ROI calculation, the same method used to justify any serious L&D spend. It ties a soft-sounding capability, AI fluency, to a hard number: this team spends X hours a week on work AI can accelerate, and closing the fluency gap could give back Y of those hours every week. Say that sentence and the AI training conversation stops being a cost debate and becomes an investment decision. Your own finance team knows what an hour of that team costs; the point of handing over hours rather than a rupee total is that the assumption stays theirs, in the open, instead of being buried inside our arithmetic.
How to measure it in two weeks
You do not need a year-long study. You need three things.
- Score real fluency, not a quiz. A short conversational diagnostic reveals how a person actually works with AI: what they use it for, where they stop, whether they check the output, whether they can tell a good result from a plausible wrong one.
- Place everyone on a maturity map. Five levels, from not using AI to a multiplier who builds and teaches. Now the whole team is on one picture and the gaps are obvious.
- Count the hours. Map each level to AI-exposable hours and a realistic lift, and you have the CFO number, per team and per role.
Two weeks, and you move from we think it is going well to here is where we stand and what it is worth to move up.
What a board-ready AI ROI gap report contains
- A team maturity heatmap: every person, every role, one glance.
- The top three gaps holding the team back.
- The hours a week that closing them could realistically give back, per team and per role.
- The specific workshop that closes the biggest gap.
- A re-measure plan so you can show realized ROI, not just projected.
That is the document you hand your CFO. It is the difference between asking for a training budget and presenting a return. Our AI Litmus diagnostic builds it in two weeks, grounded in your team's real workflows and fluency, with every figure labelled by how firm it is.
Producing that number is what AI Litmus is for. It is AI adoption software that measures how effectively, not just how much, each role uses the AI you already own, then reports the hours those gaps cost. The companion reads are how to measure AI adoption and whether your team is actually using AI well.
See this on your own teams.
A private walkthrough, calibrated to your roles. About two weeks.
Frequently asked
How do you measure AI literacy objectively? Through a structured conversational diagnostic that scores behaviour, not opinions. It looks at how someone actually uses AI, whether they verify output, and how they handle tasks near the edge of what AI does well. Scoring is consistent across people, so teams are comparable.
Isn't AI ROI just soft benefits? No. Time is the hard benefit. If a team spends measurable hours a week on work AI can accelerate, and you know a realistic lift, the value is a count of hours, not a feeling. What an hour is worth is your finance team's number to apply, not ours to assume.
How is this different from a training completion report? A completion report counts attendance. A gap report measures capability before and after, and attaches a number of hours to the change. One proves people showed up. The other proves the training was worth doing.
What size team do you need to start? It works from a single team of ten. Most companies start with one function, prove the number, then expand.
Shobhit Khandelwal
Founder, VMS Culture Labs
Shobhit Khandelwal is the founder of VMS Culture Labs, on a mission to measure what most leaders only guess at: how fluently their teams truly work with AI, and the hidden cost of how people behave at work. He is out to replace workplace guesswork with evidence, and build the kind of workplaces the next generation deserves.
Get the intelligence briefing
Occasional, high-signal notes on measuring AI fluency and culture cost. No spam.
Keep reading