AI adoption is only the beginning. See the whole transformation.Explore the platform
VMS Culture LabsTalk to our team

AI LITMUS / INSIGHTS

Your team uses AI. But are they using it well?

Adoption tells you how many people opened the tool. It cannot tell you whether the work got better. Here is what AI effectiveness actually means, why no usage dashboard can measure it, and how to read it role by role.

Your team uses AI. But are they using it well?

Adoption answers how much. It cannot answer how well.

Two different questions get collapsed into one word. The first is how much AI is being used: seats filled, people logged in, messages sent. The second is how well it is being used: whether the output is any good, whether the person can tell when it is wrong, whether the job actually changed shape. Only the first one is easy to count, so it is the one that ends up on the slide.

Prosci splits AI measurement into four layers. Layer one is activation and access, which asks whether employees are using the tool and how many active users there are. Layer two is behavioural change and proficiency, which asks whether employees are working differently and whether they use AI tools effectively in their real workflows. Layers three and four cover business impact and then culture, governance and trust. Most organisations report layer one and call it adoption. (Prosci)

The monitoring vendors describe the same split from the inside. Insightful, which sells workforce analytics, puts it bluntly: "AI dashboards report adoption, not absorption." Their own definition of absorption is when AI stops being a side tool and becomes part of real workflows, and they are direct about the consequence: a workforce can show high adoption and low absorption at the same time. (Insightful)

That sentence is worth sitting with, because it is a vendor in this category describing the limit of the category. A seat report is a record that a tool was opened. It was never evidence that anyone got better at their job.

  • What a usage log can see: who has a licence, who opened it, how often, how many messages, which features were clicked.
  • What a usage log cannot see: whether the answer was any good, whether the person checked it, whether they would have caught it if it were wrong, and whether the work they produced is better than the work they produced before.

Proficiency is the biggest blocker, and the one least likely to be on your dashboard

38%named user proficiency as the single largest challenge in AI adoption, more than double technical integration at 16 percent and organisational adoption at 15 percent. (Prosci, Keys to Unlocking AI Adoption, 1,107 participants)

There is an uncomfortable symmetry in that finding. The problem most likely to be holding your AI programme back is the one least likely to appear in any system you already own. Integration problems surface in error logs. Access problems surface in licence reports. Proficiency surfaces nowhere, because there is no event a tool can emit that means the person using it did not really know what they were doing.

So the measurement gap is not an oversight. It is structural. You cannot instrument your way to this number from the outside of the work, which is why the companies with the most detailed AI dashboards are often the least able to say whether any of it is working.

What "using AI well" actually means

Effectiveness is how well a person uses AI for their particular job, rather than how much they know about AI in general. It is not a quiz score and it is not a certificate. It shows up in the work or it does not exist.

In practice it decomposes into a small number of capabilities that fail independently of each other. Someone can be excellent at one and hopeless at another, and the two have nothing to do with one another:

  • Prompting and direction. Can they get the tool to do the thing they meant, including giving it the context it needs rather than the context they happen to have open.
  • AI and tool literacy. Do they know which of the tools they already have is right for this task, and do they know what each one is bad at.
  • Workflow integration. Has the tool moved into the actual sequence of their week, or does it sit beside the work as an occasional detour.
  • Critical thinking and judgment. Can they tell when the answer is confident and wrong, and do they check the things worth checking.
  • Growth and team influence. Do they get better over time, and does anyone else get better because of them.

The fourth one is the expensive one. AI rarely fails loudly. It hands back a fluent, well formatted answer that is quietly incorrect, and a person who cannot catch that is not a light user of AI, they are a liability with a licence. We wrote about that failure mode separately in why AI fails confidently rather than loudly.

Why a single effectiveness score hides the answer

Because those capabilities fail independently, averaging them destroys the information you needed. Two people can land on exactly the same overall number and need opposite interventions.

The first is fast, productive, and ships work nobody checked. That is the more expensive of the two, because the damage arrives downstream and arrives looking finished. The second is careful and reliable and has never moved AI into the shape of their week, so whatever they gain stops at their own desk and never reaches the team. One needs a verification habit. The other needs their workflow redesigned. Neither needs what the other needs, and the average needs nothing at all.

Two people with the same average fluency score and opposite problemsOne person scores high on prompting and low on judgment, so they ship fast and ship wrong. The other scores high on judgment and lower on prompting and workflow, so their work is trustworthy but does not compound. Both average 3.6, and each needs a completely different intervention.Ships fast, ships wrongStrong prompting, weak judgmentPrompting5Judgment2Tool literacy4Workflow4Influence3AVERAGE 3.6Trustworthy, not compoundingStrong judgment, weak workflowPrompting3Judgment5Tool literacy3Workflow3Influence4AVERAGE 3.6
The same average, two different problems, two opposite fixes. One number would send both people to the same training session.

Effectiveness is role-shaped, so the bar has to be too

What counts as good differs by job, and not slightly. For a salesperson the thing that matters most might be verifying AI generated account research before it reaches a customer. For a marketer it is reviewing claims before they are published. For someone in operations it is spotting the exception that an automated workflow got wrong.

Score all three against one company-wide bar and you get a number that describes none of them. This is the same reason a generic readiness test tells you so little, which we covered in why generic AI readiness tests fail.

The practical consequence: an effectiveness read is only useful if the thing it is compared against is the role, not the company. That means someone has to know what each role needs before anyone is scored.

How to actually measure it

The method follows from the constraint. If the signal is not in the logs, it has to come from the work. Four steps:

  • Define what good looks like for each role first. Not for the company. Which tasks in this job can AI carry, which must stay human, and what would a strong performer do differently from a weak one.
  • Have a real conversation with each person about the work they actually do. Not a survey with a scale on it. People describe their week accurately when asked about it and inaccurately when asked to rate themselves out of five.
  • Watch them do real tasks from their own role, using the tools your company already provides. This is where the difference between a confident user and a good one becomes visible, and it is the only part that cannot be faked.
  • Score each capability separately, against that role's bar, and never average them into one number you then report.

Two things to avoid. Do not use a general knowledge quiz about AI, because it measures reading, not doing. And do not let a manager see an individual result, because the moment people believe the output is going into a performance review, you stop measuring effectiveness and start measuring impression management.

What changes once you can answer how well

Enablement stops being generic. Instead of one AI session for everyone, each team gets the specific move that fits what is actually missing, which is usually not the thing anybody guessed.

Budget arguments get easier. The adoption gap has two halves, licences nobody opens and tools people use but have outgrown, and only the second is visible through an effectiveness read. Knowing which one you have tells you whether the next spend should be a new tool or the people you already have. That is the argument in the hidden AI adoption gap.

And the ROI conversation changes character. Once you know which stages of a role's week are slowed by a capability gap rather than by a missing tool, you can say how many hours are recoverable and where from, instead of asserting a percentage improvement nobody can trace.

Start with one team

This does not need a programme. It needs one team, a definition of what their roles require, and an honest read of how well each person works with the tools you already pay for.

That is what AI Litmus is built to do: AI adoption software that measures how effectively, not just how much, your people use the AI you already own, and turns it into the five things an AI transformation lead is asked for. If you want the other half of the picture first, start with how to measure AI adoption.

Frequently asked

What is the difference between AI adoption and AI effectiveness? Adoption is how much AI is being used: licences filled, people logged in, messages sent. Effectiveness is how well it is being used: whether the output is good, whether the person catches it when it is wrong, and whether the work itself changed. Adoption can be at one hundred percent in a team where nothing has improved. Prosci separates these as layer one, activation and access, and layer two, behavioural change and proficiency, and most reporting stops at layer one.

How do you measure whether employees are using AI well? Not with a quiz and not with a usage report. Define what good looks like for each role first, have a real conversation with each person about the work they actually do, then watch them carry out real tasks from their own job using the tools the company already provides. Score the capabilities separately rather than averaging them, because prompting, tool literacy, workflow integration, judgment and influence fail independently of one another.

Why can't a usage dashboard measure AI effectiveness? Because there is no event a tool can emit that means the person using it did not really know what they were doing. A log can record that a licence was opened, how often, and which features were clicked. It cannot record whether the answer was correct, whether the person verified it, or whether the work produced was better than before. As one workforce analytics vendor puts it, AI dashboards report adoption, not absorption, and a workforce can show high adoption and low absorption at the same time.

What is the biggest barrier to AI adoption in organisations? User proficiency. In Prosci's Keys to Unlocking AI Adoption study of 1,107 participants, 38 percent named user proficiency as the single largest challenge, more than double the 16 percent who named technical integration and the 15 percent who named organisational adoption challenges. It is also the barrier least likely to appear in any system a company already owns, because proficiency does not surface in a log file.

Should AI effectiveness be scored per role or company-wide? Per role. What counts as competent differs by job: verifying AI generated research before it reaches a customer matters most in sales, reviewing claims before publication matters most in marketing, and catching exceptions in automated workflows matters most in operations. Scored against one company-wide bar, three people with three unrelated problems can produce the same number, and all three get sent to the same training.

Can you measure AI effectiveness without monitoring employees? Yes, and you should. Effectiveness is read from a private conversation and from real tasks a person completes in their own role, not from surveillance of their screen or their keystrokes. Individual results should not be visible to a manager. The moment people believe the output feeds a performance review, they manage the impression rather than do the work, and the measurement stops being worth anything.

Related: how AI Litmus reads AI effectiveness and fluency by role.

Explore all articles

START WITH ONE TEAM.

A clearer picture.
A practical next move.

Bring your roles and the AI tools you already own.

Talk through these ideas

A QUICK REFLECTION

How well does your team use its AI tools?

On a scale of 1 to 10, where would you put your team today?

Your team’s use of AI, from 1 to 10
1 · Just getting started10 · Working really well