AI adoption is only the beginning. See the whole transformation.Explore the platform
VMS Culture LabsTalk to our team

AI LITMUS / INSIGHTS

How to find the automation that's actually worth doing

Most automation lists come out of a workshop and describe what annoys people rather than what costs money. Here are the criteria that hold up, the one every standard checklist leaves out, and how to find candidates from the week your team actually has.

How to find the automation that's actually worth doing

Most automation lists describe what is annoying, not what is expensive

The usual way this gets done is a workshop. People are asked what they would like to stop doing, the answers go on a board, and the board becomes a roadmap. The output is reliably a list of the most irritating tasks in the room rather than the most costly ones, and those two overlap far less than anyone expects.

Irritation scales with how unpleasant something is. Cost scales with how long it takes multiplied by how often it happens. A task that takes four minutes and happens twice a day costs more than one that takes an hour and happens twice a year, and only one of them gets mentioned in a workshop.

The criteria that hold up

The RPA era produced a genuinely useful checklist and it did not stop being true when the tooling changed. Solvexia's version is a good statement of it. A candidate tends to be a process that:

  • Takes a non-trivial amount of hands-on manual time, typically more than a couple of hours a month.
  • Runs frequently, typically at least monthly.
  • Is made up of ten or more steps.
  • Works with multiple data files or data from multiple legacy systems.
  • Has a high cost of error, in real or perceived terms.
  • Carries significant key-person dependency risk, or a high risk of human error because of its complexity.

And the counter-criterion, which gets ignored more often than any of the others: if a process is unstable and changes frequently, it may not be worth automating. Automating something mid-redesign means building twice and getting blamed once.

What every one of those checklists leaves out

Read that list again and notice what is not on it. There is no criterion for judgment. Nothing on it distinguishes a step where a person is following a rule from a step where a person is making a call, and a step where someone is making a call can satisfy every other criterion on the list perfectly.

That is the shape of the expensive mistake. A judgment step that is frequent, multi-system, error-prone and time-consuming scores extremely well on the standard checklist, which is precisely why it gets automated. Then it produces confident, plausible, well formatted decisions, and because the failure is quiet nobody notices for a quarter.

So the criteria need one more line before they are safe to use: what has to stay human, and why. Marking that explicitly before anything is ranked is the cheapest insurance in this whole exercise.

The second gap: the checklist cannot tell you which processes you have

A checklist is a filter. It is useless without a list to filter, and the list is the hard part. Most companies do not have an accurate account of how their teams actually spend a week, which means the automation exercise starts by guessing at the input and then applying rigorous criteria to the guess.

The steps that qualify best are usually the ones nobody wrote down. Reformatting something for a system that should have accepted it as it was. Chasing an approval. Rebuilding the same weekly summary from four places. None of those appear in a process document, because none of them were designed. They accumulated, and that is exactly why they are inefficient enough to be worth automating.

Which is why this has to run in the order it does. Map the real week first, then filter it. The method for the first half is in map how the work actually happens.

A one-off audit ages out. The work does not stand still.

The other structural problem with the workshop model is that it produces a snapshot. Six months later the tools have changed, two processes have been redesigned, a team has reorganised, and the list describes a company that no longer exists. It gets quietly abandoned rather than formally retired, which is why so many organisations have an automation backlog nobody has looked at since the offsite.

The alternative is to attach the discovery to something that already happens continuously. If someone is regularly talking to every individual about how their work goes, the automation candidates surface as a by-product, and they update as the work updates rather than needing another offsite.

Rank by recoverable hours, not by enthusiasm

Once you have real candidates, the ranking is arithmetic rather than debate. Time per run, multiplied by runs per period, multiplied by how many people do it. That produces hours. Hours are comparable across teams in a way that a priority score assigned in a meeting is not.

Two disciplines keep that number honest. Separate hours that are theoretically recoverable from hours something has actually given back, and never report the two as one figure. And hold the ranking against the human-only list, so that a high-scoring candidate that turns out to be a judgment call is removed rather than argued about later.

One more constraint worth planning for: an automated process still has people around it, and they need to be good enough with the tools to supervise it. Prosci's research across 1,107 participants found user proficiency is the single largest challenge in AI adoption, named by 38 percent of respondents. Automating a step in front of a team that cannot check the output moves the risk rather than removing it.

Where the candidates come from

AI Litmus produces this list as a by-product of something it does anyway. Its first job is coaching each person to use the tools they already have properly, and while it does that it is continuously watching for work in the process that should not be done by a person at all. Because the signal comes from every individual rather than from a workshop, the list reflects how the work really runs, it marks what must stay human, and it keeps changing as the work does. Start with whether your team is using AI well for the capability half.

Frequently asked

How do you identify automation opportunities in a team? Map how the week actually runs first, then filter it through the established criteria: the step takes real manual time, runs at least monthly, has ten or more steps, crosses multiple files or legacy systems, and carries a high cost of error or key-person risk. The filter is easy; the list to filter is the hard part, because the best candidates are usually undocumented steps that accumulated rather than being designed.

Which processes should not be automated? Two kinds. Processes that are unstable and change frequently, because you will build it twice. And any step where a person is making a judgment rather than following a rule, which the standard checklists do not screen for at all. A judgment step can score perfectly on every other criterion, which is exactly why it gets automated, and its failures are quiet and confident rather than loud.

What makes a good automation candidate? Non-trivial hands-on time, typically more than a couple of hours a month. Frequent, typically at least monthly. Ten or more steps. Multiple data files or several legacy systems. A high cost of error, real or perceived. Significant key-person dependency, or a high risk of human error from complexity. And, crucially, rule-following rather than judgment.

Why do automation workshops produce the wrong list? Because they surface irritation rather than cost. Irritation scales with how unpleasant a task feels; cost scales with duration multiplied by frequency multiplied by headcount. A four-minute task done twice a day costs more than an hour-long task done twice a year, and only one of those gets raised in a room. Workshops also produce a snapshot that ages out as soon as tools or teams change.

How should automation opportunities be ranked? By recoverable hours: time per run multiplied by runs per period multiplied by the number of people doing it. Hours compare across teams in a way that a priority score assigned in a meeting does not. Keep hours that are theoretically recoverable separate from hours something has actually given back, and check every high-ranking candidate against the list of what must stay human before it goes on a roadmap.

Related: how AI Litmus finds automation worth doing.

Explore all articles

START WITH ONE TEAM.

A clearer picture.
A practical next move.

Bring your roles and the AI tools you already own.

Talk through these ideas

A QUICK REFLECTION

How well does your team use its AI tools?

On a scale of 1 to 10, where would you put your team today?

Your team’s use of AI, from 1 to 10
1 · Just getting started10 · Working really well