PHASE 08

SEVEN WAYS THIS
GOES WRONG.

These are not exotic. The same small set of causes accounts for nearly every operations project that quietly stops being used, and every one of them is visible in advance if you know the shape.

AI-First Operations › Why It Fails

One: automating an undefined process

The root cause underneath most of the others. A process that lives in people's habits gets handed to a system, and the ambiguity goes with it.

What you get is not a defined process running automatically. It is an undefined process running faster, in more places, with less visibility. When it produces something wrong in month three, nobody can say what it was supposed to do, because nobody ever wrote that down.

The tell is early and obvious. If you cannot describe the process in steps, with a definition of done at each one, on a page, it is not ready to be automated. That is true of a new hire as well, and nobody would hand a new hire a job they could not describe.

Two: buying tools first

A demo is impressive, a decision gets made, and the process is then bent around whatever the software assumes.

Sometimes that works, because the software's assumptions are reasonable and your process was arbitrary. More often you end up with a tool configured around confusion, and the conclusion drawn six months later is that it was the wrong tool. So a different one is bought and configured around the same confusion. I have seen businesses do this three times with three platforms and diagnose it as bad luck with vendors.

Software executes a process. It never supplies one. The order is not stylistic.

Three: no owner for the system

The project has a champion during implementation. Afterward it belongs to nobody, which means it decays and nobody is watching.

Systems need continuous ownership, not a launch. Someone has to notice when a sync has been failing for two weeks, handle the case the design did not anticipate, update it when the business changes, and answer the questions new people ask. Without that, drift is not a risk. It is a certainty on a schedule.

One name, not a committee, and not a general expectation that the team will look after it. And they need authority: to change it, to fix it, and to switch it off without asking permission first.

Four and five: nothing measured, everything at once

Four: no measurement

Nothing was counted before, so nothing can be compared after. The project is then defended by whoever feels most strongly about it.

This has a specific downstream consequence beyond not knowing. Without evidence, there is no basis for deciding what to do next, so the second cycle is chosen by instinct, and the instinct that produced the current problems is the one doing the choosing.

Two weeks of counting one thing, written down with a date, before you change anything. That is the entire requirement and it is skipped almost every time.

Five: too much at once

Enthusiasm produces scope. Since we are fixing follow-up, we should redo the CRM. And while we are in there, the job tracking. And a reporting dashboard.

Three failure modes arrive together. Nothing ships for months, so momentum dies. When something breaks you cannot isolate the cause among twelve simultaneous changes. And the team, asked to absorb everything at once, absorbs very little and reverts to the old way for the parts they can.

One handoff. Then one automation. Then measure. The narrowness is not timidity. It is what makes diagnosis possible.

Six: the team is not brought along

The system is designed in a room the people who do the work were not in, announced as an improvement, and quietly worked around within a month.

They work around it for good reasons, which is the part owners miss. The design missed the exception that happens twice a week. It adds a step that serves a dashboard rather than the work. Nobody explained why, so it reads as surveillance. Or they were never asked, and the message received was that their knowledge of the process did not matter.

Involve them in the mapping, be straight about what changes for them, and make the first improvement remove something they find tedious. Then listen when they say it is wrong, because they are usually right and they will stop telling you if the first few reports go nowhere.

Seven: no plan for when the automation is wrong

The last one, and the most dangerous, because it converts a small error into an incident.

Everything is built for the path where it works. Nobody asked the other question: when this does something wrong, who notices, how fast, and what happens then.

Automation changes the shape of error. A person making a mistake makes one, and usually notices. A system making a mistake makes it on every item in the queue, consistently, until someone checks the queue. The mistakes also look normal, because the output is well-formatted by construction.

Four things, decided before launch, not after the first bad week.

  • Detection. Who looks at what, how often. A log nobody reads is not oversight.
  • A stop. One named person who can switch it off immediately without escalating.
  • A fallback. How the work gets done manually while it is off, written down before you need it.
  • Recovery. What you do about the items already processed wrong, including what you say to the people affected.

Anything irreversible — money moving, a contract, a message to a customer in a sensitive situation — keeps a person on the approval step permanently. Not as a starting precaution. As the design.

Frequently asked

Questions people actually ask

What is the single most common cause of failure?

Automating a process that was never defined. It sits underneath most of the other six, because an undefined process also makes measurement impossible, ownership ambiguous, and team resistance entirely reasonable.

How do I know if my process is defined enough?

Write it down in steps with a definition of done at each one. If someone who has not done the job could follow it and produce the right result, it is defined. If it needs interpretation at three points, it is not.

What if my team resists the new system?

Assume they have a reason and go find it. The usual causes are a design that missed a real exception, a step added for reporting rather than for the work, or nobody having explained why. All three are fixable and none is fixed by pushing harder.

Who should own an automation after launch?

One named person with authority to change it, fix it and shut it off. Not a committee, not the vendor, and not a general expectation that everyone will keep an eye on it.

What should I do if an automation has been sending the wrong thing?

Switch it off first, then determine the scope of what went out, then contact the people affected directly rather than hoping it goes unnoticed. The recovery plan should have been written before launch, which is why it is on the list.

Is it ever too late to fix a failed implementation?

No, but restarting usually means going back to the mapping phase rather than reconfiguring what exists. A system built on an undefined process cannot be repaired by adjusting settings, because the missing piece was never a setting.

Make your next move

A year from now, what will you be glad you started today?

You don't need another promise that everything will be easy. You need something useful to learn — and a next step you're willing to take.