Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

From 'firefighting' to learning: root causes that protect margin and lead time when the defect is not an accident, but a pattern.

11 min
Hernán Villalba Muzzin

Hernán Villalba Muzzin

Article author

Reading

# Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

From 'firefighting' to learning: root causes that protect margin and lead time.

Cover: abstract cause-effect visual board in a premium workshop, no legible text

Most companies don't have a quality problem. They have a repetition problem.

The first failure hurts. The second makes you angry. The third is no longer a failure: it's a system.

When I step into an operation where "something is always happening," I almost always find the same thing: a huge amount of effort goes into resolving the urgent, but very little into avoiding the recurrent.

And the recurrent, in project-based businesses, has a clear consequence: it eats up margin and eats up lead time, but it does so in small installments, hard to attribute, easy to normalize.

This article isn't a tool tutorial. You won't see templates, closed steps, or "apply this technique and you're done." What you will see is what truly matters: how to distinguish an RCA that learns from one that only documents, and what signals tell you if your organization is solving "so it won't come back" or simply to "get through today."

## The fundamental error: confusing “resolving” with “learning”

Resolving is putting out the fire.

Learning is redesigning the forest so it doesn't burn the same way.

In operations, resolving usually means:

    • replacing a part
    • adjusting a hinge
    • touching up a lacquer
    • resending a team
    • asking for “more care.”

That calms the client and saves the delivery. Good.

But learning means something else:

    • understanding why that defect became possible
    • why it repeats
    • why the system didn't detect it sooner
    • and what decision (or lack thereof) is fueling it.

If an organization confuses both, it becomes an expert in emergencies. And that, at first, even generates bridge. Until that pride becomes debt.

## The clearest signal that you need real RCA

When the team uses phrases like:

    • “we've seen this before”
    • “the same thing again”
    • “it depends on who does it”
    • “in the workshop it looks good, in the house it changes”
    • “we don't know why, but it happens”

That last phrase is especially dangerous. Because it doesn't describe ignorance; it describes normalization.

When a defect is normalized, it stops being an "incident" and becomes part of the process, but without control or recognized cost.

## Why RCA protects margin even if it doesn't “produce” anything

Imagen 1 — Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

There is a type of loss that almost nobody sees because it isn't invoiced as a loss:

    • calls
    • emails
    • coordination
    • rescheduling
    • second visits
    • urgent parts
    • tension between areas
    • mental load.

That is COPQ, even if nobody registers it as such.

The RCA doesn't “manufacture” parts, but it reduces that noise. And noise is margin, just wearing a different mask.

## The RCA that fails: the one done to close a file

I have seen impeccable RCAs… that change nothing.

Documents with perfect diagrams, long meetings, correct conclusions… and a month later, the same defect with another name.

When that happens, there is almost always a hidden cause: the RCA was done to demonstrate diligence, not to change the system.

And if the goal is to demonstrate, the result is paper.

## The most common blind spot: chasing the most visible symptom

In installation, for example, the final defect is seen: the edge, the adjustment, the finish.

Then the “why” becomes:

    • “the installer didn't adjust it well”
    • “the part came damaged”
    • “there was a rush”
    • “information was missing”

It might be true. But if you stay there, you are blaming the last hand.

I prefer a more uncomfortable question:

what system condition made it likely that the last hand had to improvise?

That's where the truth usually appears: tolerances, interfaces, variability, incomplete data, unmanaged changes, packaging, assembly sequence, site coordination.

## The difference between root cause and “reasonable” cause

A reasonable cause sounds good and ends the conversation.

A root cause changes decisions.

Examples of reasonable causes:

    • “lack of training”
    • “lack of attention”
    • “lack of communication”
    • “there was a rush”

I don't say they are false. I say they are cheap: they explain everything and explain nothing.

A root cause, on the other hand, usually takes this form:

    • a poorly defined interface
    • a fuzzy rule
    • a data point that is not a single source of truth
    • a control point that arrives too late
    • an incentive that rewards speed over precision
    • a variability that is accepted as normal.

The root cause isn't always technical. Often it's organizational design.

Short meeting in front of abstract visual panel in workshop, no legible text

## The trade-off no one wants to accept: learning requires stopping “a bit”

Real RCA demands something uncomfortable: a micro-pause.

I'm not talking about stopping the factory. I'm talking about stopping the autopilot.

Imagen 2 — Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

If every incident is resolved with total urgency, there is no space to:

    • look at the pattern
    • compare cases
    • separate coincidence from causality
    • decide what gets standardized.

That's why RCA fails in heroic cultures: if the hero is the one who runs, learning is what gets in the way.

And when the culture rewards running, the system punishes thinking.

## Maturity signals: when RCA is working

I recognize a good RCA system by signals, not by documents:

    • the same defect doesn't return with another name
    • the team talks in terms of system conditions, not culprits
    • subsequent actions are integrated into living standards
    • management protects the minimum time to learn
    • operations and after-sales share a language (not just tickets)
    • it is decided what IS NOT investigated (because not everything deserves the same focus).

That last point is key: maturity is not investigating everything. It's choosing well.

## The most expensive error: investigating “the spectacular” and ignoring “the frequent”

There are large incidents that are impressive.

And there are small incidents that kill.

RCA usually goes for the spectacular because it hurts more in immediate reputation. But margin usually escapes through the frequent: small defects, small reschedules, small urgencies.

I usually ask:

    • what makes you lose more energy in a month?
    • what steals more of your schedule without anyone seeing it?
    • what defect forces you to “explain” more than to “deliver”?

That's where the real focus usually lies.

## Diagnostic questions I use to detect if there is real learning

Without looking for culprits, I look for architecture:

    • What part of the defect could have been detected earlier and wasn't? Why?
    • What process “assumption” broke without anyone knowing?
    • Where is variability created: design, data, manufacturing, logistics, site?
    • What decision is made late, always late?
    • What rule exists “in the head” but not in the system?
    • What change is accepted without impact assessment (because “it's small”)?
    • If the same case is executed by another team, does the result change? Why?

When those questions have no answer, the system lives on memory and luck.

## The underlying problem: root causes usually cross borders

The final defect is seen in one place, but the cause lives in another.

    • The installer “suffers” from a part that left the workshop with fragile tolerances.
    • The workshop “suffers” from an ambiguous product data point coming from design/sales.
    • Purchasing “suffers” from a substitution that changes assembly behavior.
    • After-sales “suffers” from a commercial decision that promised something without operational margin.

If the RCA is done within a silo, it almost always ends in “lack of coordination.”

And that is the elegant way of saying: nobody has ownership of the complete system.

## How a useful RCA looks without becoming bureaucracy

Imagen 3 — Root Cause Analysis (RCA): solve it so it doesn't come back, not just for today

A useful RCA has two visible effects:

    1. reduces recurrence
  1. increases clarity.

The bureaucratic, on the other hand, increases paperwork and maintains recurrence.

I consider there to be bureaucracy when:

    • there are more meetings than changes
    • there is more “follow-up” than learning
    • there is more language than evidence
    • there are more “responsible parties” than real ownership.

And there is learning when the standard changes and is sustained.

## The place where RCA turns into margin

The margin appears when the RCA reaches decisions such as:

    • which interface is not touched (because it maintains compatibility)
    • what variability is allowed and what is not
    • which control point is brought forward (to detect earlier)
    • what information stops being debatable
    • what criterion becomes a “rule” and stops being a debate.

I don't need numbers to know if that is happening. I see it in the atmosphere: fewer urgencies, fewer surprises, more predictability.

Detail of subtle defect in premium finish, no legible text

## What I do NOT recommend (because it creates theater)

    • RCA for every incident, without prioritization.
    • RCA as a search for a culprit (it kills the truth).
    • RCA with generic actions (“training,” “communicate,” “review”).
    • RCA without a clear owner of the subsequent change.
    • RCA without connection to standards (if the system doesn't change, it returns).
    • RCA that depends on one “good” person (if they leave, it disappears).

Quality theater looks a lot like quality… until you look at the recurrence.

## Quick checklist: signals of “firefighting” vs “learning”

Mark mentally:

Firefighting

    • the speed of closing the ticket is celebrated
    • the same defect repeats
    • actions are generic
    • learning doesn't change standards
    • the system depends on heroes.

Learning

    • the reduction of recurrence is celebrated
    • variability is bounded
    • actions change decisions
    • the standard lives and is respected
    • the system works without heroism.

If your organization lives more in the first column, your problem isn't quality: it's system.

Closing of learning with abstract visual filing, no legible text

## Closing

RCA is not a technique. It's a posture: solve so it doesn't come back, not just for today.

When that posture exists, it protects margin, protects lead time and, above all, protects the operational calm that a premium business needs to fulfill its promise.

If you want, I can help you diagnose where recurrence is leaking into your operation, what patterns are draining COPQ without anyone adding them up, and what system decisions you are missing to move from firefighting to learning.

Express Diagnosis

Strategic Audit

Key control points: Root Cause Analysis (RCA)

  • Is there a defined standard for this operation?
  • Do the same dependencies repeat weekly?
  • Does the team know the exact decision criteria?
  • Is there visibility into the real process bottleneck?

Share

Other articles you might like

TPM and OEE without makeup: your factory can be busy… and still losing moneyOPERACIONES

TPM and OEE without makeup: your factory can be busy… and still losing money

When output depends on ‘let’s hope the machine behaves today,’ you don’t have capacity—you have luck. How I diagnose whether the leak is maintenance, changeovers, micro-stops, or quality, and why badly measured OEE misleads you when you need control most.

Psychological safety on the shop floor and on site: the KPI nobody tracksOPERACIONES

Psychological safety on the shop floor and on site: the KPI nobody tracks

How to detect a fear culture (silence, hiding, repeated incidents) and why it hits quality, lead time, and margin harder than many tools.

Robotization and cobots in SMEs: automating without breaking flexibilityOPERACIONES

Robotization and cobots in SMEs: automating without breaking flexibility

Robots are no longer just for automotive giants. Discover how cobots (collaborative robots) can automate repetitive tasks in your furniture SME.

Advanced visualization: renders and VR that sell without creating operational debtOPERACIONES

Advanced visualization: renders and VR that sell without creating operational debt

Renders, VR and configurators are not ‘marketing’: they are a visual promise. I treat them as a decision system, not as polish. Signals, risks and criteria to detect whether visualization protects margin or destroys it.

Omnichannel and phygital experience: when the customer lives a project, not a channelOPERACIONES

Omnichannel and phygital experience: when the customer lives a project, not a channel

Real omnichannel in high-ticket projects: channel drift signals, operational risks, and how I diagnose whether your phygital experience builds trust… or triggers price comparisons.

The technical office as a margin guardian before you sellOPERACIONES

The technical office as a margin guardian before you sell

How to spot (and stop) projects that sell well but are born broken: technical decisions, promises, and variability that destroy margin before execution starts.

Warehouse design and physical flow: margin is lost by walkingOPERACIONES

Warehouse design and physical flow: margin is lost by walking

A warehouse doesn’t ‘get messy’: it’s quietly designed to create extra travel, extra touches, and extra urgency. I diagnose physical flow with simple signals—how many times you touch an item, how far it travels, and where the promise breaks.

Cybersecurity in the connected factory: when downtime starts with a clickOPERACIONES

Cybersecurity in the connected factory: when downtime starts with a click

A connected factory doesn’t fail only because of machines. It fails because of identities, permissions, and inconsistent truths. I don’t sell fear; I care about continuity, margin, and promises that don’t rely on heroics.