← Back to Blog

Predictive Maintenance: When Is the Right Time to Stop a Healthy Machine?

Every maintenance decision trades visible downtime today against uncertain failure tomorrow. The real opportunity is learning when your equipment—not the calendar—is asking for intervention.

The hardest machine to stop is the one that appears to be working perfectly.

Pulling it from production creates an immediate, measurable loss: labor waits, output pauses, and a maintenance bill arrives. The failure you are trying to prevent is invisible. If nothing breaks afterward, the intervention can even look unnecessary.

That is the maintenance trade-off no one applauds: pay a certain cost now, or accept an uncertain—and potentially much larger—cost later.

Operations managers make this decision under pressure. They are asked to protect throughput today and equipment health tomorrow, even though those objectives can point in opposite directions. Run too long and a small degradation can become a line event. Intervene too often and healthy component life, labor, spares, and production windows are consumed for little benefit.

Predictive maintenance is not simply “more maintenance.” It is an attempt to make that timing decision with better evidence.

The goose, the egg, and production capability

Stephen R. Covey uses Aesop’s goose-and-golden-egg fable in The 7 Habits of Highly Effective People to distinguish production from the capability that produces it. A farmer, impatient for more golden eggs, destroys the goose—and loses both future output and the asset that created it.

A factory can make the same mistake without touching an axe. Every additional hour of output is another egg. The equipment, trained maintenance team, process knowledge, and spare-parts system are the goose. Maximizing today’s production while allowing that capability to deteriorate is not productivity; it is borrowing output from the future.

The opposite extreme also fails. Constantly disturbing a healthy machine can introduce defects, consume component life prematurely, and use the very capacity maintenance was meant to protect. The goal is balance: protect production and production capability.

Four maintenance questions—not one ladder

Maintenance strategies are often presented as a maturity ladder, but they answer different questions:

Approach Trigger Question it answers
Reactive / corrective Failure What must we fix now?
Preventive Calendar or usage threshold What does experience say we should service routinely?
Condition-based Measured condition crosses a limit Is the asset showing evidence of degradation?
Predictive Condition and history estimate future risk When is failure risk likely to become unacceptable?

The equipment manual is prior knowledge. It captures design intent, operating limits, known wear mechanisms, inspection methods, and recommended intervals. That makes it an excellent preventive-maintenance baseline—not a generic document to ignore once sensors arrive.

Your operating data supplies the local evidence. It tells you whether the manual’s assumptions fit your utilization, recipes, environment, product mix, and history. Two nominally identical machines can therefore deserve different maintenance dates.

The progression is better understood as:

OEM knowledge + maintenance history + operating conditions + condition signals

                  a tool-specific decision

Why “just follow the interval” leaves value behind

Time- and usage-based maintenance are sensible starting points.

  • Time-based: intervene after a fixed duration—monthly, quarterly, or annually.
  • Usage-based: intervene after cycles, wafers, batches, operating hours, or another exposure measure.
  • Condition-based: intervene when vibration, temperature, pressure, contamination, deposition, current, or quality behavior crosses a meaningful limit.
  • Predictive: combine those signals and their trends to estimate the remaining useful window and the consequence of waiting.

A manual might recommend service every 1,000 operating hours. But 1,000 hours at moderate load in a clean, stable environment is not necessarily equivalent to 1,000 hours at high temperature, aggressive chemistry, frequent changeovers, or repeated excursions.

The calendar sees elapsed time. The machine experiences stress.

This is the gap predictive maintenance tries to close. NIST describes the core distinction similarly: preventive maintenance responds to time or routine readings, while predictive maintenance uses actual condition to anticipate failure and schedule service before a critical event. The U.S. Department of Energy likewise defines predictive maintenance around detecting the onset of degradation before significant deterioration. (NIST, DOE O&M guide)

The cost is larger than the repair invoice

Maintenance cost is frequently discussed as though it were only technician time and replacement parts. A review of maintenance-cost research found wide variation in reported estimates and in what organizations include in those estimates—one reason simplistic percentages should be treated as context, not universal benchmarks. (Ran et al., 2023)

The consequence of a failure can include:

  • lost production during diagnosis, repair, qualification, and ramp-up;
  • scrap, rework, yield loss, or latent quality risk;
  • expedited parts and contractor premiums;
  • starved downstream tools and blocked upstream work-in-process;
  • missed delivery commitments;
  • secondary damage to connected components; and
  • safety or environmental exposure.

That last operational chain is the duplicative effect of neglected maintenance: one degraded component changes the conditions experienced by everything that depends on it. The failed part may be cheap. The system response may not be.

Over-maintenance has its own cost:

  • labor spent on work that did not need to happen yet;
  • premature replacement of healthy components;
  • avoidable planned downtime;
  • maintenance-induced failures from incorrect assembly, contamination, or poor calibration; and
  • excess spare-parts inventory held for poorly understood “what if” scenarios.

Spares are not free insurance. They tie up capital, occupy controlled space, age, and can become obsolete. Better forecasts do not eliminate spares; they help distinguish strategically necessary coverage from fear-driven inventory.

Put a number on the argument

The simplest defensible comparison is expected cost:

Planned intervention cost = PM work + planned downtime

Expected cost of waiting = probability of failure
                         × (corrective work + unplanned downtime)

It is incomplete—quality, safety, cascading damage, and recovery uncertainty may need explicit terms—but it exposes the decision hiding behind “the tool seems fine.” It also reveals the break-even probability: how likely failure must be before planned intervention has the lower expected cost.

Try a tool you know. If you do not know the failure probability, that discomfort is the point.

A five-minute maintenance decision

What does “keep it running” actually cost?

Compare one planned intervention with the expected exposure of waiting through your next operating window.

Planned intervention
Expected cost of waiting
Break-even failure probability

This is an expected-cost screen, not a failure prediction. Safety, quality, cascading damage, and customer impact can make the real consequence much larger.

The product takes the next step: replace the probability you guessed with evidence from tool history, condition signals, utilization, and the production schedule.See what is being built →

The calculator asks you to estimate the number that matters most. A predictive-maintenance product should help you earn that number from evidence.

What this looks like in semiconductor manufacturing

Semiconductor fabs make the timing problem unusually visible. Equipment from OEMs such as Applied Materials, Lam Research, Tokyo Electron, ASML, KLA, ASM, and Hitachi High-Tech supports deposition, etch, lithography, inspection, metrology, clean, implant, thermal, and other process steps. Architectures range from batch furnaces to single-wafer platforms with multiple chambers.

Each tool has its own maintenance logic. A platform-level event may stop every chamber; a chamber-level event may reduce capacity while the rest of the tool continues. Some PMs are calendar-based, others depend on RF hours, wafer count, film thickness deposited, chemical exposure, or clean cycles. Product mix and recipe severity can make raw wafer count a misleading proxy for wear.

Now add the factory context. Pulling a non-constraint tool during a low-load window may be inexpensive. Pulling the current bottleneck can increase queue time and WIP across the line. Waiting for failure on that same bottleneck may be far worse. Maintenance timing is therefore not only an equipment-health problem; it is a line-management decision.

A useful system must connect both sides:

  1. Asset risk: What is changing, how quickly, and with what confidence?
  2. Operational consequence: If this asset stops, what happens to throughput, WIP, yield, delivery, and cost?

Without the first, scheduling is guesswork. Without the second, every alert looks equally urgent.

What good prediction should unlock

The goal is not a dashboard full of red indicators. It is a smaller set of better decisions:

  • pull a PM forward when degradation accelerates;
  • safely push it out when condition remains stable;
  • stagger maintenance across equivalent tools;
  • protect bottleneck capacity and plan recovery windows;
  • forecast technicians, kits, and critical spares;
  • compare risk consistently across tools; and
  • learn whether the intervention actually restored performance.

Those decisions can protect OEE, throughput, yield, quality, and equipment life—but the metrics should not be promised automatically. They improve only when prediction leads to a timely, correct, and well-executed action. The DOE’s guidance stresses that predictive technologies still require system knowledge, training, monitoring, and proper implementation. Sensors do not replace maintenance discipline; they make disciplined maintenance more precisely timed.

The product question

Most teams already possess fragments of the answer: OEM intervals in manuals, work orders in a CMMS, alarms in equipment logs, process traces in historians, yield in quality systems, and production demand in planning tools. The problem is that the decision is made between those systems.

The product I am building is aimed at that gap. Imagine opening one tool record and seeing:

  • the OEM baseline and the tool’s actual exposure;
  • its condition trend and maintenance history;
  • an explainable risk window—not merely a red light;
  • the cost and line impact of acting now versus waiting;
  • candidate maintenance windows across the schedule; and
  • the evidence behind every recommendation.

The calculator above gives you the economics with a probability you supply. The product should help estimate that probability, show what is driving it, and turn it into a maintenance window an operations team can defend.

Visit the product section to see where this is going →

Predictive maintenance does not ask, “Can this machine run one more hour?” It asks a better question: Is the value of that hour greater than the risk we are adding—and do we have enough evidence to know?