GeoBusinessIQGeoBusinessIQ

Reliability-centred maintenance: choosing a policy for each way a machine fails

What this answers

For each failure mode on this asset, is there a maintenance task that is both technically effective and worth its cost?

Reliability-centred maintenance asks a narrow question repeatedly: for each way this equipment can fail, what are the consequences, and is there a task worth doing to prevent or detect it. The answer is sometimes a scheduled task, sometimes monitoring, sometimes a design change, and quite often a deliberate decision to run to failure. Its discipline is valuable and its cost is real, so where it is applied matters as much as how.

Written for: reliability engineers, maintenance managers, plant engineering teams.

Consequence decides the policy, not criticality of the asset

The analysis sorts failure modes by what happens when they occur: harm to people, breach of an environmental or regulatory limit, loss of production, or simply repair cost. Safety and compliance consequences demand a task or a design change regardless of economics. Production and cost consequences are judged commercially. This distinction is what stops a plant applying uniform intensive maintenance to a whole machine when only two of its failure modes actually matter, and it explains why two identical machines in different duties legitimately end up with different maintenance policies.

The room matters more than the worksheet

A useful analysis needs the technician who repairs the asset, the operator who runs it, someone who knows the process consequences and a facilitator who keeps the group answering the question rather than reminiscing. Missing the operator is the most common error, because operators know the failure modes that get worked around informally and never appear in any record. Sessions run by an engineer alone with a manual produce a plausible document and miss the real failure history. Budget the time properly; a rushed analysis produces a task list nobody believes and the effort is wasted twice.

Deciding where the full method is worth its cost

A rigorous analysis is expensive in skilled people's time, so apply it where the return justifies it: the assets that limit output, those with safety or environmental consequences, those with a history of expensive repeated failure, and new installations where no history exists. Elsewhere, a lighter review of the existing routines against actual failure history captures much of the benefit for a fraction of the effort. Plants that attempt to analyse everything typically stall part way through, leaving an inconsistent programme where some assets have been rigorously examined and the rest carry inherited routines.

Running to failure is a legitimate answer

Where a failure mode has no safety or environmental consequence, gives no useful warning, and costs little more to fix after the event than before it, the correct decision is to let it fail and repair it. Recording that decision explicitly is important. It prevents the item being quietly added back into a routine, it justifies stocking the relevant spare so the repair is quick, and it makes clear that the failure when it happens is an accepted outcome rather than a maintenance lapse. Unrecorded, the same conversation repeats every time the component fails.

Keeping the output alive after the project ends

The characteristic failure is that an analysis produces a revised task list, it is loaded into the maintenance schedule, and then nothing revisits it as the plant changes duty, modifies the equipment or accumulates new failure evidence. Set a review trigger rather than a calendar habit: a new failure mode not in the analysis, a modification, a change in operating conditions, or a run of repeat failures. Keep the reasoning accessible, not just the resulting tasks, because a task list without its justification is indistinguishable from any other inherited routine within a couple of years.

Frequently asked questions

Is this approach only for large or heavily regulated plants?
The full formal method suits complex assets and regulated environments where the reasoning must be defensible. Smaller plants get most of the benefit from a simplified version: list how the machine has actually failed, ask what each failure costs, and decide for each whether a task would prevent or detect it in time. That conversation between a technician, an operator and a manager takes an afternoon and usually improves an inherited routine substantially.
How does this relate to total productive maintenance?
They address different questions and can coexist. One decides which maintenance policy each failure mode deserves; the other is a shop-floor improvement approach concerned with operator ownership, equipment condition and eliminating losses. A plant can use the analysis to set its technical maintenance policy while running operator-based care and improvement activity alongside. Problems arise only when both are run as competing programmes by different departments with separate reporting.
What if we have no failure history to analyse?
Use the equipment manufacturer's failure mode information, the experience of technicians who have worked on similar machines, and any evidence from sister sites, then treat the resulting policy as provisional. Start recording failures properly from installation, including what was found rather than only what was replaced. New assets often need a review after their first period of operation, since early failures tend to reflect installation and commissioning issues rather than the long-run pattern.

Data limitations

  • Manufacturing figures are operator-supplied inputs, not market data. GeoBusinessIQ holds no factory costs, production volumes, yields, cycle times, tooling prices or capacity data and does not estimate them — every result reflects only the figures you enter.

Explore the graph

Sources

  • United Nations Industrial Development Organization UNIDO (accessed )
    Covers: Industrial development analysis, industrial statistics methodology, and manufacturing capability programmes across member states.
    Does not cover: Company-level data, factory costs, supplier information, or real-time production statistics.
    Why it matters: The United Nations agency for industrial development; used for structural framing of how manufacturing sectors develop, never for point figures.
    Review cadence: annual
  • National Institute of Standards and Technology NIST (accessed )
    Covers: Measurement science, manufacturing technology research, cybersecurity frameworks, and industrial standards support.
    Does not cover: Certification of products, endorsement of vendors, or costs for any specific implementation.
    Why it matters: A United States federal research institute whose public material covers measurement, manufacturing technology and control-system security.
    Review cadence: annual
  • European Agency for Safety and Health at Work EU-OSHA (accessed )
    Covers: Information on European Union occupational safety and health legislation and workplace risk management practice.
    Does not cover: National implementation detail, workplace-specific risk assessments, or enforcement decisions.
    Why it matters: Cited for the European framework on worker and machinery safety in manufacturing settings.
    Review cadence: annual

Educational and operational information only — not legal, engineering, safety, customs, tax, or financial advice. Requirements vary by jurisdiction, product, process, and contract; confirm with the relevant authority or a qualified professional before acting.

Last updated: