Skip to main content
KreupAI Logo
BLOG GUIDEApplies to: GlobaleAMS
KB-380

FMEA in Practice: Turning a Workshop Output Into an Actual Maintenance Plan

How to convert FMEA failure modes, effects and controls into maintenance tasks, intervals, job plans, spares, ownership and measurable feedback.

Author:Bosco Sabu John
7 min read

FMEA in Practice: Turning a Workshop Output Into an Actual Maintenance Plan

Turn each credible FMEA failure mode into a maintenance decision: prevent it, detect degradation, find a hidden failure, redesign, stock for recovery or accept run-to-failure. For every selected action define the asset scope, trigger or interval, method, acceptance limit, labour, tools, parts, safety controls, responsible role and evidence. Load approved tasks into the EAM, then use work history and failures to test assumptions and revise the FMEA.

The workshop is finished only when a maintenance decision exists

An FMEA row with a component, failure mode, effect, cause and score is analysis—not yet a maintenance plan. The team must decide whether a technically feasible task can prevent the cause, detect degradation with useful warning, reveal a hidden protective failure or reduce the consequence. If no suitable task exists, the answer may be redesign, spare strategy, contingency or explicit run-to-failure.

Do not convert every high risk-priority number into a monthly inspection. Ranking helps focus discussion, but task choice depends on failure behaviour, consequence, detectability, degradation interval, operating context and whether the proposed work actually changes risk.

The nine statuses below are a practical canonical lifecycle. Map the native codes in your CMMS or EAM to them for cross-site reporting; do not force every site to rename its system.

Translate each failure mode through a task-selection path

StatusWhat it meansMinimum evidence before moving onOwner of the next move
1. DraftA record exists, but has not entered the controlled queueAsset or location, short description, sourceRequester or planner
2. RequestedThe need is submitted and awaiting triageReported time, requester, symptom, operational impactSupervisor or helpdesk
3. ApprovedThe work is valid, in scope and assigned a priorityPriority rationale, work type, responsibility, target datePlanner
4. PlannedThe job is ready in method and resourcesJob steps, craft, estimated hours, parts, permits, isolationsScheduler
5. ScheduledA crew and execution window are committedScheduled start, assigned labour, material availabilitySupervisor or technician
6. In ProgressPhysical work has startedActual start, attending technician, initial conditionTechnician
7. On HoldWork has stopped for a named constraintHold reason, responsible party, next review dateNamed constraint owner
8. CompletedPhysical work is finished and technical history submittedActual labour, parts, readings, action, failure data, downtimeSupervisor or reliability reviewer
9. ClosedTechnical and commercial checks are complete; the record is historyCompletion review, follow-on work, cost and reservation reconciliationAuthorised closer

Cancelled is a terminal outcome, not a stage. It should be reachable only before physical work starts, require a reason code, and remain reportable. Rejected and duplicate requests are also outcomes of triage. Treating them as ordinary lifecycle stages distorts cycle-time reporting because they never had an execution path.

1–3: Draft, Requested and Approved

Draft is useful where a mobile user may save an incomplete record or an integration creates a shell before enrichment. It should not count in backlog, SLA response or demand reporting. Requested is the first controlled state: the clock starts, the record cannot quietly disappear, and triage has an owner.

Approval is not permission to start an unprepared job. It means the request is genuine, in scope and prioritised against a defined consequence matrix. The approval gate should resolve duplicates, warranty responsibility, landlord-versus-tenant scope and whether the record is actually maintenance. A request for a new installation may be a project; a comfort complaint may be an operating adjustment.

Do not ask requesters to diagnose failure modes. They can reliably describe symptoms: leaking, noisy, stopped, hot, intermittent. Diagnosis belongs after inspection. Forcing a failure code at request time produces confident fiction that survives into history.

4–5: Planned and Scheduled are different

Planned means ready to execute: the scope is understood, method is safe, skills and duration are estimated, parts are identified, and access, permits and isolations are known. Scheduled means that ready work has been placed into a dated labour window.

Combining the two hides the constraint. A job can be fully planned but unscheduled because capacity is unavailable. It can also appear on next Tuesday's schedule while lacking a bearing, permit or isolation plan. Calling both states “Assigned” makes schedule compliance impossible to interpret.

Material availability belongs as a condition or hold reason rather than an uncontrolled free-text note. IBM Maximo makes the distinction explicit with Waiting for Materials and Waiting for Plant Conditions. Whether these are separate statuses or reason codes under On Hold depends on reporting needs, but the cause must be structured.

6: In Progress — the first place data quality dies

The transition to In Progress should be triggered by arrival at the asset or the first physical task, not by dispatch and not by a supervisor bulk-starting the morning's list. It establishes the actual start time and identifies who attended.

This is where the first data loss occurs because consumption happens faster than administration. A technician takes another filter from the van, receives help from an electrician for forty minutes, changes the repair method and observes that the unit was already unavailable. If labour, materials, readings and downtime are deferred until the end of the shift, detail is reconstructed from memory or omitted.

The answer is not a longer completion form. Capture actuals at the point they occur:

  • start, pause and resume timestamps from the mobile job;
  • parts issued or scanned against the work order, including van stock;
  • meter readings with units and acceptable ranges;
  • photos tied to a defined step, not dumped into a general attachment list;
  • additional labour booked by the person who supplied it;
  • follow-on defects raised as linked work, rather than buried in notes.

A work order left In Progress overnight must mean one of two things: work genuinely continues, or the record is wrong. If access, materials, production or a permit blocks the job, move it to On Hold and name the constraint.

7: On Hold needs a clock and an owner

On Hold is usually where overdue work goes to become invisible. A useful hold state requires three fields: a reason code, the person or team responsible for clearing it, and a review date. “Pending” and “Other” are not reason codes.

Separate clock treatment from operational truth. An SLA may legitimately pause for denied access or customer authorisation while continuing for a missing internal spare. That is a contractual calculation based on the hold reason; it should not change the factual timestamps. Never overwrite elapsed time to make compliance look better.

Report hold ageing by reason and owner. A queue dominated by Waiting for Material is a planning or supply problem. One dominated by Access Required is a scheduling and stakeholder problem. A single On Hold count diagnoses neither.

8: Completed — the second place data quality dies

Completed should mean the physical task is finished and the technical record is ready for review. It is not the same as Closed. IBM describes Completed as the point at which physical work is finished, while Closed finalises the order, removes unused inventory reservations and turns it into history. SAP similarly distinguishes technical completion from business completion.

The second—and larger—data loss happens here. The technician is under pressure to reach the next job, so the close-out becomes “checked and fixed”. That sentence cannot support reliability analysis, warranty recovery, repeat-failure investigation, spares forecasting or maintenance strategy.

For corrective work, require structured answers to four different questions:

  1. What was observed? The symptom or condition as found.
  2. What failed? The maintainable item or component.
  3. How did it fail? The failure mode: leak, seizure, open circuit, drift, blockage.
  4. Why did it fail? The cause, where evidence supports one. “Unknown” is more honest than a guessed cause.

Also capture action taken, actual labour and materials, service-restored time, downtime, readings after work, and whether follow-on work is required. ISO 14224:2016 provides the underlying discipline: equipment, failure and maintenance data, including cause, consequence, action, resources used and downtime, in a standardised format. It is industry-specific, but the data-quality principle travels well.

Do not make every field mandatory for every job. A statutory inspection, lubrication task and breakdown repair need different completion schemas. Conditional forms by work type improve completeness and reduce meaningless values entered only to satisfy validation.

9: Closed is a review gate

Closure is the point of no casual editing. An authorised reviewer checks that the correct asset was used, completion codes are coherent, follow-on defects have linked orders, unused reservations are released, service entries or contractor costs are reconciled, and any required customer acceptance is attached.

This should be a risk-based control. Automatically close low-risk routine work that passes validation. Sample completed work by technician and work type. Route safety-critical, statutory, high-cost and repeat-failure jobs for explicit review. Making a supervisor manually close every filter change creates a rubber stamp; reviewing none exports bad data directly into the asset history.

If history must be corrected later, keep an audit trail showing the original value, revised value, reason, person and timestamp. Silent edits destroy trust in every reliability metric downstream.

The controls that make the lifecycle auditable

TransitionControl that matters
Draft → RequestedRequired identification and reported timestamp
Requested → ApprovedScope, duplicate and priority check
Approved → PlannedJob plan and resource readiness
Planned → ScheduledCommitted labour window and constraint check
Scheduled → In ProgressActual start by the attending worker
In Progress → On HoldStructured constraint, owner and review date
In Progress → CompletedWork-type-specific technical close-out
Completed → ClosedRisk-based quality and cost review

Measure transition quality, not just the number of jobs in each status. Useful controls include the share of jobs started before approval, scheduled without a complete plan, completed with no actual labour or parts, closed with free-text-only failure data, held past their review date, and reopened within 30 days.

Cycle time should be split by state. One average from request to close cannot tell whether delay sits in approval, planning, scheduling, execution, constraint clearance or review.

FAQ

Must a work order have exactly nine statuses? No. Nine is a useful canonical reporting model, not a standard. A simple operation may combine Draft with Requested or Planned with Scheduled. Add a status only when it represents a real decision, changes ownership or controls what can happen next.

Should Waiting for Parts be its own status? Use a distinct status where material delay is operationally significant and someone manages that queue. Otherwise use On Hold with a mandatory reason. The important point is structured, owned and aged constraint data.

When should the SLA stop? According to the contract and the hold reason. Preserve factual timestamps, then calculate paused and unpaused duration separately. Do not let users manually change elapsed time.

Who should close the work order? The technician should submit technical completion. Closure can be automated for validated low-risk work and assigned to a supervisor, planner or contract administrator for exceptions and higher-risk work.

Convert the worksheet into a task register

Create one implementation record for every approved maintenance response. The minimum fields are the FMEA revision, asset class and applicability, failure mode, selected policy, task description, maintenance type, interval or trigger, acceptance criterion, craft, estimated duration, tools, parts, permits, isolation, safety precautions and technical owner. A score alone is not enough to generate work.

Write the task so a competent technician can distinguish normal from abnormal. “Inspect bearing” is incomplete. State the method, measurement point, operating state, unit, limit and action when the limit is exceeded. If judgement is necessary, specify the competence required and capture a structured finding with supporting notes or evidence.

Separate the generic task from the asset-specific schedule. One approved motor-bearing inspection template may apply to many assets, but each application needs a start date, interval, operating context and responsibility. This prevents hundreds of copied job plans from diverging while still allowing justified local variation.

Choose intervals from failure behaviour

The workshop often selects monthly or annual intervals because those periods fit the calendar. A defensible interval comes from the reason the task works. For condition monitoring, estimate the interval between detectable degradation and functional failure, then inspect often enough to leave time for confirmation, planning, material and repair. Record uncertainty and begin conservatively where evidence is weak.

For age-related restoration or replacement, use service history, manufacturer evidence and operating severity rather than an arbitrary age. For hidden protective functions, set failure-finding frequency from the required availability and consequence. Usage-based triggers may be better than calendar time where starts, cycles, distance, throughput or operating hours drive degradation.

Intervals are controlled assumptions. Define who may change them, what evidence is required and how the revised FMEA and maintenance programme remain linked. Repeated “no defect found” inspections may justify review, but they do not prove the task is unnecessary if the protected consequence is severe or the sample is too small.

Design the feedback loop before release

The completed work order must capture evidence that can test the FMEA. Use consistent failure-mode and cause codes, as-found condition, measurement, action, parts, downtime and functional impact. Preserve free text for unexpected learning, but do not rely on narrative alone for analysis.

Review leading and lagging signals together: task compliance, overdue exposure, defects detected, warning time, repeat work, emergency failures, false alarms, intrusive-maintenance defects, parts consumption and operational consequence. A task can be completed on time and still be ineffective because it detects too late, measures the wrong point or creates more risk than it controls.

Set review triggers rather than waiting for an annual meeting. Reopen the analysis after a significant failure, repeat defect, design modification, change in duty, new consequence, obsolete part, altered operating environment or credible new technical evidence. Close the loop by versioning the FMEA, approving the maintenance change and updating affected schedules.

Govern implementation as a change

Before activation, check workforce capacity, access windows, spare availability, tools, permits and training. Adding hundreds of tasks without removing weak work increases backlog and can displace higher-value maintenance. Pilot new tasks on a representative asset set, examine whether instructions and findings are usable, and adjust before fleet deployment.

Assign one owner for technical validity and another for programme execution. Reliability or engineering may own the failure logic; maintenance planning owns job-plan quality and scheduling; supervisors assure execution; technicians provide field evidence; operations owns access and operating consequences. The EAM should show these handoffs instead of burying ownership in meeting minutes.

Where a system helps

Good workflow makes the right action easier at the moment it matters: mobile actuals during execution, a completion form that changes by work type, structured hold ownership, and a review queue driven by risk rather than volume. The result is not merely a cleaner dashboard. It is asset history reliable enough to change a PPM interval, challenge a repeat contractor failure or defend a replacement decision. See eAMS for asset management.

Related reading: What Is a CMMS? (KB-129) and What Is Planned Preventive Maintenance? (KB-130).

Sources