Reliability
FMEA worksheet
Work through how an asset can fail, what it would cost, and what would catch it first. Severity, occurrence and detection go in; priority and the gaps come out.
This browser is not saving anything, so the worksheet will be gone when you leave. Private browsing usually causes this.
Rating guide
Agree what a 7 means before anyone types one. These scales are written for maintenance rather than for a production line, and they are ours: if you are working to the AIAG-VDA standard, use its published tables instead.
Severity
How bad the effect is. Rate the effect, not the failure.
- 10Injury or a regulatory breach, with no warning
- 9Injury or a regulatory breach, with some warning
- 8Asset out of service, production stopped
- 7Major loss of function, output severely reduced
- 6Partial loss of function, output degraded
- 5Function impaired, operators work around it
- 4Minor loss of function, operators notice
- 3Minor nuisance, only maintenance notices
- 2Slight effect, no real consequence
- 1No effect
Occurrence
How often this cause actually produces this failure here.
- 10Weekly or more often
- 9About monthly
- 8Every few months
- 7About twice a year
- 6About once a year
- 5Every two years
- 4Every three years
- 3Every five years
- 2Every ten years
- 1Not seen on this asset or its siblings
Detection
Whether your current controls would catch it first. Low is good.
- 10No control at all. Found when it fails
- 9Very unlikely to be caught beforehand
- 8Unlikely. Relies on somebody happening to notice
- 7Low chance, and only on a manual check
- 6Moderate chance on a routine inspection
- 5Caught by a scheduled inspection, sometimes
- 4Usually caught by a scheduled inspection
- 3Likely caught by condition monitoring
- 2Almost always caught by monitoring or an alarm
- 1An automatic control stops it, or an alarm always fires first
Import a worksheet
Paste a file exported from this tool. It replaces what is on screen, so export the current worksheet first if you want to keep it.
| Item and function | Failure mode | Effect | S | Cause | O | Current controls | D | RPN | Priority | Recommended action | Owner | Due | After |
|---|
Getting an FMEA to be worth the afternoon it costs
Most FMEAs fail the same way. A team fills in a large grid, calculates a column of numbers, sorts by it, and then nobody looks at the file again. The grid was the deliverable rather than the decisions.
Rate the effect, not the failure
Severity is the most commonly misread column. It rates what happens as a result of the failure, not how dramatic the failure looks. A seal that weeps into a bund is a low severity failure however unpleasant it is to clean up. The same seal releasing hot water where someone stands is a nine, and the difference is the bund.
Detection means before, not after
A detection rating is a claim about the controls you run today. If the answer is that the operator hears it when the bearing goes, that is a ten, not a three. Rating detection well tends to be the part that changes a PM schedule, because it forces the team to say out loud that a quarterly route would not catch something that develops in a fortnight.
Watch what happens to severity after the action
Occurrence and detection are the two you can usually move. Severity is set by what the failure does, so the only way to change it is to change the design or the situation: a guard, a bund, a relief path, a different place to stand. When a proposed action leaves severity where it was, the worksheet is telling you that the failure can still hurt somebody and you have made it rarer rather than safer.
Finish the rows you started
The review panel below the worksheet counts the rows that scored high with nothing recommended, the actions with no owner, and the detection ratings that credit a control nobody wrote down. Those are the rows that quietly turn a worksheet back into a document.
Frequently asked questions
What is an RPN and why is it criticized?
Risk priority number: severity times occurrence times detection, so 1 to 1000. It is the number most maintenance teams use and the one this worksheet calculates.
The criticism is fair. The three ratings are ordinal, meaning a 10 is worse than a 5 but not twice as bad, so multiplying them does not produce a quantity. Severity 10, occurrence 2, detection 5 and severity 4, occurrence 5, detection 5 both give 100. The first can injure somebody and the second cannot.
So how does this tool prioritize?
Severity first, RPN second. Anything rated severity 9 or 10 is high priority whatever its RPN, and the worksheet prints that reason next to the row instead of leaving you to spot it.
Below that, the RPN threshold you set does the work. Rows at or above it are high, rows over half of it are medium, and severity 7 or 8 pulls a row up to medium on its own. The default threshold is 100 because that is the number most teams already argue about, and it is yours to change.
Are these the AIAG-VDA rating tables?
No. The 2019 AIAG-VDA standard replaced RPN with Action Priority tables, and if you are working to that standard you should use its published tables rather than this.
The scales here are written for maintenance and are ours: they talk about assets stopping, inspections catching things and people getting hurt, rather than defects reaching a customer. They are on the Rating guide panel so your team can agree what a 7 means before anyone types one.
What makes a row worth writing down?
One item, one failure mode, one cause. The most common mistake is a row that bundles three causes under one mode, which makes the occurrence rating meaningless because it is rating whichever cause the person was thinking about.
The review panel flags a few of these. A mode with no cause named is usually several rows. A detection rating of 4 or better with nothing in the controls column is crediting an inspection that may not exist.
Does this save anything or send it anywhere?
It saves to this browser and nowhere else. There is no account and no server call. Export gives you a JSON file you can import again or pass to someone else, and CSV gives you the table for a report.
One worksheet at a time. FMEAs are kept per system and per revision, so export the one you have finished before starting the next.
When is an FMEA the wrong tool?
After a failure has already happened. FMEA is for working out what could fail and deciding what to do about it in advance; once something has broken, a root cause analysis is the tool that fits, because it works backwards from a specific event.
It is also the wrong tool for an asset nobody can describe. If the team cannot name the functions, the worksheet turns into a list of components with guesses beside them.
How often should it be revised?
When something changes: a modification, a new failure nobody predicted, or an action completing and changing the ratings. The revised severity, occurrence and detection columns exist for that last case, so you can show what the work bought rather than asserting it.
A worksheet that has not been touched since commissioning is a document, not a control.
Turn the actions into work that actually runs
The recommended actions from an FMEA are PMs, inspections and condition monitoring. UpKeep holds them against the asset, schedules them, and records whether they were done, so the next revision starts from evidence.