Reliability

FMEA worksheet

Work through how an asset can fail, what it would cost, and what would catch it first. Severity, occurrence and detection go in; priority and the gaps come out.

Getting an FMEA to be worth the afternoon it costs

Most FMEAs fail the same way. A team fills in a large grid, calculates a column of numbers, sorts by it, and then nobody looks at the file again. The grid was the deliverable rather than the decisions.

Rate the effect, not the failure

Severity is the most commonly misread column. It rates what happens as a result of the failure, not how dramatic the failure looks. A seal that weeps into a bund is a low severity failure however unpleasant it is to clean up. The same seal releasing hot water where someone stands is a nine, and the difference is the bund.

Detection means before, not after

A detection rating is a claim about the controls you run today. If the answer is that the operator hears it when the bearing goes, that is a ten, not a three. Rating detection well tends to be the part that changes a PM schedule, because it forces the team to say out loud that a quarterly route would not catch something that develops in a fortnight.

Watch what happens to severity after the action

Occurrence and detection are the two you can usually move. Severity is set by what the failure does, so the only way to change it is to change the design or the situation: a guard, a bund, a relief path, a different place to stand. When a proposed action leaves severity where it was, the worksheet is telling you that the failure can still hurt somebody and you have made it rarer rather than safer.

Finish the rows you started

The review panel below the worksheet counts the rows that scored high with nothing recommended, the actions with no owner, and the detection ratings that credit a control nobody wrote down. Those are the rows that quietly turn a worksheet back into a document.

Frequently asked questions

What is an RPN and why is it criticized?

Risk priority number: severity times occurrence times detection, so 1 to 1000. It is the number most maintenance teams use and the one this worksheet calculates.

The criticism is fair. The three ratings are ordinal, meaning a 10 is worse than a 5 but not twice as bad, so multiplying them does not produce a quantity. Severity 10, occurrence 2, detection 5 and severity 4, occurrence 5, detection 5 both give 100. The first can injure somebody and the second cannot.

So how does this tool prioritize?

Severity first, RPN second. Anything rated severity 9 or 10 is high priority whatever its RPN, and the worksheet prints that reason next to the row instead of leaving you to spot it.

Below that, the RPN threshold you set does the work. Rows at or above it are high, rows over half of it are medium, and severity 7 or 8 pulls a row up to medium on its own. The default threshold is 100 because that is the number most teams already argue about, and it is yours to change.

Are these the AIAG-VDA rating tables?

No. The 2019 AIAG-VDA standard replaced RPN with Action Priority tables, and if you are working to that standard you should use its published tables rather than this.

The scales here are written for maintenance and are ours: they talk about assets stopping, inspections catching things and people getting hurt, rather than defects reaching a customer. They are on the Rating guide panel so your team can agree what a 7 means before anyone types one.

What makes a row worth writing down?

One item, one failure mode, one cause. The most common mistake is a row that bundles three causes under one mode, which makes the occurrence rating meaningless because it is rating whichever cause the person was thinking about.

The review panel flags a few of these. A mode with no cause named is usually several rows. A detection rating of 4 or better with nothing in the controls column is crediting an inspection that may not exist.

Does this save anything or send it anywhere?

It saves to this browser and nowhere else. There is no account and no server call. Export gives you a JSON file you can import again or pass to someone else, and CSV gives you the table for a report.

One worksheet at a time. FMEAs are kept per system and per revision, so export the one you have finished before starting the next.

When is an FMEA the wrong tool?

After a failure has already happened. FMEA is for working out what could fail and deciding what to do about it in advance; once something has broken, a root cause analysis is the tool that fits, because it works backwards from a specific event.

It is also the wrong tool for an asset nobody can describe. If the team cannot name the functions, the worksheet turns into a list of components with guesses beside them.

How often should it be revised?

When something changes: a modification, a new failure nobody predicted, or an action completing and changing the ratings. The revised severity, occurrence and detection columns exist for that last case, so you can show what the work bought rather than asserting it.

A worksheet that has not been touched since commissioning is a document, not a control.

Turn the actions into work that actually runs

The recommended actions from an FMEA are PMs, inspections and condition monitoring. UpKeep holds them against the asset, schedules them, and records whether they were done, so the next revision starts from evidence.