This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Quality Risk Management — How Much Evidence Is Enough

ICH Q9(R1) as a loop, not a form: the risk-management toolbox (FMEA, FTA, HACCP, HAZOP, risk ranking and filtering, Ishikawa/PHA), FMEA in action, and how a risk assessment becomes a control strategy — worked through the nitrosamine risk assessments.
A banner titled 'Quality Risk Management — How Much Evidence Is Enough?' with the tagline 'Focus Effort. Control What Matters. Protect Patients.' and the note that ICH Q9(R1) is a science- and risk-based approach to identify, evaluate, control, and review risks across the product lifecycle. Panels: (1) The One Idea — the two governing principles: risk evaluation is grounded in scientific knowledge and links to patient protection; effort, formality, and documentation are proportionate to risk, with a note that risk management focuses budget on failures that would actually hurt a patient without gold-plating the rest; (2) The ICH Q9(R1) Process — a continuous four-stage loop around 'Quality Risk Management': Risk Assessment (identify, analyze, evaluate — what could go wrong, how likely, how severe, is it acceptable), Risk Control (reduce or accept — change the design, implement controls, accept residual risk), Risk Communication (share and discuss with all stakeholders), Risk Review (monitor and revise when new information emerges) — captioned as a living process across the product lifecycle, not a form to be filed; (3) From Risk Assessment to Control Strategy — a five-step vertical flow (understand the product and process per Q8/Q9/Q10; identify and assess risks — CQAs, method parameters, process steps; implement controls — method design, robustness, specifications, monitoring; control strategy — the planned set of controls that assures quality and performance; lifecycle management — review and adapt per Q12/Q14) paired with its outputs (CQAs and CPPs, analytical method controls, specifications under Q6, the stability program under Q1, monitoring and continued verification, regulatory submissions) and the quote that risk management turns knowledge into decisions, and decisions into patient protection; (4) Hazard vs. Risk — a hazard is the potential to cause harm ('this solvent is toxic'), risk is the probability of that harm and its severity ('at the residual level detected by this method, exposure is under 1% of the PDE'), with the reminder that formality is a dial, not a switch — one page or a full FMEA can both be appropriate if justified by the risk; (5) Risk-Management Toolbox — a table of six tools: FMEA/FMECA (failures of a process or method built from many steps, the workhorse in analytical development), fault tree analysis/FTA (working backward from one defined failure, good for OOS root-cause work), HACCP (identifying and controlling critical points, origin in food safety, useful in manufacturing), HAZOP (deviations from design intent, a guided-word approach common in process/engineering), risk ranking and filtering (comparing many risks across a portfolio, useful for site- and portfolio-level decisions), Ishikawa (fishbone)/PHA (structuring a first-pass hazard identification, often the front end of an FMEA); (6) FMEA in Action, an analytical-method example — a table walking five steps (sample preparation, chromatography, detection, data analysis, system suitability) each with a failure mode, its effect on the patient/decision, Severity, Occurrence, Detection, the resulting RPN, and a corrective action, with the reminder that RPN = Severity × Occurrence × Detection, a high Detection score means poorly detected (the scale runs backward), and every action gets a re-score to confirm risk reduction; (7) Worked Example — Nitrosamine Risk Assessment, a real-world QRM case walking six numbered steps (identify the hazard — potent mutagenic carcinogens present in multiple products, triggered by the valsartan recalls; analyze the risk — synthetic route, nitrite sources, secondary amines, recovered solvents, water, confirmatory testing; evaluate the risk against acceptable-intake thresholds; control the risk — route changes, nitrite scavengers, tighter ppb-level specifications; communicate — share with regulators, document decisions, meet deadlines like EMA Article 5(3); review — ongoing monitoring and periodic reassessment as new information emerges), beside a chemical structure of NDMA (N-nitrosodimethylamine) and a chromatogram showing sensitive LC–MS/MS detection at nanogram levels, with the note that the ability to detect a nitrosamine at its acceptable intake was a fundamental part of the risk conclusion — an analytical issue; (8) A Typical Risk Matrix — a 5×5 severity-by-occurrence grid color-coded from green (low risk, acceptable) through yellow (medium risk, consider action) to red (high risk, needs action); (9) Key Takeaways — use scientific knowledge to focus effort on what matters for the patient; match the level of effort and documentation to the level of risk; risk management is a loop — assess, control, communicate, and review; FMEA is powerful but has limitations, don't over-trust the number; a control strategy is the output of risk management, not a separate exercise; (10) Connection to Other ICH Guidelines — a six-box chain: Q8 Development (build knowledge and design quality into the product), Q9 Risk Management (identify, evaluate, control, and review risks), Q10 Lifecycle (a lifecycle approach to quality), Q12 Change Management (manage changes based on risk), Q14 Analytical Development (method design, MODR, and control strategy), Q6/Q1 Specifications & Stability (numeric risk decisions); (11) For Discussion — five questions: where to draw the line between a hazard and a risk in analytical work; how to determine the right level of formality for a given decision; which failure modes in the FMEA example concern you most and why; how the nitrosamine case illustrates the importance of analytical detection; how a risk assessment translates into method design and specification. Footer: 'Science + Risk-Based Decisions = Better Medicines for Patients,' Temple University branding, and the tagline 'All science ultimately serves people.'

The one idea

Two principles govern quality risk management: the evaluation of risk is grounded in scientific knowledge and ultimately links to protection of the patient; and the level of effort, formality, and documentation is proportionate to the level of risk.

Every analytical decision spends a finite budget of time, money, and attention. Risk management is how you point that budget at the failures that would actually hurt a patient, and stop gold-plating the ones that wouldn’t. It is the machinery behind “scientifically justified” — the phrase that appears in almost every ICH guideline and is doing a lot of quiet work.

The ICH Q9 framework

ICH Q9(R1) — Quality Risk Management (the R1 revision, adopted 2023, added guidance on subjectivity, the hazard-versus-risk distinction, formality, and risk-based decision-making). The process is a loop, not a form:

StageWhat happensAnalytical example
Risk assessment — identificationWhat could go wrong?A co-eluting degradant is not resolved from the API
Risk assessment — analysisHow likely, how severe, how detectable?Estimate occurrence from forced-degradation data; severity from the degradant’s qualification threshold; detection from method specificity
Risk assessment — evaluationIs that acceptable against defined criteria?Compare against a risk threshold agreed before the assessment
Risk control — reductionChange the design to lower likelihood or raise detectionSwitch to an orthogonal column; add a peak-purity check
Risk control — acceptanceSome residual risk is accepted, explicitly and on the recordDocument the residual and the justification
Risk communicationThe assessment and decisions are shared with everyone who acts on themThe control strategy, the filing, the SOP
Risk reviewRevisit when something changesA new impurity at month 9 of stability reopens the assessment

Two ideas from Q9(R1) matter for the analyst:

  • Hazard is not risk. A hazard is the potential to cause harm; risk combines the probability of that harm with its severity. “This solvent is toxic” is a hazard statement; “at the residual level this method can detect, the exposure is X% of the PDE” is a risk statement.
  • Formality is a dial, not a switch. A one-line rationale, a risk-ranking table, and a full cross-functional FMEA are all valid quality risk management — the guideline asks you to match the formality to what is at stake, and to say why.

The toolbox

Each tool below gets its own full walkthrough — mechanics, a worked analytical example, and where it breaks down:

ToolBest forNotes
FMEA / FMECAFailures of a process or method built from many stepsThe workhorse in analytical development — full walkthrough →
Fault tree analysis (FTA)Working backward from one defined failure to its contributing causesGood for OOS root-cause work
HACCPIdentifying and controlling critical points in a processOrigin in food safety; maps well to manufacturing
HAZOPDeviations from design intent, guided-word by guided-wordMore common in process/engineering than in the QC lab
Risk ranking and filteringComparing many risks that don’t share a scalePortfolio-level and site-level decisions
Ishikawa (fishbone) / PHAStructuring a first-pass hazard identificationOften the front end of an FMEA

FMEA in action

Failure Mode and Effects Analysis decomposes a method or process into steps, and for each step asks: what could fail (failure mode), what would that do (effect), why would it happen (cause), and how would we catch it (controls). Each mode is scored:

Risk Priority Number = Severity × Occurrence × Detection

  • Severity — how bad the effect is for the patient or the decision (a wrong release decision scores high; a re-run scores low).
  • Occurrence — how often the cause is expected to produce the failure.
  • Detection — how likely the existing controls are to catch it before it matters. High detection score = poorly detected — this scale runs backward, and it is where most FMEAs go wrong.

Modes with a high RPN, or a high severity regardless of RPN, get an action; then the mode is re-scored to show the action worked. The number is easy to game and easy to over-trust — see the full FMEA walkthrough for a worked multi-failure-mode example and the known weaknesses worth teaching so students don’t over-trust it.

From risk assessment to control strategy

A control strategy is the planned set of controls — derived from current product and process understanding — that assures performance and quality. It is the output of risk management, not a separate exercise:

  • Attribute risk assessment decides which quality attributes are critical (CQAs) and therefore need a specification and a method.
  • Method risk assessment (an FMEA against the analytical target profile) decides which method parameters need to be controlled, and how tightly — this is where a robustness study is a risk-control activity, not a validation checkbox.
  • The specification (Q6) and the stability program (Q1) are risk decisions in numeric form.
  • Under the analytical procedure lifecycle, what counts as a reportable change to a method is set by the risk it carries.

Worked example — nitrosamine risk assessments. Between 2018 and 2023 every marketing authorization holder had to assess every product for the risk of N-nitrosamine impurities (NDMA, NDEA, and drug-specific nitrosamines), triggered by the valsartan recalls. The assessment is a textbook QRM: identify the hazard (potent mutagenic carcinogens), analyze the risk (synthetic route, nitrite sources, secondary amines, recovered solvents, water; then confirmatory testing), control it (route changes, nitrite scavengers, tightened limits at ppb levels), and communicate it (to the agency, on a deadline). It also shows the analyst’s exposure directly: the risk conclusion depended entirely on whether a method existed that could see a nitrosamine at its acceptable intake — a detection problem.


Source note. Risk management is anchored in ICH Q9(R1), with ICH Q8(R2), Q10, and Q14. FMEA methodology follows IEC 60812 and the AIAG-VDA FMEA handbook. The nitrosamine case follows the EMA/FDA guidance and Article 5(3) referral outcomes. (Instructor: confirm the Q9(R1) adoption date and current EMA nitrosamine guidance revision.)

1 - FMEA in Detail — Scoring, Scaling, and Where It Breaks

Failure Mode and Effects Analysis worked end to end on an HPLC assay method: the RPN formula, a full failure-mode table with a before/after action, why detection runs backward, and the known weaknesses that make RPN easy to over-trust.
A banner titled 'FMEA in Detail — Scoring, Scaling, and Where It Breaks' with the tagline 'Find the important failures. Take the right action. Don't over-trust the number.' and the note that this is a worked example from an HPLC assay method, with practical guidance for real decisions. Panels: (1) The One Idea — RPN is a prioritization tool, not a measurement; it tells you which failure mode to look at first, not how much worse one is than another, and treating it like it does is the single most common way an FMEA goes wrong, with the note to focus effort where it matters for the patient and avoid gold-plating the rest; (2) How FMEA Works — a five-step chevron: list the process steps (e.g., sample prep, separation, detection, data analysis, reporting), identify failure modes (what could fail at each step), analyze effects and causes (what would it do, why would it happen, how would we catch it), score S, O, D (Severity, Occurrence, Detection), take action and re-score (implement risk controls and re-score to show the improvement) — captioned as a living tool, updated when methods, instruments, materials, or knowledge change; (3) The RPN Formula — RPN = S × O × D, with Severity (how bad the effect is on the patient or the decision, 1 = negligible to 10 = catastrophic e.g. potential patient harm), Occurrence (how often it is expected to happen, 1 = remote to 10 = very high frequency), and Detection (how likely current controls are to catch it before it matters, with the callout that a high score means poorly detected because the scale runs backward — 1 = almost certain to detect, 10 = very unlikely to detect); (4) Typical 1–10 Scales (Examples) — a table mapping score bands (10, 5, 1) to example Severity, Occurrence, and Detection descriptions, with a note to use defined, documented criteria tailored to the method, product, and patient risk; (5) Worked Example — HPLC Assay Method FMEA, a five-row table (mis-integrated peak, wrong diluent used, column-to-column carryover, drifting calibration curve, co-eluting unknown degradant) each with process step/failure mode, effect, cause, current control, S, O, D, RPN, a risk control action, and a re-scored RPN, with the note that carryover (RPN 200) outranks drifting calibration (RPN 108) even though a wrong release decision from drifting calibration may seem worse — because detection was poor (D = 8) — RPN doing its job of surfacing the blind spot, not just the scariest-sounding failure; (6) Why the Same RPN Can Mean Very Different Things — two failure modes (A: rare but severe, S=9 O=2 D=5; B: more frequent, less severe, S=5 O=3 D=6) both scoring RPN 90, with the point that a severity-first reviewer would act on A first regardless of the tied RPN, which is exactly why you shouldn't rank by RPN alone and why any mode with severity ≥ 9 should be automatically flagged for action; (7) Known Limitations — RPN is an ordinal product, not a true measurement (100 is not twice as bad as 50); detection and occurrence are often guesses, and Q9(R1) highlights this subjectivity, asking for defined scales, cross-functional input, and documented assumptions; different (S,O,D) combinations can give the same RPN with very different meaning; it can miss low-probability, high-severity events (use severity-first rules); it is not a substitute for scientific judgment — it's a tool to structure it; (8) FMEA vs. FMECA — a side-by-side comparison: FMEA uses S×O×D (RPN), prioritizes failure modes, is simple and widely used, good for method development; FMECA adds criticality analysis (e.g., a severity/probability matrix), often includes failure-mode ratios, is better for high-risk or regulated products, and most analytical FMEAs are effectively FMECAs in practice; (9) When to Use Other Risk Tools — a table of five tools (fault tree analysis for working backward from a failure, useful for OOS root-cause investigation; HACCP for identifying and controlling critical points, useful in manufacturing or sample handling; HAZOP for deviations from design intent, useful in process/engineering systems; risk ranking and filtering for comparing many unrelated risks, useful for site or portfolio decisions; Ishikawa/fishbone/PHA for first-pass hazard identification, useful for early method or process review); (10) From Risk to Control Strategy — a five-step numbered flow: identify CQAs (attribute risk assessment), assess method parameters (method FMEA), implement controls (e.g., robustness, system suitability), set specifications (Q6) and stability program (Q1), monitor and manage change (Q14), with the note that a control strategy is the output of risk management, not a separate exercise; (11) Worked Case — Nitrosamine Risk Assessment, a bulleted walkthrough: identify hazard (potent mutagenic carcinogens, e.g. NDMA, NDEA, drug-specific), analyze risk (synthetic route, nitrite sources, secondary amines, recovered solvents), control risk (route changes, nitrite scavengers, tighter limits at ppb levels), communicate (to the agency, on a deadline), analytical challenge (need methods sensitive enough to detect at the acceptable intake). Footer: 'Science + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'All science ultimately serves people.'

The one idea

RPN is a prioritization tool, not a measurement. It tells you which failure mode to look at first — it does not tell you how much worse one failure mode is than another, and treating it like it does is the single most common way an FMEA goes wrong.

Mechanics

Failure Mode and Effects Analysis decomposes a method or process into steps, and for each step asks: what could fail (failure mode), what would that do (effect), why would it happen (cause), and how would we catch it (controls)? Each mode is scored on three independent 1–10 scales and multiplied:

Risk Priority Number = Severity × Occurrence × Detection

  • Severity — how bad the effect is for the patient or the decision. A wrong release decision (a failing batch shipped, or a good batch scrapped) scores high; a re-run that costs a day scores low.
  • Occurrence — how often the cause is expected to produce the failure, from historical data or, absent that, engineering judgment.
  • Detection — how likely the existing controls are to catch the failure before it matters. High detection score = poorly detected — this scale runs backward from the other two, and it is where most FMEAs go wrong: a “10” means “we would almost certainly miss this,” not “we’d definitely catch it.”

Modes with a high RPN, or a high severity regardless of RPN, get a corrective action; the mode is then re-scored to show the action actually moved the number, not just noted “action taken.”

Worked example — an HPLC assay method

Failure modeEffectCauseCurrent controlSODRPNActionRe-scored RPN
Mis-integrated peakWrong reported assay valueManual integration override without documented rationalePeer review of chromatograms846192Require documented integration parameters; lock auto-integration settings8 × 4 × 2 = 64
Wrong diluent usedLow or erratic recoverySimilar-looking bottles stored adjacent on the benchAnalyst training63590Segregate diluent storage; barcode-scan verification at weigh-in6 × 3 × 2 = 36
Column-to-column carryoverGhost peak misread as an impurityInsufficient wash gradient between injectionsNone — relies on visual inspection558200Add a blank injection after each sample series; extend wash time5 × 5 × 3 = 75
Drifting calibration curveSystematic bias in reported resultStandard degraded between preparation and useSystem suitability at run start only926108Add a mid-run suitability check; shorten standard hold time9 × 2 × 3 = 54
Co-eluting unknown degradantImpurity result reported lowInsufficient resolution between API and degradantResolution check in system suitability934108Switch to an orthogonal column for confirmatory testing9 × 3 × 2 = 54

Two things worth noticing in this table: the carryover mode (RPN 200) outranks the drifting-calibration mode (RPN 108) even though a wrong release decision from a drifting curve is arguably worse — because carryover’s detection score was so bad (8: nobody was actually looking for it). That is RPN doing its job: surfacing the blind spot, not just the scariest-sounding failure.

Why the same RPN can mean very different things

Failure modeSODRPN
A92590
B53690

Both score 90. Mode A is a rare but severe failure that’s moderately well detected; mode B is a more frequent, less severe failure that’s poorly detected. A severity-first reviewer would act on A first regardless of the tied RPN — which is exactly the argument for not ranking a whole FMEA by RPN alone, and for flagging any mode with severity ≥ 9 for action independent of its RPN.

FMEA vs. FMECA

FMECA adds a formal criticality analysis on top of FMEA — instead of (or alongside) the RPN product, each failure mode’s criticality is assessed against a defined severity/probability matrix, often with failure-mode ratios when one cause can produce several distinct failure modes. In practice, most analytical-development FMEAs are really FMECAs in miniature: teams already flag “any severity ≥ 9 regardless of RPN” as an action trigger, which is a criticality rule, not a pure RPN rule.

Known weaknesses — worth teaching so students don’t over-trust the number

  • RPN is an ordinal product treated as if it were interval data; an RPN of 100 is not “twice as bad” as 50, and — as shown above — different (S, O, D) triples give the same RPN with very different meaning.
  • Detection and occurrence are often guessed. Q9(R1) explicitly flags this subjectivity and asks for it to be managed (defined scales, cross-functional scoring, documented assumptions).
  • Many programs now supplement or replace RPN with a severity-first criticality matrix, or with risk ranking and filtering when comparing failure modes across unrelated processes.

When to reach for something else

FMEA decomposes one process step by step and scores every mode on the same three scales — it’s the right tool when the process is defined and you’re building or revising its control strategy. Reach for fault tree analysis instead when you’re working backward from a failure that has already happened and need to trace its root cause; reach for risk ranking and filtering when you’re comparing risks that don’t share a process or a scale at all.

2 - Fault Tree Analysis — Working Backward From a Failure

FTA starts from a failure that already happened and works backward through AND/OR logic to its contributing causes — the standard tool for an OOS root-cause investigation, and the mirror image of FMEA’s forward-looking approach.
A banner titled 'Fault Tree Analysis — Working Backward From a Failure' with the tagline 'Find the causes. Fix the system. Prevent recurrence.' and the note that FTA starts from a failure that already happened and works backward through logic to its contributing causes — a standard tool for OOS investigations and a key part of a robust quality system. Panels: (1) The One Idea — FMEA asks, before anything has gone wrong, 'what could fail in this process?'; FTA asks, after something already has, 'what chain of causes could have produced exactly this failure?'; they run in opposite directions through the same failure space, and a mature quality system uses both, captioned 'Start with the failure. Work backward. Find the real cause. Prevent it from happening again.'; (2) Worked Example — OOS Assay Result (HPLC), a fault tree with the top event 'Reported assay result outside the specification range,' branching through an OR gate into Analytical/Laboratory Error (False OOS) and True Failure (Product Out of Spec); the error branch further ORs into standard out of date or degraded, system suitability failed but overridden or missed, sample preparation error, and instrument malfunction, each with a way to check it; the true-failure branch ANDs a manufacturing process producing an out-of-spec batch with no analytical error found in the investigation; captioned that an OOS investigation follows this logic — Phase I (laboratory investigation) works the analytical-error branches first, Phase II (full investigation) proceeds to the true-failure branch only if no assignable analytical cause is found; (3) Steps to Build a Fault Tree — a seven-step numbered list: define the top event (be specific), identify immediate causes, use AND/OR logic gates to build branches for all plausible pathways, continue decomposition by asking 'why' until reaching basic events, evaluate with data to confirm or rule out each basic event, identify root cause(s) (there may be more than one), implement and verify corrective actions; (4) Logic Gates — an OR gate (any one input can cause the event above, 'this or this') and an AND gate (all inputs must be present for the event above), with the note to use OR and AND gates to map all plausible causes, continuing until reaching basic events that can be confirmed or ruled out with data; (5) Example Basic Events (HPLC Assay) — a bulleted list: wrong diluent used, column contamination or carryover, incorrect mobile phase composition, detector wavelength mis-set, integration parameters changed, analyst transcription error, software/processing error, degraded reference standard, sample instability, environmental factor (temperature); (6) FTA vs. FMEA — Complementary Tools, a comparison table across direction (backward/reactive vs. forward/prospective), starting point (a specific observed failure vs. a process/method/system), purpose (find root causes vs. identify and prioritize potential failures), output (causal logic tree vs. Risk Priority Numbers and an action plan), use case (OOS investigations and deviations vs. analytical method development, process design, control strategy), and when to use (after a failure has occurred vs. before failures occur, and to check coverage after an event); (7) Strengths and Limitations — strengths (structured logical approach, ensures all plausible causes are considered, visual and easy to communicate, drives data-based investigation, links directly to corrective actions and prevention) and limitations (only as good as the top event's definition, can become large and complex, does not rank or prioritize causes, requires disciplined use of basic events, may miss systemic issues if the scope is too narrow, typically used alongside FMEA or risk ranking for a complete view), captioned 'Define the failure clearly. Use the data. Keep it focused. FTA finds the cause — your quality system prevents the next one.'; (8) Key Takeaways — FTA works backward from a defined failure using AND/OR logic; it is the standard tool for OOS root-cause investigations; the quality of the analysis depends on a clear top event and disciplined decomposition; FTA does not score or rank — it is a diagnostic tool, not a replacement for FMEA; use FTA and FMEA together to build a stronger, more resilient quality system. Footer: 'Science + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'All science ultimately serves people.'

The one idea

FMEA asks, before anything has gone wrong, “what could fail in this process?” FTA asks, after something already has, “what chain of causes could have produced exactly this failure?” They run in opposite directions through the same failure space, and a mature quality system uses both.

Mechanics

A fault tree starts with a single, precisely defined top event — the failure that occurred — and branches downward through logic gates to the conditions that could produce it:

  • An AND gate means every branch beneath it must be true for the event above to occur (e.g., a wrong result reaches release and the reviewer misses it).
  • An OR gate means any one branch beneath it is sufficient (e.g., a degraded standard, a mis-set instrument parameter, or a transcription error could each independently cause a wrong reported value).

The tree bottoms out in basic events — causes you either confirm or rule out with data, not further decomposition. No formal Boolean notation is required to use this at the bench; the value is in the discipline of writing every “or this could have happened” branch down before deciding which one is true.

Worked example — an out-of-specification (OOS) assay result

Top event: Reported assay result outside the specification range.

Reported assay OOS
 └─ OR: Result is a true failure vs. a lab/analytical error
     ├─ OR (analytical/lab error branch)
     │    ├─ Standard was out of date or degraded
     │    │    → check standard prep date, storage conditions, prior QC data
     │    ├─ System suitability failed but was overridden or missed
     │    │    → review the suitability data logged that run
     │    ├─ Sample preparation error (dilution, weighing, transcription)
     │    │    → re-check the prep worksheet against the raw balance/pipette record
     │    └─ Instrument malfunction (detector drift, pump seal, injector carryover)
     │         → review instrument maintenance and diagnostic logs
     └─ AND (true-failure branch)
          ├─ Manufacturing process produced an out-of-spec batch
          │    → review batch record deviations, in-process controls
          └─ No analytical error found in the OOS investigation above
               → confirms the result should stand

An OOS investigation under Q7/GMP follows exactly this shape: Phase I (laboratory investigation) works the analytical-error branches first, because a confirmed lab error can invalidate the result without ever reaching the manufacturing branch; Phase II (full investigation) only proceeds down the true-failure branch once Phase I finds no assignable analytical cause.

When to reach for it vs. FMEA

FTA is reactive — it exists because a specific, already-observed failure needs a root cause, and it only makes sense once that top event is precisely defined. FMEA is prospective — it exists to find failure modes before they happen, and it doesn’t require anything to have gone wrong yet. In practice, a documented FMEA is often what an OOS investigation checks against: “was this failure mode already identified, and if so, why did the existing control not catch it?”

Known weaknesses

  • FTA is only as good as the top event’s definition — a vaguely stated failure (“something went wrong with the assay”) produces an unusably broad tree.
  • Trees for a complex, multi-step method can become large fast; without discipline about what counts as a “basic event,” the tree can sprawl without converging on an actionable root cause.
  • FTA doesn’t score or prioritize the way RPN does — it’s a diagnostic tool for one failure, not a ranking tool across many, which is why it’s typically paired with an FMEA or risk ranking rather than used as the whole risk program.

3 - HACCP — Critical Control Points, Borrowed From Food Safety

Hazard Analysis and Critical Control Points asks a narrower question than FMEA: not every failure mode in a process, but where the few points are whose failure directly threatens the patient — worked through a sterile-fill bioburden-control example.
A banner titled 'HACCP — Critical Control Points, Borrowed From Food Safety' with the tagline 'Prevent Hazards. Protect Patients.' and the note that this is a practical, risk-based approach to focus on the few points where control is essential. Panels: (1) The One Idea — instead of scoring every failure mode in a process, HACCP asks a narrower, sharper question: where in this process is a critical control point, a step where losing control means the hazard reaches the patient with nothing downstream left to catch it, captioned 'Find the few make-or-break points. Control what matters. Protect the patient.'; (2) The Seven Principles of HACCP — a seven-step chevron: hazard analysis (what biological, chemical, or physical hazards could occur at each step), identify CCPs (which steps are the last point where the hazard can be prevented, eliminated, or reduced to an acceptable level), establish critical limits (define a measurable threshold for each CCP), establish monitoring (how and how often the critical limit will be checked, and by whom), establish corrective actions (what happens when a critical limit is exceeded), verification (show the system works — trend data, media fills, audits), record keeping (document everything, the backbone of an auditable system) — captioned that the first five principles define the control strategy, and verification and record-keeping make it sustainable; (3) HACCP Decision Logic — Is It a CCP?, a flowchart: does a hazard exist at this step that could affect the patient (No → not a CCP); is this step the last point where the hazard can be prevented, eliminated, or reduced to an acceptable level (No → not a CCP, consider other controls; Yes → this step is a CCP), with the quote 'A downstream test that only detects a hazard is not a CCP if it cannot remove the hazard'; (4) Origins and Relevance — HACCP was developed for NASA's manned space program to ensure astronaut food had zero tolerance for contamination, and maps directly to sterile and biologic manufacturing, which share that same 'no downstream catch' property, captioned 'From space food to patient medicines — same principle: prevent the hazard'; (5) Worked Example — Sterile Fill/Finish Bioburden Control, a five-row table (raw material receipt, compounding, sterilizing-grade filtration, aseptic fill, final inspection) each with hazard, whether it's a CCP, critical limit, monitoring, and corrective action — sterilizing-grade filtration is the only 'Yes (CCP)' row (a non-sterile filter passing organisms into the final fill, critical limit is a filter integrity/bubble-point test pre- and post-use, 100% integrity testing every batch, corrective action is fail the batch, do not release, investigate filter lot and process), captioned that filtration is the CCP because it is the last point where the hazard can still be prevented — everything downstream has no way to remove it; (6) HACCP vs. FMEA — Different Questions, a comparison table across direction (focused/narrow vs. comprehensive/broad), key question (where can the hazard reach the patient vs. what can fail at each step), scope (few critical control points vs. all failure modes), best for (manufacturing and process risk, e.g. sterility, cross-contamination vs. analytical methods and detailed process analysis), output (control strategy — CCPs, critical limits, monitoring vs. prioritized list of failure modes (RPN) and actions), captioned 'Use HACCP when a small number of make-or-break points exist. Use FMEA when you need full coverage of every failure mode.'; (7) Strengths and Limitations — strengths (focuses resources on what matters most, simple/structured/easy to communicate, well-suited for processes with zero tolerance for patient risk e.g. sterility, drives clear measurable control strategies) and limitations (works best when there are a small number of CCPs, requires deep process understanding to identify the true CCP, can be misapplied if detection steps are labeled as CCPs, less natural for analytical-method risk than FMEA); (8) Key Takeaways — HACCP asks where the hazard can reach the patient, focus on CCPs; a CCP is the last point to prevent, eliminate, or reduce the hazard; define measurable critical limits and monitor them; corrective actions must be specific and immediate; use HACCP for manufacturing/process risk and FMEA for method risk. Footer: 'Science + Quality + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'Prevent Today. Protect Tomorrow.'

The one idea

Instead of scoring every failure mode in a process, HACCP asks a narrower, sharper question: where in this process is a critical control point — a step where losing control means the hazard reaches the patient, with nothing downstream left to catch it?

Mechanics

HACCP originated in food safety (developed for NASA’s manned space program, to guarantee astronaut food had zero tolerance for contamination) and maps cleanly onto sterile and biologic manufacturing, which share that same “no downstream catch” property. The full method has seven principles; the ones that matter for a control-strategy discussion are:

  1. Conduct a hazard analysis — what biological, chemical, or physical hazards could occur at each process step?
  2. Identify critical control points (CCPs) — of all the steps, which ones are the point where the hazard can still be prevented, eliminated, or reduced to an acceptable level? A step downstream of the true control point is not itself a CCP, even if a hazard could theoretically show up there.
  3. Establish critical limits — a measurable threshold for each CCP (a temperature, a pressure differential, a bioburden count) that separates “in control” from “out of control.”
  4. Establish monitoring — how and how often the critical limit is checked, and by whom.
  5. Establish corrective action — what happens, specifically, the moment a critical limit is exceeded.

(The remaining two principles — verification and record-keeping — are the documentation backbone that makes the first five auditable, and aren’t specific to any one CCP.)

Worked example — sterile fill/finish bioburden control

StepHazardIs it a CCP?Critical limitMonitoringCorrective action
Raw material receiptContaminated excipientNo — caught downstream—Certificate of analysis reviewReject lot
CompoundingMicrobial ingress during mixingNo — bioburden reducible later—Environmental monitoring (routine)Investigate, re-clean
Sterilizing-grade filtrationA non-sterile filter passes organisms into the final fillYes — nothing downstream removes a missed organismFilter integrity test (bubble point) passes pre- and post-use100% integrity testing, every batchFail the batch; do not release; investigate filter lot and process
Aseptic fillEnvironmental contamination during fillingPartially — mitigated by isolator/RABS design, not a single measurable limit— (engineering control, not a CCP in the classic sense)Continuous particle counts, media fillsHalt line, investigate
Final inspectionVisible particulateNo — a quality check, not a hazard-elimination point—Visual inspectionReject unit

The filtration step is the CCP because it is the last point where the hazard (a non-sterile product) can still be prevented — everything upstream can be caught or corrected later in the process, and everything downstream has no way to remove an organism that already got through. That is the test for “is this a CCP,” not “could something go wrong here.”

When to reach for it vs. FMEA

FMEA decomposes an entire process into every failure mode and scores each one — useful when you want comprehensive coverage of a method or process. HACCP deliberately does the opposite: it narrows attention to the small number of points where losing control is unrecoverable, which is exactly right for manufacturing and process risk (sterility assurance, allergen control, cross-contamination) but a poor fit for analytical method risk, where FMEA’s step-by-step, fully-scored decomposition is what regulators and most labs actually expect.

Known weaknesses

  • Works best when there really are a small number of make-or-break points; forcing a HACCP structure onto a process with many, roughly-equally-important risks just reproduces an FMEA with extra steps.
  • Identifying the true CCP takes real process understanding — misidentifying a downstream inspection point as a CCP gives false confidence, since it doesn’t actually prevent the hazard, only detects it after the fact.
  • Less natural for analytical-method risk (where FMEA dominates) than for manufacturing/process risk, where it originated and still fits best.

4 - HAZOP — Deviations From Design Intent

Hazard and Operability study asks, guided word by guided word, what happens if a process parameter is too much, too little, reversed, or accompanied by something unintended — a process/engineering tool applied here to a chromatography example.
A banner titled 'HAZOP — Deviations From Design Intent' with the tagline 'Ask "What if?" before it happens.' and the note that this is a structured, guide-word approach to identify how process parameters can deviate, what could happen, and how to keep the process safe, robust, and in control. Panels: (1) The One Idea — HAZOP doesn't start from a list of known failure modes, it starts from the process's own design intent and systematically asks what happens if reality deviates from it, one guide word at a time, parameter by parameter, captioned 'Use guide words to challenge assumptions, uncover what could go wrong, and strengthen the design before it happens.'; (2) How a HAZOP Works — a five-step chevron: define scope and team (process section, e.g. chromatography; multidisciplinary team — process, analytical, engineering, quality, EHS), list design intent (process steps, key parameters like flow/temperature/pressure/pH/time/concentration, normal operating ranges), apply guide words (ask what happens for each parameter using NO, MORE, LESS, AS WELL AS, REVERSE, OTHER THAN), identify causes and consequences (what could cause the deviation, what are the consequences for safety/quality/operability/regulatory), assess safeguards and actions (what safeguards already exist, are they sufficient, define actions for gaps) — captioned to document, track actions, and follow through to closure; (3) The HAZOP Guide Words — a table of six guide words with meaning and a generic example: NO (completely absent, no flow — pump failure), MORE (higher than intended, more pressure than the system is rated for), LESS (lower than intended, less temperature than required), AS WELL AS (something additional is present, an unexpected contaminant enters the intended feed), REVERSE (opposite direction, reverse flow through a failed check valve), OTHER THAN (completely different than intended, a different reagent is charged instead of the intended one); (4) Worked Examples — two side-by-side tables applying guide words: Example 1, HPLC Flow Rate (Analytical Process) — MORE (flow rate too high, pump set point drifts high, column overpressure/seal failure/resolution loss, system pressure alarm), LESS (flow rate too low, partial pump blockage, retention times shift/poor resolution, system suitability retention-time check), NO (no flow, pump stalls, no separation occurs/run aborts, run-sequence software flags failed injection), AS WELL AS (contaminant in mobile phase, impurity or wrong solvent present, interfering peaks/method failure, incoming solvent specification/UV scan), REVERSE (reverse flow, check valve failure, column damage/carryover, check valve and system pressure direction check), OTHER THAN (different solvent, wrong solvent selected, unexpected selectivity/no separation, barcode verification/method review); Example 2, Bioreactor Temperature (Manufacturing Process) — MORE (temperature too high, heating control fails open, reduced cell viability/altered glycosylation — a CQA hit, independent high-temperature interlock), LESS (temperature too low, cooling jacket over-corrects, reduced growth rate/extended run time, continuous temperature logging with trend alarms), NO (no temperature control, control system failure, loss of culture control/batch failure, alarm and automated shutdown), AS WELL AS (contaminant introduced, leaking line or open port, microbial contamination, closed system design/sterility assurance), REVERSE (reverse flow of coolant, valve mispositioned, overheating risk, valve position interlocks), OTHER THAN (wrong medium added, operator error, cell stress/off-spec product, barcode scanning/double-check procedure), with the note that a higher temperature may not cause an obvious failure but can silently change a critical quality attribute like glycosylation — exactly the kind of risk HAZOP is designed to uncover; (5) HAZOP vs. FMEA — Different Starting Points, Complementary Tools, a comparison table across direction (starts from design intent/deviations vs. starts from process steps/failure modes), key question (what happens if this parameter deviates vs. what could fail at each step), focus (process/engineering design and operability vs. analytical methods and detailed processes), output (list of credible deviations, causes, consequences, safeguards, actions vs. prioritized failure modes (RPN) and actions), best for (manufacturing and process design vs. QC methods and laboratory processes), captioned 'Use HAZOP for process and engineering risk. Use FMEA for analytical methods. They often inform each other.'; (6) Strengths and Limitations — strengths (structured/systematic way to challenge the design, uncovers non-obvious deviations using guide words, focuses on patient safety/product quality/operability, ideal for complex processes and new designs, multidisciplinary — brings different perspectives together) and limitations (can be time-consuming and exhaustive, requires good process understanding to identify true CCPs, no built-in scoring — prioritization is a separate step, can overlap with FMEA if both are used, less natural for analytical-method risk where FMEA is usually preferred); (7) Key Takeaways — HAZOP uses guide words to explore deviations from design intent; not every step is a CCP — only where the hazard cannot be caught downstream; focus on causes, consequences, and existing safeguards; works best for process and engineering risk (e.g., manufacturing); use HAZOP and FMEA together for a stronger, more complete risk program. Footer: 'People + Process + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'Science Today. Healthier Tomorrows.'

The one idea

HAZOP doesn’t start from a list of known failure modes the way FMEA does — it starts from the process’s own design intent and systematically asks what happens if reality deviates from it, one guide word at a time, parameter by parameter.

Mechanics

For each parameter at each step of a process (flow rate, temperature, pressure, pH, concentration, time), a HAZOP team applies a fixed set of guide words and asks what a deviation of that kind would actually cause:

Guide wordMeaningGeneric example
NOThe parameter is completely absentNo flow — pump failure
MOREThe parameter is higher than intendedMore pressure than the system is rated for
LESSThe parameter is lower than intendedLess temperature than the reaction requires
AS WELL ASSomething additional is presentAn unexpected contaminant enters with the intended feed
REVERSEThe parameter or flow runs backwardReverse flow through a check valve that has failed
OTHER THANSomething completely different happens insteadA different reagent is charged than intended

Unlike FMEA, HAZOP doesn’t score every deviation on Severity/Occurrence/Detection — the output is a qualitative list of credible deviations, their causes, consequences, and existing safeguards, with follow-up actions where the safeguards look thin.

Worked example — HPLC flow rate and a bioreactor’s temperature

Guide wordParameterDeviationConsequenceSafeguard
MOREHPLC flow ratePump set point drifts highColumn overpressure, potential seal failure, resolution lossSystem pressure alarm, method-defined pressure limit
LESSHPLC flow ratePartial pump blockageRetention times shift, poor resolution between API and impuritySystem suitability retention-time check
NOHPLC flow ratePump stallsNo separation occurs at all; run abortsRun-sequence software flags a failed injection
MOREBioreactor temperatureHeating control fails openReduced cell viability, altered glycosylation profile (a CQA hit)Independent high-temperature interlock, separate from the control loop
LESSBioreactor temperatureCooling jacket over-correctsReduced growth rate, extended run timeContinuous temperature logging with trend alarms

Notice the bioreactor row: a MORE temperature deviation doesn’t just risk an obvious failure (dead cells) — it can silently shift a critical quality attribute (glycosylation) while the culture still looks healthy, which is exactly the kind of consequence a guide-word walk-through is designed to surface deliberately, rather than relying on someone to have already thought of it.

When to reach for it vs. FMEA

HAZOP and FMEA overlap heavily in outcome — both end up identifying deviations and their consequences — but HAZOP is organized around the process’s design intent, parameter by parameter, which makes it a natural fit for engineering and process-design teams examining a new unit operation (a reactor, a filtration skid, a chromatography skid) before it’s ever run. Most QC labs default to FMEA for method risk because the “steps” of a method are already well defined; HAZOP earns its keep more in process/engineering contexts where the parameters, not discrete process steps, are the natural unit of analysis.

Known weaknesses

  • Applying every guide word to every parameter at every step can be slow and exhaustive for a complex process — teams often scope it to the parameters most likely to matter, which reintroduces some of the same judgment calls HAZOP is meant to avoid.
  • Without a scoring step, prioritizing which deviations to act on first is a separate, later exercise — HAZOP tells you what could deviate, not which deviation matters most.
  • The overlap with FMEA means running both on the same process is often redundant; most sites pick one as the primary tool for a given risk type (HAZOP for process design, FMEA for methods) rather than running both routinely.

5 - Risk Ranking and Filtering — Comparing Risks That Don't Share a Scale

When risks come from different processes, products, or sites and don’t share a common scale, risk ranking and filtering normalizes them against weighted criteria to build one prioritized list — worked through a site quality council’s quarterly resourcing decision.
A banner titled 'Risk Ranking and Filtering — Comparing Risks That Don't Share a Scale' with the tagline 'Prioritize what matters. Make the best use of limited resources.' and the note that this is a structured, transparent way to compare different risks across products, processes, and sites — and build a defensible action plan. Panels: (1) The One Idea — FMEA scores risks within one process on one shared scale; risk ranking and filtering compares risks across processes, products, or sites that have no natural shared scale, by defining and weighting the criteria that make one risk matter more than another, captioned 'Different risks. One decision. Focus on what matters most for the patient, the business, and compliance.'; (2) How It Works — A Simple, Repeatable Process, a five-step chevron: define criteria (choose the criteria that matter across all risks, e.g. patient impact, regulatory exposure, likelihood, detectability), set weights (assign weights to reflect what matters most in this decision — patient impact typically highest), score each risk (rate each risk against every criterion using a consistent scale, e.g. 1–5), calculate weighted score (multiply scores by weights and sum to get a total weighted score), rank and filter (sort the risks, apply a cutoff e.g. top 3, and document decisions and rationale) — captioned 'From many risks to a focused action plan.'; (3) Typical Criteria and Example Weights — a table of six criteria with why it matters and an example weight: patient impact/safety/quality (direct effect on patient safety or product quality, weight 3), regulatory exposure (risk of inspection findings, warning letter, or enforcement, weight 2), likelihood (how likely the risk is to occur, weight 1), detectability (how likely it is to be detected before it impacts patients, weight 1), business impact — optional (cost, supply, reputation, weight 1), timeline pressure — optional (committed dates, customer obligations, weight 0.5–1); (4) Worked Example — Site Quality Council (Quarterly Resourcing): five risks competing for the same limited investigation and remediation budget this quarter, in a table with patient impact (×3), regulatory exposure (×2), likelihood (×1), weighted score, rank, and decision — stability OOS trend Product A (5,4,3 → 26, rank 1, fund this quarter), recurring documentation deviation/data-integrity adjacent (3,5,4 → 23, rank 2, fund this quarter), method-transfer gap Product B new receiving lab (3,3,4 → 19, rank 3, fund as tie-break), pending inspection commitment due date approaching (2,4,5 → 19, rank 4, deferred/tie), aging HPLC fleet increasing downtime (2,1,5 → 13, rank 5, defer/documented); (5) Ranking and Filtering Result — a funnel: Top 2 fund now (scores 26, 23), Next 2 tie at 19 (apply a secondary criterion, e.g. regulatory due date), Defer (score 13, document rationale and review next cycle), captioned 'The goal is a short, defensible action list — not a long table that sits on a shelf.'; (6) Risk Ranking vs. FMEA — Different Purposes, a comparison table across scope (multiple unrelated risks across products/processes/sites vs. one process or method), key question (which of these do we fix first vs. what could fail at each step), scale (weighted criteria, custom vs. S×O×D, a shared scale), output (prioritized list and resourcing decisions vs. list of failure modes and actions), best for (portfolio decisions, site priorities, limited resources vs. detailed method or process analysis); (7) Key Takeaways — define the right criteria and weight them transparently; use a consistent scoring scale; rank, filter, and document the decisions; use a secondary criterion to break ties (e.g. regulatory due date); this is a decision-making tool, not a measurement of absolute risk; (8) Known Weaknesses — weighting is subjective, different stakeholders often disagree; scores can create false precision — a 26 vs. 23 looks decisive but both rest on judgment calls; not a replacement for detailed analysis — use FMEA (or HAZOP) to understand each risk in depth; requires good input data and cross-functional agreement; (9) Who Should Be Involved? — Quality (QA/QC), Manufacturing/Technical Operations, Regulatory Affairs, Supply Chain/Business, Site Leadership, with the note that different perspectives lead to better decisions, and a stronger quality system delivers healthier patients. Footer: 'Science + Risk-Based Thinking = Better Medicines for Patients,' Temple University branding, and the tagline 'Assess risks. Prioritize wisely. Advance together.'

The one idea

FMEA scores risks within one process on one shared scale. Risk ranking and filtering compares risks across processes, products, or sites that have no natural shared scale at all, by explicitly defining and weighting the criteria that make one risk matter more than another.

Mechanics

  1. Define criteria that matter across every risk being compared — typically patient impact, regulatory exposure, likelihood, and detectability, though a portfolio-level exercise might add business impact or timeline pressure.
  2. Weight the criteria to reflect what actually matters most in this decision (patient impact usually carries the most weight; timeline pressure usually carries the least, if it’s included at all).
  3. Score each risk against every criterion, using whatever scale is practical (often 1–5, sometimes qualitative bands converted to numbers).
  4. Compute a weighted score and rank — then filter: set a threshold or a headcount/budget cutoff and act on what clears it, explicitly documenting why anything below the line is being deferred.

The “filtering” half is as important as the ranking half — the exercise exists to produce a short, defensible action list, not just a long sorted table nobody acts on.

Worked example — a site quality council’s quarterly resourcing decision

Five unrelated findings are competing for the same limited investigation and remediation budget this quarter:

RiskPatient impact (×3)Regulatory exposure (×2)Likelihood (×1)Weighted score
Stability OOS trend, Product A5435×3 + 4×2 + 3×1 = 26
Method-transfer gap, Product B (new receiving lab)3343×3 + 3×2 + 4×1 = 19
Aging HPLC fleet (increasing downtime)2152×3 + 1×2 + 5×1 = 13
Recurring documentation deviation (data-integrity adjacent)3543×3 + 5×2 + 4×1 = 23
Pending inspection commitment (due date approaching)2452×3 + 4×2 + 5×1 = 19

Ranked and filtered against a “fund the top three this quarter” cutoff: the stability OOS trend (26) and the documentation deviation (23) fund first regardless of tiebreaks; the method-transfer gap and the inspection commitment tie at 19 and need a secondary criterion (e.g., regulatory due date) to break the tie for the third slot. The aging-fleet risk (13) is explicitly deferred — not ignored, documented as deferred, with the reasoning on record for the next review cycle.

When to reach for it vs. FMEA

Use risk ranking and filtering when the decision spans multiple unrelated risks competing for the same finite resource — funding, staffing, audit time — not when you’re working through the failure modes of a single process or method, which is FMEA’s job. It’s the tool for “which of these five different problems do we fix first,” not “what could go wrong in this one method.”

Known weaknesses

  • The weighting scheme is itself a subjective judgment call — this is the same criticism Q9(R1) raises about FMEA’s Severity/Occurrence/Detection scoring; risk ranking and filtering doesn’t remove that subjectivity, it just moves it up a level, from scoring individual failure modes to weighting the criteria that compare them.
  • Different stakeholders (quality, manufacturing, regulatory affairs) often disagree on the weights themselves — reaching agreement on the weighting is frequently the harder part of the exercise, not the scoring.
  • A weighted score can create false precision — a 26 vs. a 23 looks decisive, but both numbers rest on the same soft inputs as any other risk score, and the ranking should be sanity-checked qualitatively before being treated as a tiebreaker.

6 - Ishikawa / Fishbone / PHA — Structuring the First Pass

Before an FMEA can score failure modes, it needs a reasonably complete list of them — fishbone diagrams and Preliminary Hazard Analysis are how that list gets brainstormed systematically, worked through an unexpected-peak example that feeds directly into an FMEA.
A banner titled 'Ishikawa / Fishbone / PHA — Structuring the First Pass' with the tagline 'Start broad. Capture the possibilities. Feed the FMEA.' and the note that this is a structured way to brainstorm what could go wrong, category by category, so nothing important is missed. Panels: (1) The One Idea — an FMEA is only as complete as its failure-mode list, and that list has to come from somewhere; Ishikawa (fishbone) diagrams and Preliminary Hazard Analysis (PHA) are how you brainstorm it systematically, category by category, instead of relying on whoever's in the room to remember everything, captioned 'Start with a structured brainstorm. Capture the possibilities. Then prioritize with FMEA.'; (2) Worked Example — Fishbone for 'Unexpected Peak in a Stability Sample,' a fishbone diagram with five category branches (Method/procedure: gradient resolution, wrong wavelength, integration parameters, injection volume, sample prep procedure; Materials/reagents/consumables: column degradation, contaminated mobile phase, reference standard cross-contamination, impure reagents/solvents, vial septa/leachables; Machine/instrumentation: detector lamp aging, carryover from prior injection, autosampler needle wash, pump composition error, calibration out of date; Manpower/people: sample preparation error, mislabeled vial/sample mix-up, incorrect method execution, data processing/integration, fatigue/training gap; Environment/lab conditions: temperature excursion, humidity effects, vibration, power interruption, sample storage conditions) all pointing to the effect 'Unexpected Peak in a Stability Sample'; (3) Example Causes by Category — a table repeating the five categories (Method, Materials, Machine, Manpower, Environment) each with a bulleted list of candidate causes matching the fishbone diagram; (4) What Are They? — Ishikawa/Fishbone (a visual diagram to organize possible causes of a problem, uses standard categories like Method/Materials/etc., great for team brainstorming and root-cause thinking) and Preliminary Hazard Analysis/PHA (first-pass identification of what could go wrong, a simple hazard/cause/effect table or checklist, used early in development or for new processes to scope what needs deeper analysis); (5) From Brainstorm to Action — a four-step chevron: brainstorm (Ishikawa/PHA — capture as many plausible causes as possible), convert to failure modes (turn key causes into FMEA failure modes), score and prioritize (FMEA — assess S, O, D and identify high-risk items), implement controls (take action and re-score if needed) — captioned 'Fishbone and PHA provide the input. FMEA provides the prioritization.'; (6) When to Use Each Tool — a two-column comparison: use Ishikawa/PHA when the process or method is new or unfamiliar, a cross-functional team needs a shared brainstorm, you want broad coverage before scoring, or it's an early stage of development; go directly to FMEA when failure modes are already well understood, methods or processes are mature and well-characterized, you need to prioritize and assign actions, or regulatory/management expects a scored risk assessment; (7) Known Limitations — purely qualitative, no built-in scoring or prioritization; coverage depends on who is in the room; does not guarantee completeness; not a complete QRM record — typically feeds into an FMEA or risk-ranking exercise; (8) Key Takeaways — start with a clearly defined effect or hazard; use standard categories to ensure a complete brainstorm; fishbone/PHA is qualitative — no scoring; convert key causes to FMEA failure modes for prioritization; it's a front end to QRM, not a replacement for FMEA or risk ranking, with the summary equation 'Many Perspectives + Better Ideas + More Complete Risk Assessment = Safer Patients' and the quote 'A structured brainstorm today prevents surprises tomorrow.' Footer: 'Science + People + Process = Better Medicines for Patients,' Temple University branding, and the tagline 'Identify. Understand. Control. Deliver.'

The one idea

An FMEA is only as complete as its failure-mode list, and that list has to come from somewhere. Ishikawa (fishbone) diagrams and Preliminary Hazard Analysis are how you brainstorm it systematically, category by category, instead of relying on whoever’s in the room to remember everything from experience.

Mechanics

An Ishikawa diagram starts from a defined effect (an observed or feared problem) and branches into standard categories of contributing cause. Adapted for an analytical lab, the categories are usually:

  • Method — the procedure itself: parameters, steps, order of operations
  • Materials — reagents, standards, reference materials, columns, consumables
  • Machine — instrumentation: hardware, software, calibration state
  • Manpower — analyst training, technique, fatigue, handoffs
  • Environment — temperature, humidity, lighting, vibration, power quality

Preliminary Hazard Analysis (PHA) is a lighter, earlier-stage cousin — a first-pass brainstorm of what could possibly go wrong before a process even exists in detail, often just a simple hazard/cause/effect table, used to scope what a later, more formal risk assessment needs to cover.

Worked example — “unexpected peak in a stability sample”

CategoryCandidate causes brainstormed
MethodInsufficient gradient resolution; wrong wavelength selected; integration parameters too aggressive
MaterialsColumn degradation; contaminated mobile phase; reference standard cross-contamination
MachineDetector lamp aging (baseline drift creating false peaks); carryover from a prior injection; autosampler needle wash insufficient
ManpowerSample prep error introducing a degradant precursor; mislabeled vial swapped with another study
EnvironmentLab temperature excursion affecting sample stability between prep and injection

This is deliberately a long, unfiltered list — the point of the fishbone pass is coverage, not judgment. Three or four of these branches then become the failure-mode column of a follow-on FMEA: “contaminated mobile phase” becomes a scoreable failure mode with its own severity, occurrence, and detection; “detector lamp aging” becomes another. The fishbone did the brainstorming; the FMEA does the prioritizing.

When to reach for it vs. FMEA directly

Skip straight to FMEA when the failure modes are already well understood from experience — a mature, well-characterized method rarely needs a fresh fishbone pass. Reach for Ishikawa or PHA first when the process or method is new or unfamiliar, or when a cross-functional team is starting from very different mental models of what could go wrong and needs a shared, structured brainstorm before anyone starts scoring anything.

Known weaknesses

  • Purely qualitative — a fishbone diagram or PHA table has no scoring or prioritization built in; it can surface a long list of contributing factors without telling you which ones actually matter.
  • Coverage depends heavily on who’s in the room; the category headings help structure the brainstorm, but they don’t guarantee completeness the way a systematic top-down decomposition (like FMEA’s step-by-step structure) does.
  • It is not, on its own, a complete quality risk management record — it’s the front end that typically feeds into an FMEA or risk ranking and filtering exercise, not a substitute for either.