Root Cause Analysis for Turbomachinery: an introduction
- Fernando E. Romero, P.E.

- Jun 13
- 14 min read

When a turbomachinery failure happens, the questions that follow are expensive ones...
I have had the privilege of working for an independent service provider of turbomachinery repairs for almost 25 years.
I work at a place that has built itself to be like a hospital for machines.
I use this analogy when giving tours to visitors.
It is easy to describe what we do if I say:
“Imagine this place is like a hospital.
Most of our work we do, can be considered routine, or scheduled work. Just like you or I would go to the hospital for a yearly checkup, or a yearly eye appointment, or dental cleaning. This is work that is scheduled in advance and if you are like me, you would have tried taking good care of yourself and would not expect any surprises.
But some of our work is emergency work. For those “unplanned” events, like a sudden loss of lubrication in a steam turbine, or a sudden high vibration event, or the sudden stop of a machine.We have built our shops to have skilled mechanics, inspectors, technicians, engineer specialists, metallurgists. And we have equipped them with the best tools so we can provide quick response.”
I’ve seen both types of work for almost 25 years.I’ve seen beautiful machines, some exquisitely designed and extremely well taken care of. I’ve seen some unique machines where there are no more than a handful alike in the world. And I have seen some incredible catastrophic failures, that as an engineer have left me intrigued and flabbergasted on how a machine could have come to that end.
The hospital and finding out why machines get sick
Just like the hospital analogy, the equipment owners (refineries, petrochemical plants and power plants) need their equipment repaired, brought back to life. In terms of the refining industry, when a gas compressor goes down and it affects the throughput of a distilling column, or fractionator, or cracker unit, the losses are measured in millions in loss of production per day.
By the same token, even on planned outages, you should account that the cost of a scheduled maintenance event will be a multiple of the cost of loss of production per day. So one day delay on an outage means not just the loss of production, but you have to add the cost to keep all your contractors working.
I say all of this, because there is immense economical pressure to repair the equipment, and to do it quickly.
My role in my company is to serve as an engineering specialist. Like that doctor in the movies walks to the emergency room and tells the family how an operation is going and informs them of the possible outcomes.
I am often present in the mechanical assembly shop (emergency room) when a unit arrives, to observe the condition of the equipment as it arrives, and as it is disassembled.
I review inspection documents; I may even indicate what type of additional tests to perform.
I may remind personnel how careful they need to be, to treat the equipment as if it was a “crime scene” investigation.
I gather all this insight into the equipment failure, work with a fantastic team of engineers in coming up with options to move forward with a repair, BUT first we discuss performing a root cause failure analysis first.
Before we can get hasty with fixing a machine, we should understand the contributing causes to the failure. We don’t want to fix something, and have it fail again.
Obviously the greater the failure, the higher the risk of personal damage, or equipment damage, the higher we must raise this flag to proceed with caution.
I’ve mentioned this before, but time between maintenance events where you completely open a turbine or compressor are often in the range of 4-8 years.
This means that for that entire time, no one can visually and thoroughly look inside the equipment while it is running.
Every now and then, if the owner has an opportunity and the equipment is not running, and accessible, they can perform the equivalent of a colonoscopy. In the industry, we call it a “bore scope” or “videoscope” inspection. And it is when we insert a tiny camera inside the machine to see if we can appreciate the internal condition.
These inspections are more useful on certain types of machines, like a gas turbine for instance, where there may be multiple entry points to check the condition of hot section blades.
But it is impractical for most steam turbines and compressors in petrochemical service.
There simply is not enough room to snake a tiny camera around blades and impellers and all the guts of the machine.
So, for 4-8 years, most machines are running and no one knows what the inside looks like.
We can measure temperatures and pressures, we can measure shaft displacement, and case vibrations. But these are mostly external or lateral indicators that we have desperately spent decades trying to use to help predict the real conditions inside a machine.
We try our best to monitor machine condition and to use predictive methods to prevent failures. But sometimes, most times, when there is a failure, we learn about it when it is obviously too late. We learn about it once something has already failed and broken.
There is an interesting term used by a pioneer in the field of machinery failure analysis, when something starts to go wrong inside a machine. Mr. Heinz Bloch called this incipient failure.
*side note:
I spent my formative years in Ecuador, so my daily vocabulary is Spanish.
So when I hear an English word that I don’t hear often or never heard before, I have this instant thought in my brain.
I imagine a spelling be, where the kids say: ”can you tell me the root? Can you tell me the origin of the word?” I don’t know why, I cannot help it or stop it.
So, when I read: Incipient. I repeated it in my mind three times, and then I had to look it up.
Incipient, comes from Latin! “incipere”, TO BEGIN!
Mr. Bloch defines a set of stages that culminate in failure. The first two stages include the term “incipient”.
Incipient Failure -> Incipient damage -> distress -> deterioration -> damage-> failure
Incipient failure: the process has begun at the material or component level. A crack has initiated. The surface is beginning to change. The failure mechanism is active. But the component is still performing its function. No one would call it failed if they could see it.
Incipient damage: the first physical evidence of that process becomes detectable, if you know where to look and what to look for. Microscopic. Not visible to the naked eye in most cases. Still no functional consequence.
Distress: the damage is now visible and measurable. The component is under stress from the developing condition. Performance may begin to degrade slightly. A very attentive analyst might detect something in the operating data.
Deterioration: the condition is advancing. Performance degradation becomes measurable. The component is losing capability.
Damage: significant physical change has occurred. The component is compromised. Detection is now straightforward if someone is looking.
Final failure: the component can no longer perform its intended function. This is when most people find out something has gone wrong.
So, a quick recap to tie things together.
Equipment failures can cause great damage to people or the equipment itself.
The cost and impact of failed steam turbines, compressors and gas turbines is measured in millions per day in loss of production.
There is a need to respond quickly with repairs, but also a need to understand what lead to the failure to prevent it from happening again.
Another legitimate reason that goes alongside prevention of a future failure is, that owners and insurance companies need to know who is going to pay the bill?
Did the failure occur because of a material defect, a design defect, improper operation of the equipment, a faulty previous repair, by external events?
I am a big fan of English detective dramas and lured by those storied of insatiable vicars and detective inspectors looking for clues to find a murderer or a criminal.
Perhaps I like them so much, because it reminds me of performing root cause analysis.
Root Cause Analysis and Root Cause Failure Analysis
Root Cause Analysis RCA: is a broad term for the analytical process of tracing a problem backward. Understanding what happened and why it happened until you understand the cause. It can apply to processes, systems, safety incidents, etc.…
Root Cause Failure Analysis RCFA: is the same practice but specifically applied to a physical component or machine. RCFAs are used to investigate machine failures, and physical evidence that requires specialized engineering interpretation.
Both terms are used often to describe the detective work done by engineers, metallurgists, or other investigators.
For half of my career, I was around RCFAs and worked alongside engineers that led these investigations. I collected evidence, data from control rooms, used coordinate measure machines to measure parts, created 3D models to be later used for finite element analysis. Worked with metallurgists cutting samples to be analyzed under microscopes.
I sort of organically grew in experience on how to be a part of an RCFA, but not until the middle of my career as an engineer that I wanted to formally learn a method or a process for conducting an RCA.
So, I had seen the British detective dramas, I know some colleges offer studies in forensic sciences. I have been exposed and learned about Lean methodologies, and the work of Sakichi Toyoda and the 5 Why analysis. I have learned about Ishikawa or fishbone diagrams. I have seen YouTube videos where pilots dissect accidents and NTSB reports. I have been in the room with experts from companies like Chevron or Exxon and observed them direct work related to RCFAs. But I have not found the perfect class.
There are 3 published references that I think are a useful starting point for an engineer that wants to learn more about this. But these books explain some of the tools, present the general concept, but the rest really comes down to practice. And working with someone that has gained this experience and have them as a mentor. I really see no other way.
Let’s dive into the references:
ASM International - Handbook Vol 11
API 585
Machinery Failure Analysis and Troubleshooting: Practical Machinery Management for Process Plants
American Society for Metals International
Is a professional organization founded in 1913 around the steel industry.
They publish encyclopedia size books called the ASM Handbooks, and these things are heavy and deep with knowledge.
And let me explain something, ASM Handbooks are really sort of encyclopedias. They have no regulatory weight, they are not standards or specifications that you are required to follow if cited.
I say this, because there are other societies or associations, like the ASME, or ASTM that do write testing procedures and specifications, that become standards which when cited as a requirement must be followed.
The ASM Handbook Vol11 is titled Failure Analysis and Prevention, and it has so much information on the tools and processes used for investigating material failures, that it spans over two volumes.
This handbook contains a basic guide on how to investigate, the methods used to collect, examine and interpret the evidence.It describes all the lab tools used: microscopy, chemical composition analysis, mechanical testing, etc..And it dives deep into explaining the 4 top failure mechanism affecting metals.
Fatigue and Fracture
Corrosion
Wear
Distortion
The handbook goes deep into explaining what these failures look like, what drives them, and how to recognize them.
From the list of both tools and mechanisms, you will realize all of these are within the expert jurisdiction of metallurgists.
In our analogy of a detective drama, the metallurgist is the doctor or forensic pathologists conducting the autopsies and issuing a report at the most opportune time indicating things like: blunt trauma to the head, of 2 shots at close range to the head, or corrosion assisted fatigue, or ductile overload.
This handbook is an excellent must-have guide for all metallurgists wanting to become acquainted with the methods.And it is good for any other engineer to understand the context, methods and capabilities of each test and tool.
That report that the metallurgist delivers is usually the beginning of the root cause investigation.
Metallurgists describe what they find in the crime scene. They read the broken surface of the metals and describe meticulously what they observe, what chemistry they detect, and what mechanical properties they measure.
This is all evidence, that describes what the material looks like now that it has failed.
They provide a basic description of the failure mechanism. The cause of death.
But they do not explain the motive, the method, or the reason. They do not explain why?
American Petroleum Institute – Recommended Practice 585
If you read my previous posts on my blog, you will know I usually make references to API 687: Equipment Repairs.
This document unfortunately does not explain any aspect of conducting or performing RCFAs.They only mention that, if necessary, the owner should request one.
There is no API document that explains RCFAs for rotating equipment. But there is an API Recommended Practice number 585, titled: Pressure Equipment Integrity Incident Investigation
And this document, although it does not mention rotating equipment, does provide a nice introduction and guidelines to the method.
What I really appreciate from API 585 are how they define a three-level investigation framework:
Level 1: Low consequence
One-or-two-person team, short timeline, experience-driven judgment. You talk to the people involved, look at the evidence, and document your findings. No formal team or structured process beyond what that individual already knows how to do.
Level 2: Medium consequence
A small team (about 5) with different areas of expertise. This is where structured tools come in, and by tools I do not mean software. I mean specific techniques for organizing what you know and what you still need to find out. A written timeline sequencing events from the moment of failure back through the weeks before it. A causal factor chart showing how contributing conditions connect to each other. An evidence matrix tracking every data point, document, and witness account, and what each one confirms or leaves open.
Level 3: High consequence
Full formal investigation. A trained lead investigator, a multidisciplinary team, and potentially external specialists. Timeline running weeks to months. All the Level 2 tools plus deeper analytical work. Metallurgical lab analysis, finite element modeling, rotordynamic studies if the failure mode requires it. The findings need to be defensible under scrutiny from insurers, regulators, or legal counsel. This is forensic investigation in the fullest sense of the word.
In terms of a turbomachinery repair facility, or a day in the life of any equipment or reliability engineer, I would say we use Level 1 investigations on every routine repair job.
Basically, every seal or bearing change, we would be level 1 investigating the condition of the parts we removed from the machine.
I would say that if we repair some equipment that shows some vibration instability, there is a change in process conditions that is causing vibrations, see some wear on some blades, we would be conducting a Level 2 analysis.
And lastly, if we are losing 1 million dollars a day because our gas turbine is not running after a major inspection, that should launch a Level 3 investigation.
Another cool three level thing from API 585 is they present a three-cause layer model. They determine three categories:
Physical cause: what actually failed. The bearing overload. The impeller cover-plate crack. The shaft failure. This is what the metallurgist hands you in the lab report. The cause of death.
Human cause: the act or omission that allowed physical failure to occur. The wrong clearance was used at the last overhaul.Someone raised the trip level in the machinery protection system.The technician who installed a component backwards. This layer requires interviewing people and reviewing decisions, not just examining hardware.
Latent cause: the organizational or institutional condition that made human error possible in the first place. The procedure that had not been updated in fifteen years. The experienced engineer whose position was eliminated to cut costs. The culture that prioritized getting the machine back online over documenting what was found. This is the hardest layer to reach and the most important one to fix. Leave it unaddressed and the same failure repeats itself, with a different machine and a different crew, on a different day.
API RP 585 is sort of a hidden gem, that I would recommend all reliability engineers or repair engineers get acquainted with.
Machinery Failure Analysis and Troubleshooting: Practical Machinery Management for Process Plants by Heinz P. Bloch and Fred K. Geitner.
This book has a lot of relevance to me because it was written by two very prolific engineers.I must admit I do not know much about Mr. Geitner, but I did get to hear Mr. Bloch once and I have read a lot of his other publications.
Mr. Bloch spent part of his career as a machinery specialist at Exxon Chemical in the USA, Mr. Geitner the same at Imperial Oil in Canada. And their book is the closest thing there is to a practitioner’s manual for machinery failure investigation.
The table of contents looks like this:
Chapter 1: The Failure Analysis and Troubleshooting System
Chapter 2: Metallurgical Failure Analysis
Chapter 3: Machinery Component Failure Analysis
Chapter 4: Machinery Troubleshooting
Chapter 5: Vibration Analysis
Chapter 6: Generalized Machinery Problem-Solving Sequence
Chapter 7: Statistical Approaches in Machinery Problem Solving
Chapter 8: Formalized Failure Reporting as a Teaching Tool
Chapter 9: The Seven Cause Category Approach to Root-Cause Failure Analysis
The book is full of wisdom and pictures.
If you plan to purchase this book on Amazon, you will find one complaint about the illustrations being old and pictures being black and white. You must look beyond that!
Here is my version of their Figure 1 as a funnel of a Materials based RCFA.

This book sits somewhere between ASM and API 585. Between these three books you should be able to build a solid understanding of what a metallurgist will do, what evidence he will collect and his findings. And understand the philosophy and method you should follow as an engineer to conduct and finish the investigation.
And if you noticed, I said “understand”. These references will allow you to understand. But there is a gap between understanding and successfully conducting an investigation.
This is where experience and practice come in. And if you are lucky, working with someone that has more experience in this area.
The greatest challenge in this business of conducting RCFAs is gaining experience and being successful in solving the mysteries.
I have participated and conducted some RCFAs on my own. Level 1 type simple investigations. I have had to assist on Level 2 and 3. And I have to tell you, sometimes the keys to solving a multi-million-dollar failure mystery are in the details.
I have spent countless hours looking at drawings, data trends, pictures, assemblies. For someone with more experience or perspective to come and in 30 seconds point out something I had not noticed.
I have been in the room where experts have had to interview operators. I’ve been in control rooms asking questions looking at PI trends, asking if I can get a download of a trend data set, only to find out IT policies or configurations don’t let that data to be easily downloaded.
Imagine having an airplane blackbox, that is designed to record valuable information. And when you go look inside of it, you realize it was configured to get one data sample every 10 seconds, instead of every 0.1 seconds.
It is having the experience and a discerning eye, confidence and the courage to look, to ask for information, to ask for parts, to ask for samples, to ask to be given permission to look yourself.
And then comes the analysis, having the time and the patience to organize the collected information, to formulate the hypothesis, the restraint not to jump to conclusions.
And then in the end there is sometimes the realization that it may not be possible to single out one root cause. Often what we find is there are multiple contributing factors.
It may be that the amount of collateral damage extending from a failure is such that all the evidence from the original failure is destroyed.
Sometimes the evidence gets destroyed when the machine is being dismantled, or because parts were not preserved and were cleaned instead.
This can be both frustrating for the engineer and the owner, and especially the insurer.
The engineer is driven to find the truth, to understand why a failure occurred. The owner has an interest in getting the unit back in service and avoiding another failure. The insurer definitively wants to understand where the responsibility lies as far as covering the losses.
It is a difficult game sometimes. The only way to master it is to practice and seek guidance.
I am still learning. And I suspect I will be for as long as machines keep breaking.
The mystery does not always get easily solved. But the ones who keep asking why are the ones who solve it most often.
If you are working through a failure and need a second set of eyes, or simply want to compare notes, reach out. This is exactly the kind of conversation I enjoy.
Comments