How to Do Turbine Oil Failure Root Cause Analysis
Considering Turbine Oil as an Asset — with ISO 55000 and ICML 55 in Mind
By Khash — MLE Perspective
In many plants, turbine oil failure root cause analysis is treated as a laboratory problem:
“The TAN increased.”
“The MPC is high.”
“The RULER dropped.”
“The oil is dark.”
“The bearing temperature increased.”
But from an MLE and asset management perspective, this is not enough.
A turbine oil failure is rarely just an “oil problem.” It is usually a lubricated asset management failure involving the oil, machine, environment, operating condition, maintenance practices, sampling quality, filtration strategy, contamination control, human decisions, and business risk.
ISO 55000 defines asset management around realizing value from assets over their life cycle. The current ISO 55000:2024 standard provides the overview, terminology, and principles for developing a proactive asset management system. (ISO) ICML 55 is written as an enabling standard for lubricated asset management in support of ISO 55000, with ICML 55.1:2019 covering requirements for optimized lubrication of mechanical physical assets. (info.lubecouncil.org)
Therefore, when we investigate turbine oil failure, we should not ask only:
“Why did the oil fail?”
We should ask:
“Why did the turbine oil asset fail to deliver its required function, at the required risk level, for the required life-cycle value?”
That is a completely different level of thinking.
1. First Principle: Define Turbine Oil as an Asset
From an ISO 55000 mindset, an asset is something that has actual or potential value to the organization. Turbine oil clearly qualifies as an asset because it protects much higher-value assets:
- Steam turbine bearings
- Gas turbine bearings
- Generator bearings
- Load gearboxes
- Hydraulic control systems
- Trip and governor systems
- Journal bearings and thrust bearings
- Servo valves and control valves
- Pumps, coolers, filters, reservoirs, and piping
- Production availability and plant reliability
A turbine oil charge in a large machine may cost much less than the turbine, but its functional failure can cause:
- bearing wiping,
- thrust bearing distress,
- servo valve sticking,
- hydraulic instability,
- high bearing temperature,
- forced outage,
- failed start,
- failed trip response,
- production loss,
- safety risk,
- unplanned oil replacement,
- expensive flushing,
- and loss of confidence in the reliability program.
So the oil should not be treated as a consumable.
It should be treated as a managed physical asset with a life-cycle plan.
2. Define What “Turbine Oil Failure” Really Means
Many people define turbine oil failure only as:
“The oil is outside the lab limit.”
This is too narrow.
A turbine oil failure should be defined as:
The inability of the turbine oil system to perform its required functions within acceptable risk, reliability, cleanliness, chemical stability, and life-cycle cost limits.
This includes both oil condition failure and oil-system functional failure.
Turbine oil required functions
The oil is expected to:
| Function | Failure Example |
|---|---|
| Lubricate bearings | high friction, wear, bearing temperature rise |
| Remove heat | poor heat transfer, oxidized oil, cooler fouling |
| Prevent corrosion | water, acids, depleted rust inhibitor |
| Transfer hydraulic energy | servo valve sticking, poor control response |
| Separate water | poor demulsibility, stable emulsion |
| Release air | foaming, air entrainment, microdieseling |
| Resist oxidation | TAN rise, RULER depletion, sludge/varnish |
| Remain clean | high particle count, debris, filter plugging |
| Stay chemically stable | additive depletion, degradation products |
| Protect asset availability | avoiding forced outages and trips |
Therefore, a “failed” turbine oil may still look acceptable in one test but fail in its real function.
For example:
- TAN may be acceptable, but MPC is high.
- MPC may drop suddenly because varnish has deposited on surfaces.
- RULER may look acceptable after top-up, but deposits remain in the system.
- Particle count may be clean, but soluble varnish is high.
- Water ppm may be low, but demulsibility is poor.
- Oil color may be dark, but not necessarily failed.
- Oil color may be clear, but varnish potential may be severe.
This is why turbine oil RCA must be evidence-based, multi-parameter, and asset-risk-based.
3. Link the RCA to ISO 55000 Thinking
ISO 55000 encourages life-cycle asset management, alignment with organizational objectives, and value realization. (ISO) Applied to turbine oil RCA, this means the RCA must answer five asset-management questions:
1. Value
How did this oil failure affect business value?
Examples:
- lost generation,
- lost production,
- forced outage,
- bearing damage,
- trip risk,
- oil replacement cost,
- flushing cost,
- safety exposure,
- maintenance cost,
- reduced machine confidence.
2. Alignment
Were oil maintenance activities aligned with the criticality of the turbine?
A 100 MW steam turbine should not be managed with the same oil-analysis frequency and filtration strategy as a small auxiliary pump.
3. Leadership and accountability
Who owns turbine oil reliability?
Common problem:
- Operations owns the machine.
- Maintenance owns the equipment.
- Laboratory owns the sample result.
- Procurement owns the oil purchase.
- Reliability owns the failure report.
- Nobody owns the oil as an asset.
This is one of the biggest root causes.
4. Assurance
Was there a system to assure that the oil remained fit for service?
This includes:
- proper sampling point,
- correct sampling frequency,
- correct test slate,
- trend analysis,
- alarms,
- actions,
- filtration strategy,
- contamination control,
- inspection during outage,
- and follow-up verification.
5. Life-cycle management
Was the oil managed from selection to disposal?
Life-cycle stages include:
- specification,
- purchase,
- receipt inspection,
- storage,
- transfer,
- filling,
- commissioning,
- routine operation,
- sampling,
- analysis,
- purification,
- additive monitoring,
- varnish control,
- water control,
- contamination control,
- top-up management,
- outage inspection,
- partial replacement,
- full replacement,
- flushing,
- disposal,
- lessons learned.
If RCA ignores this life cycle, it becomes only a laboratory interpretation, not true asset management.
4. Link the RCA to ICML 55 Thinking
ICML 55 focuses on optimized lubrication of mechanical physical assets and supports ISO 55000 alignment. It looks at lubrication management as a structured system, not random oil changes or occasional lab tests. (info.lubecouncil.org)
For turbine oil RCA, ICML 55 thinking means we must investigate the lubrication program itself.
Not only:
“What happened to the oil?”
But also:
“Which lubrication-management process allowed this failure mechanism to develop, remain undetected, or remain uncorrected?”
Typical ICML 55-related RCA areas include:
| Lubrication Management Area | RCA Question |
|---|---|
| Lubricant selection | Was the oil suitable for turbine type, temperature, base oil group, OEM requirements, and duty cycle? |
| Storage and handling | Was the new oil contaminated before use? |
| Contamination control | Were particles, water, air, fuel gas, process chemicals, or cleaning chemicals controlled? |
| Oil analysis | Was the test slate complete enough? |
| Sampling | Was the sample representative, hot, live-zone, and repeatable? |
| Training | Did staff understand MPC, RULER, TAN, water, demulsibility, foam, and air release? |
| Condition monitoring | Were trends reviewed or only single results? |
| Maintenance strategy | Was filtration reactive or proactive? |
| Documentation | Were oil events, top-ups, filter changes, and alarms recorded? |
| Continuous improvement | Were previous failures converted into improved standards? |
This is where MLE thinking becomes very important.
An MLE should not only interpret oil analysis. An MLE should design the lubrication management system that prevents recurrence.
5. Step-by-Step Method for Turbine Oil Failure RCA
Step 1: State the Failure Clearly
Do not begin with assumptions.
Avoid weak statements like:
“The oil failed due to varnish.”
Instead, define the failure event precisely.
Examples:
Case A — Bearing temperature issue
“The DE journal bearing temperature of Steam Turbine ST-101 increased from 82°C to 101°C over six weeks, while load remained similar. MPC increased from 18 to 42 during the same period, and inspection found brown deposits on bearing pads.”
Case B — Servo valve sticking
“The hydraulic control valve response became unstable during startup. Oil analysis showed high MPC, reduced antioxidant reserve, and fine insoluble degradation products. Servo valve inspection showed sticky brown deposits.”
Case C — Oil chemical degradation
“The turbine oil TAN increased by 0.35 mg KOH/g above new oil value within 12 months, RULER phenolic antioxidant dropped below expected trend, and MPC increased sharply after hot operating periods.”
A proper failure statement should include:
- equipment tag,
- oil type,
- oil age,
- failure symptom,
- affected component,
- time frame,
- operating condition,
- lab trend,
- physical evidence,
- business consequence.
Step 2: Classify the Failure Type
Turbine oil failures can be classified into six major categories.
1. Chemical degradation failure
Examples:
- oxidation,
- nitration,
- thermal degradation,
- additive depletion,
- acid formation,
- sludge,
- soluble varnish precursors.
Key tests:
- TAN by ASTM D664,
- RULER by ASTM D6971,
- FTIR oxidation/nitration,
- RPVOT by ASTM D2272,
- MPC by ASTM D7843,
- color,
- viscosity.
2. Contamination failure
Examples:
- water,
- particles,
- fibers,
- rust,
- process contamination,
- cleaning chemicals,
- wrong oil,
- top-up contamination.
Key tests:
- Karl Fischer water,
- particle count,
- patch microscopy,
- elemental analysis,
- FTIR,
- demulsibility,
- visual inspection.
3. Varnish/deposit failure
Examples:
- bearing pad deposits,
- servo valve sticking,
- reservoir bathtub ring,
- cooler fouling,
- filter plugging,
- thrust bearing temperature rise.
MPC by ASTM D7843 is widely used for varnish potential in turbine oils. EPT describes MPC as an analytical test used to determine the tendency of lubricant to form varnish deposits, while other technical sources emphasize that MPC is valuable but should be interpreted with awareness of limitations and trend behavior. (EPT Clean Oil)
4. Physical property failure
Examples:
- viscosity change,
- poor air release,
- foam tendency,
- poor demulsibility,
- low flash point,
- poor heat transfer.
Key tests:
- viscosity at 40°C,
- viscosity index,
- flash point,
- air release,
- foam,
- demulsibility,
- density.
5. System design or operating failure
Examples:
- hot spots,
- undersized reservoir,
- poor residence time,
- excessive turbulence,
- air entrainment,
- wrong return-line design,
- poor filtration location,
- cooler leakage,
- dead zones,
- electrostatic discharge in filters,
- high bearing metal temperature.
6. Management system failure
Examples:
- wrong sampling point,
- missing baseline,
- no alarm limits,
- no trend review,
- poor ownership,
- procurement-driven oil selection,
- no contamination-control standard,
- no varnish-control strategy,
- no corrective action tracking.
This sixth category is often the real root cause.
6. Step 3: Build the Turbine Oil Failure Timeline
A serious RCA must create a timeline.
Do not look only at the latest sample.
Build a timeline including:
| Timeline Item | Why It Matters |
|---|---|
| New oil baseline | Without baseline, used-oil interpretation is weak |
| Fill date | Establishes oil age |
| Top-up events | May dilute or mask trends |
| Oil replacement or bleed-and-feed | May hide degradation history |
| Filter changes | May indicate insoluble loading |
| Cooler leaks | May explain water |
| Start/stop cycles | May accelerate oxidation and varnish |
| Trips | May create thermal stress |
| Load changes | Affect temperature and oxidation rate |
| Bearing temperature trend | Links oil health to machine behavior |
| Servo valve issues | Links varnish to control reliability |
| Oil temperature | Main oxidation accelerator |
| Reservoir temperature | May differ from bearing-zone stress |
| Lab results | Must be trended, not isolated |
| Outage inspections | Surface evidence |
| Filtration/purification history | Corrective action effectiveness |
The timeline should answer:
“When did the failure mechanism begin, when was it detectable, when was it detected, and when was action taken?”
This links RCA to the P-F curve.
The earlier the oil degradation is detected, the longer the plant has to act before functional failure.
7. Step 4: Separate Symptoms, Failure Modes, Mechanisms, and Root Causes
This is where many reports become weak.
Example
Symptom
Bearing temperature increased.
Failure mode
Loss of normal lubricating and heat-transfer condition at bearing surface.
Failure mechanism
Varnish deposit on bearing surface reduced heat transfer, changed surface energy, affected oil film behavior, and increased frictional instability.
Physical root cause
Oxidation by-products and polar degradation products accumulated in oil and deposited on cooler metal surfaces.
Systemic root cause
No proactive varnish monitoring, no MPC trend review, no hot-oil resin-based varnish removal strategy, and no asset-risk-based oil management plan.
This structure is much stronger than saying:
“Root cause: varnish.”
Varnish is often not the root cause.
Varnish is usually the visible consequence of deeper chemical, thermal, operational, and management failures.
8. Step 5: Collect Evidence in Four Layers
A strong RCA needs four evidence layers.
Layer 1: Oil analysis evidence
Minimum test slate for turbine oil RCA:
| Test | Purpose |
|---|---|
| Viscosity at 40°C | Detect wrong oil, oxidation, contamination |
| TAN ASTM D664 | Detect acidic degradation |
| RULER ASTM D6971 | Antioxidant reserve |
| RPVOT ASTM D2272 | Oxidation stability trend |
| MPC ASTM D7843 | Varnish/deposit tendency |
| FTIR | Oxidation, nitration, contamination |
| Karl Fischer water | Dissolved/free water level |
| Particle count | Solid contamination |
| ISO cleanliness code | Cleanliness trend |
| Elemental analysis | Wear, additive, contamination |
| Demulsibility | Water separation capability |
| Foam tendency/stability | Foam risk |
| Air release | Air entrainment risk |
| Patch microscopy | Visual debris/deposit evidence |
| Color ASTM D1500 | General condition, not varnish diagnosis alone |
Important point:
No single test proves the full root cause.
MPC alone is not enough. TAN alone is not enough. RULER alone is not enough. Color alone is definitely not enough.
Layer 2: Machine condition evidence
Collect:
- bearing metal temperature,
- bearing drain oil temperature,
- vibration trend,
- axial position,
- thrust bearing temperature,
- differential pressure across filters,
- hydraulic control response,
- valve stroking behavior,
- oil pressure,
- oil flow,
- cooler performance,
- reservoir temperature,
- load profile,
- number of starts and stops.
Layer 3: Physical inspection evidence
During outage or inspection, look for:
- brown/orange varnish,
- black carbonaceous deposits,
- sticky servo valve deposits,
- reservoir bathtub ring,
- deposits on sight glass,
- deposits on bearing pads,
- cooler plate fouling,
- filter element discoloration,
- sludge at reservoir bottom,
- water pockets,
- rust,
- foam marks,
- dead-leg deposits.
Take photographs.
Every RCA should include photos where possible.
Layer 4: Management system evidence
Review:
- oil specification,
- sampling procedure,
- sampling point design,
- frequency,
- lab method consistency,
- alarm limits,
- action limits,
- historical trends,
- oil top-up records,
- filter replacement history,
- purifier operation history,
- reservoir inspection records,
- training records,
- responsibility matrix,
- previous RCA actions.
This is the ISO 55000 and ICML 55 layer.
Without this layer, the RCA is incomplete.
9. Step 6: Use a Structured RCA Logic
For turbine oil RCA, I recommend combining:
- 5-Why analysis
- Fault tree analysis
- Fishbone diagram
- Barrier analysis
- Life-cycle asset review
Example: High MPC and Bearing Temperature Rise
Problem
Bearing temperature increased and MPC is high.
5-Why logic
Why did bearing temperature increase?
Because heat transfer and friction condition at the bearing surface changed.
Why did the bearing surface condition change?
Because varnish-like deposits formed on bearing pads.
Why did deposits form?
Because polar oxidation by-products exceeded oil solubility and deposited on metal surfaces.
Why did oxidation by-products accumulate?
Because antioxidant reserve declined, oil operated at elevated temperature, and varnish precursors were not removed.
Why were they not removed?
Because the plant relied on standard mechanical filtration, which removes particles but not dissolved oxidation by-products.
Why did the asset management system allow this?
Because turbine oil was treated as a consumable and not as a managed lubricated asset with a varnish-control strategy, alarm limits, ownership, and life-cycle plan.
This is a much stronger RCA.
10. Step 7: Interpret Key Failure Mechanisms
A. Oxidation
Oxidation is one of the main turbine oil degradation pathways. It is accelerated by:
- heat,
- air,
- water,
- metals,
- contamination,
- depleted antioxidants,
- high residence time at elevated temperature,
- entrained air,
- microdieseling,
- poor reservoir design.
Oxidation produces:
- acids,
- sludge,
- varnish precursors,
- polar degradation products,
- color change,
- increased TAN,
- reduced RPVOT,
- antioxidant depletion.
RCA question:
“Was oxidation the primary mechanism, or only one contributor?”
B. Varnish formation
Varnish is not just “dirt.”
It is often formed from polar, oil-degradation products that may remain dissolved at high temperature and precipitate when conditions change.
Varnish may deposit on:
- bearing pads,
- thrust shoes,
- servo valves,
- hydraulic valves,
- cooler surfaces,
- reservoir walls,
- filters,
- narrow clearances,
- low-flow areas.
MPC is important, but interpretation must consider trend, sampling temperature, storage conditions, and whether varnish has already deposited. A high MPC may show elevated insolubles or varnish potential, while a sudden MPC decrease without corrective action can sometimes indicate that material has left the oil and deposited on surfaces. (TestOil)
C. Antioxidant depletion
RULER is critical because turbine oil can look visually acceptable while antioxidant reserve is already weak.
RCA questions:
- Which antioxidant depleted faster: phenolic or aminic?
- Was depletion linear or sudden?
- Was the oil topped up and artificially refreshed?
- Was RULER compared with the new oil baseline?
- Was RULER interpreted with TAN, MPC, and RPVOT?
D. Water contamination
Water can cause:
- additive hydrolysis,
- rust,
- oxidation acceleration,
- poor demulsibility,
- filter plugging,
- microbial growth in some systems,
- bearing distress,
- reduced oil film strength.
RCA questions:
- Is the water dissolved, emulsified, or free?
- Is the source cooler leakage, steam seal leakage, condensation, washdown, breathers, storage, or poor handling?
- Is the oil still able to separate water?
- Is water a cause or a consequence of poor reservoir management?
E. Air entrainment and foam
Air can cause:
- oxidation acceleration,
- unstable oil pressure,
- poor hydraulic response,
- cavitation-like effects,
- microdieseling,
- varnish acceleration,
- bearing film disturbance.
RCA questions:
- Is return oil entering above the oil level?
- Is reservoir residence time adequate?
- Is the oil level correct?
- Are suction leaks present?
- Is the defoamant depleted or filtered out?
- Is the oil contaminated with incompatible top-up oil?
F. Wrong oil or incompatible top-up
Many turbine oil problems begin with innocent top-up.
RCA questions:
- Was the top-up oil the same brand and formulation?
- Was it from the same product family?
- Was compatibility tested?
- Was new oil baseline available?
- Did additive elements change?
- Did foam, air release, demulsibility, or MPC change after top-up?
11. Step 8: Identify Root Causes at Three Levels
A professional RCA should classify root causes into three levels.
Level 1: Technical root causes
Examples:
- oxidation,
- water contamination,
- varnish formation,
- antioxidant depletion,
- poor air release,
- high particle contamination,
- wrong oil,
- thermal stress,
- additive incompatibility.
Level 2: Equipment/system root causes
Examples:
- cooler leak,
- undersized reservoir,
- poor return-line design,
- excessive turbulence,
- dead zones,
- poor filtration location,
- inadequate purification,
- high bearing temperature,
- poor breather system,
- poor drain design.
Level 3: Management-system root causes
Examples:
- no turbine oil asset strategy,
- no oil criticality ranking,
- poor sampling practice,
- incomplete test slate,
- no MPC testing,
- no RULER baseline,
- no oil life-cycle plan,
- no ownership,
- no action limits,
- no RCA trigger criteria,
- no follow-up verification,
- procurement selected oil without reliability input,
- no ICML 55-style lubrication management system.
The third level is where recurrence prevention happens.
12. Step 9: Define Failure Consequence and Risk
Since ISO 55000 is value-focused, RCA must quantify consequence.
Use a risk matrix.
Example risk categories
| Consequence | Example |
|---|---|
| Safety | trip system malfunction, fire risk, emergency shutdown |
| Production | forced outage, derating, failed start |
| Asset damage | bearing damage, servo valve damage, cooler fouling |
| Cost | oil replacement, flushing, labor, lost generation |
| Environmental | oil disposal, leakage |
| Reputation | repeated reliability failure |
| Compliance | failure to follow internal asset management procedures |
A turbine oil with MPC 45 in a non-critical small auxiliary unit may be a medium risk.
The same MPC 45 in a critical gas turbine hydraulic/control oil system may be a severe risk.
This is why fixed lab limits are not enough.
Risk must be linked to:
- asset criticality,
- operating temperature,
- failure history,
- component sensitivity,
- oil volume,
- redundancy,
- outage window,
- replacement cost,
- and business impact.
13. Step 10: Develop Corrective Actions Using Asset Management Logic
Corrective actions should not be random.
They should be linked to root causes.
Example corrective action table
| Root Cause | Corrective Action | Verification |
|---|---|---|
| High varnish potential | Install resin-based varnish removal or chemistry management system | MPC reduction trend, patch color, bearing temp stabilization |
| Water ingress | Repair cooler leak or improve sealing/breather system | KF water trend, demulsibility recovery |
| Antioxidant depletion | Evaluate partial/full oil replacement or reconditioning strategy | RULER, RPVOT, TAN trend |
| Poor sampling | Install live-zone sampling point and train technicians | Repeatable lab results |
| Incomplete oil analysis | Add MPC, RULER, demulsibility, air release, patch microscopy | Better early detection |
| No ownership | Assign turbine oil asset owner | RCA action closure |
| Poor filtration strategy | Upgrade from reactive filtration to proactive contamination and chemistry control | ISO cleanliness, MPC, TAN trends |
| Wrong top-up | Create approved oil list and compatibility procedure | No unexplained additive/property shifts |
| Poor reservoir condition | Inspect and clean during outage | Deposit reduction, improved oil stability |
Corrective actions must include:
- action owner,
- deadline,
- technical justification,
- risk reduction target,
- verification method,
- follow-up sampling date,
- expected trend,
- acceptance criteria.
14. Step 11: Create a Turbine Oil Asset Health Index
For asset management, it is useful to convert lab and operational data into a health index.
Example:
Turbine Oil Health Reliability Index — TOHRI
| Parameter | Weight |
|---|---|
| MPC | 20% |
| RULER antioxidant reserve | 20% |
| TAN | 15% |
| RPVOT | 10% |
| Water | 10% |
| Particle count | 10% |
| Demulsibility | 5% |
| Air release/foam | 5% |
| Operating temperature history | 5% |
Then combine with asset criticality:
Oil Risk Priority = (100 – TOHRI) × Asset Criticality Factor
Example:
- TOHRI = 62
- Asset criticality factor = 5
- Risk priority = (100 – 62) × 5 = 190
This allows management to see oil risk as an asset-management risk, not only as a laboratory number.
15. Step 12: RCA Report Structure
A professional turbine oil RCA report should include:
1. Executive summary
- what failed,
- consequence,
- most probable root cause,
- risk level,
- required actions.
2. Asset information
- turbine tag,
- OEM,
- oil type,
- oil volume,
- oil age,
- reservoir volume,
- bearing type,
- hydraulic/control system details,
- criticality.
3. Failure description
- symptoms,
- dates,
- operating conditions,
- affected components.
4. Oil analysis trend
Include trend charts for:
- MPC,
- TAN,
- RULER,
- RPVOT,
- viscosity,
- water,
- particle count,
- demulsibility,
- foam,
- air release.
5. Physical evidence
- photos,
- patch images,
- bearing inspection,
- filter element inspection,
- reservoir inspection.
6. RCA logic
Use:
- 5-Why,
- fishbone,
- fault tree,
- barrier analysis.
7. Root cause classification
Separate:
- direct cause,
- contributing causes,
- systemic causes.
8. Corrective action plan
Include owner, target date, verification method.
9. Asset management improvement
Link recommendations to:
- ISO 55000 asset value,
- ICML 55 lubrication management,
- turbine oil life-cycle plan,
- risk-based monitoring.
10. Lessons learned
Convert the RCA into a plant standard.
16. Practical Example: Varnish-Related Turbine Oil Failure RCA
Event
A steam turbine experienced increasing bearing temperature and unstable control valve response.
Evidence
- MPC increased from 12 to 38 over 10 months.
- RULER aminic antioxidant dropped significantly.
- TAN increased gradually.
- Bearing temperature increased by 12°C.
- Servo valve response became slow.
- Brown deposits found on valve spool.
- Reservoir wall showed bathtub ring.
- No resin-based varnish removal system was installed.
- Oil analysis did not include MPC until after symptoms appeared.
- Sampling point was from reservoir drain, not live-zone sample.
Direct cause
Deposits formed on bearing and servo valve surfaces, affecting heat transfer and valve movement.
Failure mechanism
Oxidation by-products accumulated in oil, exceeded solubility margin, and deposited as varnish on polar metal surfaces and tight-clearance components.
Contributing causes
- elevated oil temperature,
- antioxidant depletion,
- no varnish-control technology,
- poor sampling location,
- incomplete oil-analysis slate,
- delayed action after early warning signs.
Systemic root cause
The turbine oil was not managed as a critical lubricated asset. The plant lacked a risk-based turbine oil asset management plan aligned with ISO 55000 value realization and ICML 55 lubrication management principles.
Corrective actions
- Establish new oil baseline.
- Add monthly MPC and RULER for critical turbine oils.
- Install live-zone sampling points.
- Implement resin-based varnish and acid-removal strategy.
- Inspect servo valves and bearing surfaces during next outage.
- Define turbine oil alarm and action limits.
- Assign turbine oil asset owner.
- Create turbine oil life-cycle management plan.
- Add RCA trigger criteria.
- Review results monthly in reliability meeting.
Verification
- MPC trend reduced and stabilized.
- Bearing temperature returned to normal band.
- Servo valve response improved.
- TAN stabilized.
- RULER depletion rate slowed.
- Filter differential pressure stabilized.
- No repeat event after defined monitoring period.
17. Common Mistakes in Turbine Oil RCA
Mistake 1: Calling varnish the root cause
Varnish is often the result, not the deepest cause.
Mistake 2: Looking only at the latest sample
Turbine oil RCA requires trends.
Mistake 3: Ignoring sampling quality
A poor sample creates poor diagnosis.
Mistake 4: Treating color as oil health
Oil color is useful but not sufficient. Clear oil can be unhealthy. Dark oil can still be functional depending on other parameters.
Mistake 5: Using particle filtration as varnish control
Mechanical filters remove particles. They do not fully address dissolved oxidation by-products or soluble varnish precursors.
Mistake 6: Ignoring RULER
Antioxidant depletion is often the early warning before TAN and varnish become severe.
Mistake 7: Ignoring operating temperature
Temperature history is essential for oxidation RCA.
Mistake 8: Ignoring top-up history
Top-up can dilute TAN, refresh antioxidants, change additive chemistry, or create compatibility issues.
Mistake 9: No link to business consequence
Without consequence, the RCA remains technical but not managerial.
Mistake 10: No management-system corrective action
Replacing oil without fixing the system only resets the failure clock.
18. MLE-Level RCA Questions
An MLE should ask deeper questions:
- What is the oil’s required function in this asset?
- What is the asset criticality?
- What is the oil’s life-cycle stage?
- What is the new oil baseline?
- What changed in operation?
- What changed in temperature?
- What changed in top-up oil?
- What changed in filtration?
- What changed in contamination exposure?
- What changed in bearing temperature or valve response?
- Was the sample representative?
- Was the test slate complete?
- Were the alarms suitable for this criticality?
- Was the oil analyzed as a trend?
- Was varnish soluble, insoluble, or already deposited?
- Was antioxidant depletion understood?
- Was water dissolved, emulsified, or free?
- Was air entrainment investigated?
- Was the oil management plan proactive or reactive?
- Was the corrective action verified?
- Was the lesson converted into a standard?
19. Final MLE Message
A turbine oil RCA should not finish with:
“Change the oil.”
That is not root cause analysis.
It should finish with:
“This is why the turbine oil asset failed to deliver its required function; this is how the degradation mechanism developed; this is why our management system did not detect or prevent it earlier; this is the risk to the business; and this is the corrected life-cycle asset management plan to prevent recurrence.”
That is the difference between an oil-analysis report and an asset-management RCA.
As an MLE, the target is not only to find the failed property.
The target is to protect the lubricated asset, extend the P-F interval, reduce the risk of functional failure, and maximize the life-cycle value of the turbine oil and the turbomachinery it protects.
In simple words:
Do not investigate turbine oil as a dirty fluid. Investigate it as a failed asset protection system.
Discover more from Turbine Oil Reliability
Subscribe to get the latest posts sent to your email.
