How to Do Turbine Oil Failure Root Cause Analysis

How to Do Turbine Oil Failure Root Cause Analysis

Considering Turbine Oil as an Asset — with ISO 55000 and ICML 55 in Mind

By Khash — MLE Perspective

In many plants, turbine oil failure root cause analysis is treated as a laboratory problem:

“The TAN increased.”
“The MPC is high.”
“The RULER dropped.”
“The oil is dark.”
“The bearing temperature increased.”

But from an MLE and asset management perspective, this is not enough.

A turbine oil failure is rarely just an “oil problem.” It is usually a lubricated asset management failure involving the oil, machine, environment, operating condition, maintenance practices, sampling quality, filtration strategy, contamination control, human decisions, and business risk.

ISO 55000 defines asset management around realizing value from assets over their life cycle. The current ISO 55000:2024 standard provides the overview, terminology, and principles for developing a proactive asset management system. (ISO) ICML 55 is written as an enabling standard for lubricated asset management in support of ISO 55000, with ICML 55.1:2019 covering requirements for optimized lubrication of mechanical physical assets. (info.lubecouncil.org)

Therefore, when we investigate turbine oil failure, we should not ask only:

“Why did the oil fail?”

We should ask:

“Why did the turbine oil asset fail to deliver its required function, at the required risk level, for the required life-cycle value?”

That is a completely different level of thinking.


1. First Principle: Define Turbine Oil as an Asset

From an ISO 55000 mindset, an asset is something that has actual or potential value to the organization. Turbine oil clearly qualifies as an asset because it protects much higher-value assets:

  1. Steam turbine bearings
  2. Gas turbine bearings
  3. Generator bearings
  4. Load gearboxes
  5. Hydraulic control systems
  6. Trip and governor systems
  7. Journal bearings and thrust bearings
  8. Servo valves and control valves
  9. Pumps, coolers, filters, reservoirs, and piping
  10. Production availability and plant reliability

A turbine oil charge in a large machine may cost much less than the turbine, but its functional failure can cause:

  • bearing wiping,
  • thrust bearing distress,
  • servo valve sticking,
  • hydraulic instability,
  • high bearing temperature,
  • forced outage,
  • failed start,
  • failed trip response,
  • production loss,
  • safety risk,
  • unplanned oil replacement,
  • expensive flushing,
  • and loss of confidence in the reliability program.

So the oil should not be treated as a consumable.

It should be treated as a managed physical asset with a life-cycle plan.


2. Define What “Turbine Oil Failure” Really Means

Many people define turbine oil failure only as:

“The oil is outside the lab limit.”

This is too narrow.

A turbine oil failure should be defined as:

The inability of the turbine oil system to perform its required functions within acceptable risk, reliability, cleanliness, chemical stability, and life-cycle cost limits.

This includes both oil condition failure and oil-system functional failure.

Turbine oil required functions

The oil is expected to:

FunctionFailure Example
Lubricate bearingshigh friction, wear, bearing temperature rise
Remove heatpoor heat transfer, oxidized oil, cooler fouling
Prevent corrosionwater, acids, depleted rust inhibitor
Transfer hydraulic energyservo valve sticking, poor control response
Separate waterpoor demulsibility, stable emulsion
Release airfoaming, air entrainment, microdieseling
Resist oxidationTAN rise, RULER depletion, sludge/varnish
Remain cleanhigh particle count, debris, filter plugging
Stay chemically stableadditive depletion, degradation products
Protect asset availabilityavoiding forced outages and trips

Therefore, a “failed” turbine oil may still look acceptable in one test but fail in its real function.

For example:

  • TAN may be acceptable, but MPC is high.
  • MPC may drop suddenly because varnish has deposited on surfaces.
  • RULER may look acceptable after top-up, but deposits remain in the system.
  • Particle count may be clean, but soluble varnish is high.
  • Water ppm may be low, but demulsibility is poor.
  • Oil color may be dark, but not necessarily failed.
  • Oil color may be clear, but varnish potential may be severe.

This is why turbine oil RCA must be evidence-based, multi-parameter, and asset-risk-based.


3. Link the RCA to ISO 55000 Thinking

ISO 55000 encourages life-cycle asset management, alignment with organizational objectives, and value realization. (ISO) Applied to turbine oil RCA, this means the RCA must answer five asset-management questions:

1. Value

How did this oil failure affect business value?

Examples:

  • lost generation,
  • lost production,
  • forced outage,
  • bearing damage,
  • trip risk,
  • oil replacement cost,
  • flushing cost,
  • safety exposure,
  • maintenance cost,
  • reduced machine confidence.

2. Alignment

Were oil maintenance activities aligned with the criticality of the turbine?

A 100 MW steam turbine should not be managed with the same oil-analysis frequency and filtration strategy as a small auxiliary pump.

3. Leadership and accountability

Who owns turbine oil reliability?

Common problem:

  • Operations owns the machine.
  • Maintenance owns the equipment.
  • Laboratory owns the sample result.
  • Procurement owns the oil purchase.
  • Reliability owns the failure report.
  • Nobody owns the oil as an asset.

This is one of the biggest root causes.

4. Assurance

Was there a system to assure that the oil remained fit for service?

This includes:

  • proper sampling point,
  • correct sampling frequency,
  • correct test slate,
  • trend analysis,
  • alarms,
  • actions,
  • filtration strategy,
  • contamination control,
  • inspection during outage,
  • and follow-up verification.

5. Life-cycle management

Was the oil managed from selection to disposal?

Life-cycle stages include:

  1. specification,
  2. purchase,
  3. receipt inspection,
  4. storage,
  5. transfer,
  6. filling,
  7. commissioning,
  8. routine operation,
  9. sampling,
  10. analysis,
  11. purification,
  12. additive monitoring,
  13. varnish control,
  14. water control,
  15. contamination control,
  16. top-up management,
  17. outage inspection,
  18. partial replacement,
  19. full replacement,
  20. flushing,
  21. disposal,
  22. lessons learned.

If RCA ignores this life cycle, it becomes only a laboratory interpretation, not true asset management.


4. Link the RCA to ICML 55 Thinking

ICML 55 focuses on optimized lubrication of mechanical physical assets and supports ISO 55000 alignment. It looks at lubrication management as a structured system, not random oil changes or occasional lab tests. (info.lubecouncil.org)

For turbine oil RCA, ICML 55 thinking means we must investigate the lubrication program itself.

Not only:

“What happened to the oil?”

But also:

“Which lubrication-management process allowed this failure mechanism to develop, remain undetected, or remain uncorrected?”

Typical ICML 55-related RCA areas include:

Lubrication Management AreaRCA Question
Lubricant selectionWas the oil suitable for turbine type, temperature, base oil group, OEM requirements, and duty cycle?
Storage and handlingWas the new oil contaminated before use?
Contamination controlWere particles, water, air, fuel gas, process chemicals, or cleaning chemicals controlled?
Oil analysisWas the test slate complete enough?
SamplingWas the sample representative, hot, live-zone, and repeatable?
TrainingDid staff understand MPC, RULER, TAN, water, demulsibility, foam, and air release?
Condition monitoringWere trends reviewed or only single results?
Maintenance strategyWas filtration reactive or proactive?
DocumentationWere oil events, top-ups, filter changes, and alarms recorded?
Continuous improvementWere previous failures converted into improved standards?

This is where MLE thinking becomes very important.

An MLE should not only interpret oil analysis. An MLE should design the lubrication management system that prevents recurrence.


5. Step-by-Step Method for Turbine Oil Failure RCA

Step 1: State the Failure Clearly

Do not begin with assumptions.

Avoid weak statements like:

“The oil failed due to varnish.”

Instead, define the failure event precisely.

Examples:

Case A — Bearing temperature issue

“The DE journal bearing temperature of Steam Turbine ST-101 increased from 82°C to 101°C over six weeks, while load remained similar. MPC increased from 18 to 42 during the same period, and inspection found brown deposits on bearing pads.”

Case B — Servo valve sticking

“The hydraulic control valve response became unstable during startup. Oil analysis showed high MPC, reduced antioxidant reserve, and fine insoluble degradation products. Servo valve inspection showed sticky brown deposits.”

Case C — Oil chemical degradation

“The turbine oil TAN increased by 0.35 mg KOH/g above new oil value within 12 months, RULER phenolic antioxidant dropped below expected trend, and MPC increased sharply after hot operating periods.”

A proper failure statement should include:

  • equipment tag,
  • oil type,
  • oil age,
  • failure symptom,
  • affected component,
  • time frame,
  • operating condition,
  • lab trend,
  • physical evidence,
  • business consequence.

Step 2: Classify the Failure Type

Turbine oil failures can be classified into six major categories.

1. Chemical degradation failure

Examples:

  • oxidation,
  • nitration,
  • thermal degradation,
  • additive depletion,
  • acid formation,
  • sludge,
  • soluble varnish precursors.

Key tests:

  • TAN by ASTM D664,
  • RULER by ASTM D6971,
  • FTIR oxidation/nitration,
  • RPVOT by ASTM D2272,
  • MPC by ASTM D7843,
  • color,
  • viscosity.

2. Contamination failure

Examples:

  • water,
  • particles,
  • fibers,
  • rust,
  • process contamination,
  • cleaning chemicals,
  • wrong oil,
  • top-up contamination.

Key tests:

  • Karl Fischer water,
  • particle count,
  • patch microscopy,
  • elemental analysis,
  • FTIR,
  • demulsibility,
  • visual inspection.

3. Varnish/deposit failure

Examples:

  • bearing pad deposits,
  • servo valve sticking,
  • reservoir bathtub ring,
  • cooler fouling,
  • filter plugging,
  • thrust bearing temperature rise.

MPC by ASTM D7843 is widely used for varnish potential in turbine oils. EPT describes MPC as an analytical test used to determine the tendency of lubricant to form varnish deposits, while other technical sources emphasize that MPC is valuable but should be interpreted with awareness of limitations and trend behavior. (EPT Clean Oil)

4. Physical property failure

Examples:

  • viscosity change,
  • poor air release,
  • foam tendency,
  • poor demulsibility,
  • low flash point,
  • poor heat transfer.

Key tests:

  • viscosity at 40°C,
  • viscosity index,
  • flash point,
  • air release,
  • foam,
  • demulsibility,
  • density.

5. System design or operating failure

Examples:

  • hot spots,
  • undersized reservoir,
  • poor residence time,
  • excessive turbulence,
  • air entrainment,
  • wrong return-line design,
  • poor filtration location,
  • cooler leakage,
  • dead zones,
  • electrostatic discharge in filters,
  • high bearing metal temperature.

6. Management system failure

Examples:

  • wrong sampling point,
  • missing baseline,
  • no alarm limits,
  • no trend review,
  • poor ownership,
  • procurement-driven oil selection,
  • no contamination-control standard,
  • no varnish-control strategy,
  • no corrective action tracking.

This sixth category is often the real root cause.


6. Step 3: Build the Turbine Oil Failure Timeline

A serious RCA must create a timeline.

Do not look only at the latest sample.

Build a timeline including:

Timeline ItemWhy It Matters
New oil baselineWithout baseline, used-oil interpretation is weak
Fill dateEstablishes oil age
Top-up eventsMay dilute or mask trends
Oil replacement or bleed-and-feedMay hide degradation history
Filter changesMay indicate insoluble loading
Cooler leaksMay explain water
Start/stop cyclesMay accelerate oxidation and varnish
TripsMay create thermal stress
Load changesAffect temperature and oxidation rate
Bearing temperature trendLinks oil health to machine behavior
Servo valve issuesLinks varnish to control reliability
Oil temperatureMain oxidation accelerator
Reservoir temperatureMay differ from bearing-zone stress
Lab resultsMust be trended, not isolated
Outage inspectionsSurface evidence
Filtration/purification historyCorrective action effectiveness

The timeline should answer:

“When did the failure mechanism begin, when was it detectable, when was it detected, and when was action taken?”

This links RCA to the P-F curve.

The earlier the oil degradation is detected, the longer the plant has to act before functional failure.


7. Step 4: Separate Symptoms, Failure Modes, Mechanisms, and Root Causes

This is where many reports become weak.

Example

Symptom

Bearing temperature increased.

Failure mode

Loss of normal lubricating and heat-transfer condition at bearing surface.

Failure mechanism

Varnish deposit on bearing surface reduced heat transfer, changed surface energy, affected oil film behavior, and increased frictional instability.

Physical root cause

Oxidation by-products and polar degradation products accumulated in oil and deposited on cooler metal surfaces.

Systemic root cause

No proactive varnish monitoring, no MPC trend review, no hot-oil resin-based varnish removal strategy, and no asset-risk-based oil management plan.

This structure is much stronger than saying:

“Root cause: varnish.”

Varnish is often not the root cause.

Varnish is usually the visible consequence of deeper chemical, thermal, operational, and management failures.


8. Step 5: Collect Evidence in Four Layers

A strong RCA needs four evidence layers.

Layer 1: Oil analysis evidence

Minimum test slate for turbine oil RCA:

TestPurpose
Viscosity at 40°CDetect wrong oil, oxidation, contamination
TAN ASTM D664Detect acidic degradation
RULER ASTM D6971Antioxidant reserve
RPVOT ASTM D2272Oxidation stability trend
MPC ASTM D7843Varnish/deposit tendency
FTIROxidation, nitration, contamination
Karl Fischer waterDissolved/free water level
Particle countSolid contamination
ISO cleanliness codeCleanliness trend
Elemental analysisWear, additive, contamination
DemulsibilityWater separation capability
Foam tendency/stabilityFoam risk
Air releaseAir entrainment risk
Patch microscopyVisual debris/deposit evidence
Color ASTM D1500General condition, not varnish diagnosis alone

Important point:

No single test proves the full root cause.

MPC alone is not enough. TAN alone is not enough. RULER alone is not enough. Color alone is definitely not enough.

Layer 2: Machine condition evidence

Collect:

  • bearing metal temperature,
  • bearing drain oil temperature,
  • vibration trend,
  • axial position,
  • thrust bearing temperature,
  • differential pressure across filters,
  • hydraulic control response,
  • valve stroking behavior,
  • oil pressure,
  • oil flow,
  • cooler performance,
  • reservoir temperature,
  • load profile,
  • number of starts and stops.

Layer 3: Physical inspection evidence

During outage or inspection, look for:

  • brown/orange varnish,
  • black carbonaceous deposits,
  • sticky servo valve deposits,
  • reservoir bathtub ring,
  • deposits on sight glass,
  • deposits on bearing pads,
  • cooler plate fouling,
  • filter element discoloration,
  • sludge at reservoir bottom,
  • water pockets,
  • rust,
  • foam marks,
  • dead-leg deposits.

Take photographs.

Every RCA should include photos where possible.

Layer 4: Management system evidence

Review:

  • oil specification,
  • sampling procedure,
  • sampling point design,
  • frequency,
  • lab method consistency,
  • alarm limits,
  • action limits,
  • historical trends,
  • oil top-up records,
  • filter replacement history,
  • purifier operation history,
  • reservoir inspection records,
  • training records,
  • responsibility matrix,
  • previous RCA actions.

This is the ISO 55000 and ICML 55 layer.

Without this layer, the RCA is incomplete.


9. Step 6: Use a Structured RCA Logic

For turbine oil RCA, I recommend combining:

  1. 5-Why analysis
  2. Fault tree analysis
  3. Fishbone diagram
  4. Barrier analysis
  5. Life-cycle asset review

Example: High MPC and Bearing Temperature Rise

Problem

Bearing temperature increased and MPC is high.

5-Why logic

Why did bearing temperature increase?
Because heat transfer and friction condition at the bearing surface changed.

Why did the bearing surface condition change?
Because varnish-like deposits formed on bearing pads.

Why did deposits form?
Because polar oxidation by-products exceeded oil solubility and deposited on metal surfaces.

Why did oxidation by-products accumulate?
Because antioxidant reserve declined, oil operated at elevated temperature, and varnish precursors were not removed.

Why were they not removed?
Because the plant relied on standard mechanical filtration, which removes particles but not dissolved oxidation by-products.

Why did the asset management system allow this?
Because turbine oil was treated as a consumable and not as a managed lubricated asset with a varnish-control strategy, alarm limits, ownership, and life-cycle plan.

This is a much stronger RCA.


10. Step 7: Interpret Key Failure Mechanisms

A. Oxidation

Oxidation is one of the main turbine oil degradation pathways. It is accelerated by:

  • heat,
  • air,
  • water,
  • metals,
  • contamination,
  • depleted antioxidants,
  • high residence time at elevated temperature,
  • entrained air,
  • microdieseling,
  • poor reservoir design.

Oxidation produces:

  • acids,
  • sludge,
  • varnish precursors,
  • polar degradation products,
  • color change,
  • increased TAN,
  • reduced RPVOT,
  • antioxidant depletion.

RCA question:

“Was oxidation the primary mechanism, or only one contributor?”

B. Varnish formation

Varnish is not just “dirt.”

It is often formed from polar, oil-degradation products that may remain dissolved at high temperature and precipitate when conditions change.

Varnish may deposit on:

  • bearing pads,
  • thrust shoes,
  • servo valves,
  • hydraulic valves,
  • cooler surfaces,
  • reservoir walls,
  • filters,
  • narrow clearances,
  • low-flow areas.

MPC is important, but interpretation must consider trend, sampling temperature, storage conditions, and whether varnish has already deposited. A high MPC may show elevated insolubles or varnish potential, while a sudden MPC decrease without corrective action can sometimes indicate that material has left the oil and deposited on surfaces. (TestOil)

C. Antioxidant depletion

RULER is critical because turbine oil can look visually acceptable while antioxidant reserve is already weak.

RCA questions:

  • Which antioxidant depleted faster: phenolic or aminic?
  • Was depletion linear or sudden?
  • Was the oil topped up and artificially refreshed?
  • Was RULER compared with the new oil baseline?
  • Was RULER interpreted with TAN, MPC, and RPVOT?

D. Water contamination

Water can cause:

  • additive hydrolysis,
  • rust,
  • oxidation acceleration,
  • poor demulsibility,
  • filter plugging,
  • microbial growth in some systems,
  • bearing distress,
  • reduced oil film strength.

RCA questions:

  • Is the water dissolved, emulsified, or free?
  • Is the source cooler leakage, steam seal leakage, condensation, washdown, breathers, storage, or poor handling?
  • Is the oil still able to separate water?
  • Is water a cause or a consequence of poor reservoir management?

E. Air entrainment and foam

Air can cause:

  • oxidation acceleration,
  • unstable oil pressure,
  • poor hydraulic response,
  • cavitation-like effects,
  • microdieseling,
  • varnish acceleration,
  • bearing film disturbance.

RCA questions:

  • Is return oil entering above the oil level?
  • Is reservoir residence time adequate?
  • Is the oil level correct?
  • Are suction leaks present?
  • Is the defoamant depleted or filtered out?
  • Is the oil contaminated with incompatible top-up oil?

F. Wrong oil or incompatible top-up

Many turbine oil problems begin with innocent top-up.

RCA questions:

  • Was the top-up oil the same brand and formulation?
  • Was it from the same product family?
  • Was compatibility tested?
  • Was new oil baseline available?
  • Did additive elements change?
  • Did foam, air release, demulsibility, or MPC change after top-up?

11. Step 8: Identify Root Causes at Three Levels

A professional RCA should classify root causes into three levels.

Level 1: Technical root causes

Examples:

  • oxidation,
  • water contamination,
  • varnish formation,
  • antioxidant depletion,
  • poor air release,
  • high particle contamination,
  • wrong oil,
  • thermal stress,
  • additive incompatibility.

Level 2: Equipment/system root causes

Examples:

  • cooler leak,
  • undersized reservoir,
  • poor return-line design,
  • excessive turbulence,
  • dead zones,
  • poor filtration location,
  • inadequate purification,
  • high bearing temperature,
  • poor breather system,
  • poor drain design.

Level 3: Management-system root causes

Examples:

  • no turbine oil asset strategy,
  • no oil criticality ranking,
  • poor sampling practice,
  • incomplete test slate,
  • no MPC testing,
  • no RULER baseline,
  • no oil life-cycle plan,
  • no ownership,
  • no action limits,
  • no RCA trigger criteria,
  • no follow-up verification,
  • procurement selected oil without reliability input,
  • no ICML 55-style lubrication management system.

The third level is where recurrence prevention happens.


12. Step 9: Define Failure Consequence and Risk

Since ISO 55000 is value-focused, RCA must quantify consequence.

Use a risk matrix.

Example risk categories

ConsequenceExample
Safetytrip system malfunction, fire risk, emergency shutdown
Productionforced outage, derating, failed start
Asset damagebearing damage, servo valve damage, cooler fouling
Costoil replacement, flushing, labor, lost generation
Environmentaloil disposal, leakage
Reputationrepeated reliability failure
Compliancefailure to follow internal asset management procedures

A turbine oil with MPC 45 in a non-critical small auxiliary unit may be a medium risk.

The same MPC 45 in a critical gas turbine hydraulic/control oil system may be a severe risk.

This is why fixed lab limits are not enough.

Risk must be linked to:

  • asset criticality,
  • operating temperature,
  • failure history,
  • component sensitivity,
  • oil volume,
  • redundancy,
  • outage window,
  • replacement cost,
  • and business impact.

13. Step 10: Develop Corrective Actions Using Asset Management Logic

Corrective actions should not be random.

They should be linked to root causes.

Example corrective action table

Root CauseCorrective ActionVerification
High varnish potentialInstall resin-based varnish removal or chemistry management systemMPC reduction trend, patch color, bearing temp stabilization
Water ingressRepair cooler leak or improve sealing/breather systemKF water trend, demulsibility recovery
Antioxidant depletionEvaluate partial/full oil replacement or reconditioning strategyRULER, RPVOT, TAN trend
Poor samplingInstall live-zone sampling point and train techniciansRepeatable lab results
Incomplete oil analysisAdd MPC, RULER, demulsibility, air release, patch microscopyBetter early detection
No ownershipAssign turbine oil asset ownerRCA action closure
Poor filtration strategyUpgrade from reactive filtration to proactive contamination and chemistry controlISO cleanliness, MPC, TAN trends
Wrong top-upCreate approved oil list and compatibility procedureNo unexplained additive/property shifts
Poor reservoir conditionInspect and clean during outageDeposit reduction, improved oil stability

Corrective actions must include:

  1. action owner,
  2. deadline,
  3. technical justification,
  4. risk reduction target,
  5. verification method,
  6. follow-up sampling date,
  7. expected trend,
  8. acceptance criteria.

14. Step 11: Create a Turbine Oil Asset Health Index

For asset management, it is useful to convert lab and operational data into a health index.

Example:

Turbine Oil Health Reliability Index — TOHRI

ParameterWeight
MPC20%
RULER antioxidant reserve20%
TAN15%
RPVOT10%
Water10%
Particle count10%
Demulsibility5%
Air release/foam5%
Operating temperature history5%

Then combine with asset criticality:

Oil Risk Priority = (100 – TOHRI) × Asset Criticality Factor

Example:

  • TOHRI = 62
  • Asset criticality factor = 5
  • Risk priority = (100 – 62) × 5 = 190

This allows management to see oil risk as an asset-management risk, not only as a laboratory number.


15. Step 12: RCA Report Structure

A professional turbine oil RCA report should include:

1. Executive summary

  • what failed,
  • consequence,
  • most probable root cause,
  • risk level,
  • required actions.

2. Asset information

  • turbine tag,
  • OEM,
  • oil type,
  • oil volume,
  • oil age,
  • reservoir volume,
  • bearing type,
  • hydraulic/control system details,
  • criticality.

3. Failure description

  • symptoms,
  • dates,
  • operating conditions,
  • affected components.

4. Oil analysis trend

Include trend charts for:

  • MPC,
  • TAN,
  • RULER,
  • RPVOT,
  • viscosity,
  • water,
  • particle count,
  • demulsibility,
  • foam,
  • air release.

5. Physical evidence

  • photos,
  • patch images,
  • bearing inspection,
  • filter element inspection,
  • reservoir inspection.

6. RCA logic

Use:

  • 5-Why,
  • fishbone,
  • fault tree,
  • barrier analysis.

7. Root cause classification

Separate:

  • direct cause,
  • contributing causes,
  • systemic causes.

8. Corrective action plan

Include owner, target date, verification method.

9. Asset management improvement

Link recommendations to:

  • ISO 55000 asset value,
  • ICML 55 lubrication management,
  • turbine oil life-cycle plan,
  • risk-based monitoring.

10. Lessons learned

Convert the RCA into a plant standard.


16. Practical Example: Varnish-Related Turbine Oil Failure RCA

Event

A steam turbine experienced increasing bearing temperature and unstable control valve response.

Evidence

  • MPC increased from 12 to 38 over 10 months.
  • RULER aminic antioxidant dropped significantly.
  • TAN increased gradually.
  • Bearing temperature increased by 12°C.
  • Servo valve response became slow.
  • Brown deposits found on valve spool.
  • Reservoir wall showed bathtub ring.
  • No resin-based varnish removal system was installed.
  • Oil analysis did not include MPC until after symptoms appeared.
  • Sampling point was from reservoir drain, not live-zone sample.

Direct cause

Deposits formed on bearing and servo valve surfaces, affecting heat transfer and valve movement.

Failure mechanism

Oxidation by-products accumulated in oil, exceeded solubility margin, and deposited as varnish on polar metal surfaces and tight-clearance components.

Contributing causes

  • elevated oil temperature,
  • antioxidant depletion,
  • no varnish-control technology,
  • poor sampling location,
  • incomplete oil-analysis slate,
  • delayed action after early warning signs.

Systemic root cause

The turbine oil was not managed as a critical lubricated asset. The plant lacked a risk-based turbine oil asset management plan aligned with ISO 55000 value realization and ICML 55 lubrication management principles.

Corrective actions

  1. Establish new oil baseline.
  2. Add monthly MPC and RULER for critical turbine oils.
  3. Install live-zone sampling points.
  4. Implement resin-based varnish and acid-removal strategy.
  5. Inspect servo valves and bearing surfaces during next outage.
  6. Define turbine oil alarm and action limits.
  7. Assign turbine oil asset owner.
  8. Create turbine oil life-cycle management plan.
  9. Add RCA trigger criteria.
  10. Review results monthly in reliability meeting.

Verification

  • MPC trend reduced and stabilized.
  • Bearing temperature returned to normal band.
  • Servo valve response improved.
  • TAN stabilized.
  • RULER depletion rate slowed.
  • Filter differential pressure stabilized.
  • No repeat event after defined monitoring period.

17. Common Mistakes in Turbine Oil RCA

Mistake 1: Calling varnish the root cause

Varnish is often the result, not the deepest cause.

Mistake 2: Looking only at the latest sample

Turbine oil RCA requires trends.

Mistake 3: Ignoring sampling quality

A poor sample creates poor diagnosis.

Mistake 4: Treating color as oil health

Oil color is useful but not sufficient. Clear oil can be unhealthy. Dark oil can still be functional depending on other parameters.

Mistake 5: Using particle filtration as varnish control

Mechanical filters remove particles. They do not fully address dissolved oxidation by-products or soluble varnish precursors.

Mistake 6: Ignoring RULER

Antioxidant depletion is often the early warning before TAN and varnish become severe.

Mistake 7: Ignoring operating temperature

Temperature history is essential for oxidation RCA.

Mistake 8: Ignoring top-up history

Top-up can dilute TAN, refresh antioxidants, change additive chemistry, or create compatibility issues.

Mistake 9: No link to business consequence

Without consequence, the RCA remains technical but not managerial.

Mistake 10: No management-system corrective action

Replacing oil without fixing the system only resets the failure clock.


18. MLE-Level RCA Questions

An MLE should ask deeper questions:

  1. What is the oil’s required function in this asset?
  2. What is the asset criticality?
  3. What is the oil’s life-cycle stage?
  4. What is the new oil baseline?
  5. What changed in operation?
  6. What changed in temperature?
  7. What changed in top-up oil?
  8. What changed in filtration?
  9. What changed in contamination exposure?
  10. What changed in bearing temperature or valve response?
  11. Was the sample representative?
  12. Was the test slate complete?
  13. Were the alarms suitable for this criticality?
  14. Was the oil analyzed as a trend?
  15. Was varnish soluble, insoluble, or already deposited?
  16. Was antioxidant depletion understood?
  17. Was water dissolved, emulsified, or free?
  18. Was air entrainment investigated?
  19. Was the oil management plan proactive or reactive?
  20. Was the corrective action verified?
  21. Was the lesson converted into a standard?

19. Final MLE Message

A turbine oil RCA should not finish with:

“Change the oil.”

That is not root cause analysis.

It should finish with:

“This is why the turbine oil asset failed to deliver its required function; this is how the degradation mechanism developed; this is why our management system did not detect or prevent it earlier; this is the risk to the business; and this is the corrected life-cycle asset management plan to prevent recurrence.”

That is the difference between an oil-analysis report and an asset-management RCA.

As an MLE, the target is not only to find the failed property.

The target is to protect the lubricated asset, extend the P-F interval, reduce the risk of functional failure, and maximize the life-cycle value of the turbine oil and the turbomachinery it protects.

In simple words:

Do not investigate turbine oil as a dirty fluid. Investigate it as a failed asset protection system.


Discover more from Turbine Oil Reliability

Subscribe to get the latest posts sent to your email.

Leave a Reply

Discover more from Turbine Oil Reliability

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Turbine Oil Reliability

Subscribe now to keep reading and get access to the full archive.

Continue reading