Insights from Five Experts on Long-Cycle Reliable Operation Management of Pumps (Part 2)
By Nick Li · August 10, 2026 · Technical Articles

Source: Equipment Professionals’ Perspective (WeChat: Equipment Horizon) | Translated & Polished

Figure 1: Industrial pumps and compressors with digital condition monitoring in a petrochemical plant
Chapter 3: The Operation and Maintenance Strategy Debate — Breaking the “Three-Year Major Overhaul” Curse
Xiao Hong: “We have roughly covered the selection issues. Now let us move into operational aspects. Some say equipment follows a three-year major overhaul cycle, while others argue against periodic overhauls, claiming that the three-year rule is an arbitrary mandate. Still others contend that if you do not follow the cycle and something goes wrong, who takes responsibility? Let us explore this topic in greater depth today.”
“In actual production, including inspection rounds, condition monitoring, and data analysis, are there successful experiences where monitoring data analysis helped avoid major accidents? Conversely, if the monitoring system alarms but no one dares to shut down — because stopping could cause an entire system or plant shutdown — and ultimately something serious happens because of it. This organizational barrier — fear of accountability and fear of production impact — how do we break through it? Let us first ask Director Wang to analyze this.”
3.1 Successful Practices in Condition Monitoring and the Shutdown Decision Dilemma
Wang Panfeng first shared a case where condition monitoring successfully prevented a major accident. A certain company’s circulating water system pump was a BB1-type water pump handling circulating water media with a flow rate of approximately 10,000 m³/h, supported by rolling element bearings. The condition monitoring system detected a rising trend in bearing temperature, which was consistent with on-site bearing temperature measurements. An abnormal noise was also audible at the bearing location, yet the pump vibration values showed no significant change.
The pump bearing was equipped with an automatic grease lubricator with a scheduled and quantified lubrication plan. However, such lubrication devices frequently experienced operational failures due to battery issues, electromagnetic interference, or environmental factors in practice, making the effective grease delivery unreliable. Manual inspection and supplementary lubrication or replacement were required to ensure effective bearing lubrication.
For the aforementioned bearing temperature rise issue, the equipment management team immediately coordinated with operations personnel to replenish grease promptly. However, the bearing temperature did not show a noticeable decrease; the abnormal noise diminished somewhat, and the pump vibration values and flow rate remained normal. The unit decided to continue monitoring the pump’s operating condition.
Subsequently, the pump continued operating for half a day, during which the bearing temperature continued to climb, and the abnormal noise persisted. Considering that the circulating water system affects the stable production of the entire plant, the unit decisively chose to shut down for inspection. Upon opening the bearing end cover, water ingress was found in the bearing housing. The grease had emulsified and deteriorated, resulting in bearing lubrication failure. Pitting of varying degrees was observed on the inner and outer ring raceways. New bearings and grease were replaced, and upon restart, all parameters including temperature, vibration, and flow rate operated normally. The timely handling of this bearing temperature rise issue prevented disruption to the entire plant’s stable production.
The value of this case lies in revealing the multi-dimensional interpretation of condition monitoring — one cannot ignore temperature and acoustic anomaly signals simply because vibration values are normal. It was precisely because the circulating water system involved the entire plant that the shutdown decision could be made decisively. However, when this logic is extended to other scenarios — particularly during high-load production periods where output is prioritized — shutdown decisions become exceptionally difficult.

Figure 2: Rolling element bearing cross-section showing grease lubrication system and water ingress damage with raceway pitting
3.2 The Path to Breaking Through: Sun Hongjun’s Four-Step Systematic Approach
Sun Hongjun provided a systematic exposition on condition monitoring and the shutdown decision dilemma, proposing a progressive four-step framework. He noted that the application of condition monitoring in the process industry is showing an accelerating trend of popularization — from the early days when only large rotating equipment trains were equipped with shaft vibration, shaft displacement, and bearing pad temperature monitoring, to the current state where heavy-duty centrifugal pumps, large process compressors, and even reciprocating compressors have begun to be equipped with various online monitoring sensors.
However, the widespread adoption of technical means has not fundamentally resolved the core contradictions described above. The inability to quantify the period over which a latent defect degrades into a failure, and the use of shutdown losses as the primary decision-making metric, remain universal management challenges faced by all equipment management personnel.
If equipment has a latent hazard but is not promptly shut down for remediation, and a major safety incident results, how should the relevant responsibility be defined? Conversely, if excessive overhaul is carried out blindly, resulting in dual losses of production output and maintenance costs, who should bear the corresponding responsibility? These two types of responsibility definition dilemmas cause decision-makers at all levels to be overly cautious and conservative in equipment management.
To address this, he summarized a four-step breakthrough approach:
- Step 1: Introduce condition monitoring as objective evidence. By introducing condition data into the decision chain, subjective decisions based on personal experience can be transformed into reference points based on objective information, reducing the risk of individuals bearing full responsibility for decision errors. As he stated: “Introducing it allows the machine to speak from a third-party perspective to the leadership, explaining the current operating state of the equipment and whether it should be stopped, thereby reducing individual accountability.”
- Step 2: Elevate the technical management capabilities of operations and maintenance personnel. With the continuous expansion of plant scale and the proliferation of chemical projects, the pace of O&M personnel development is disproportionate to industry demand. Without mature O&M personnel, if no one can accurately interpret the data output by condition monitoring systems, its value remains unrealized. “The key is to improve the management level of technical personnel — the ability to understand the data, present feasible solutions to leadership, and make more accurate judgments.”
- Step 3: Establish a reasonable shared accountability mechanism. He pointed directly to the breakthrough point: many enterprises need to handle accountability reasonably — “When everyone sits together and discusses a result, the responsibility must ultimately be shared.” Rather than waiting until equipment fails, when frontline technicians provide judgments, they often must weigh professional opinions against other considerations. Shared accountability transforms an individual’s burden into a team’s responsibility.
- Step 4: Improve the shutdown standards system. He advocated pre-emptively cataloging common issues before plant startup, developing tiered shutdown standards for equipment handling flammable/explosive media or with histories of problems — specifying which conditions require immediate shutdown, which allow monitored operation, and which require no action. These standards should be documented as checklists endorsed by all parties as the common reference for subsequent decisions. “Through these four aspects, I believe a comprehensive system can be established that allows those who understand the technology to speak up with confidence.”

Figure 3: Condition monitoring dashboard showing vibration spectrum analysis, bearing temperature trends, and pump performance data
3.3 The Scientific Basis of Overhaul Cycles and Selective Strategies
Xiao Hong: “For many enterprises, the three-year major overhaul may be a hard line that no one dares to break. But with modern condition monitoring, if a pump is operating in excellent condition, are there cases where it has been kept in service beyond three years? If extended, who signs off and who bears the risk? Is there a complete decision-making mechanism? I hope Director Zhang can share from the oil and gas field perspective.”
Zhang Jie provided an in-depth analysis from the perspective of someone who has personally experienced the industry’s evolution. He recalled the progression of overhaul cycles: from early refining enterprises pioneering three-year cycles, to later pushing for four-year and five-year cycles, with some oil and gas field enterprises following suit.
However, he candidly stated: “Whether you define it as a three-year or four-year cycle — honestly, from a technical and data perspective, this basis has always been a subject of repeated debate at expert committee meetings.” He further explained that overhaul cycle determination is actually the result of weighing multiple factors including regulatory inspection requirements, historical failure statistics, peer industry operating experience, and production-maintenance windows — not something that can be calculated by a single technical formula.
Zhang Jie believed that the key to solving the overhaul cycle challenge lies in developing differentiated preventive maintenance strategies based on equipment criticality classification. Following the conventional approach to equipment classification management, equipment can be categorized into three types:
- Class A Equipment: Affects the safe and stable operation of the plant; failure would cause significant losses. These are critical equipment such as compressor trains and special valves.
- Class B Equipment: Secondary importance; moderate impact on production.
- Class C Equipment: Relatively minor impact on production.
For B and C class equipment, he advocated establishing a normalized preventive maintenance mechanism, even achieving “no major overhaul during major overhaul windows” — performing maintenance during daily operation supported by condition monitoring, rather than accumulating all maintenance tasks for the overhaul window. “I can let it undergo condition-based maintenance.”
However, for Class A equipment, even if condition monitoring detects no anomalies, he still recommended adhering to disassembly and overhaul within specific cycles. “For our compressor units, special valves, or other critical equipment, many companies know from experience that even though condition monitoring can detect many problems, we must also clearly recognize that condition monitoring has its technical boundaries — minute cracks on impellers, fatigue accumulation in materials under alternating stress — these defects may not be detectable through monitoring means.” This clear-eyed recognition of hidden defects stems from the accumulation of long-term practical experience.
He simultaneously pointed out that the effective implementation of this strategy faces multiple practical obstacles:
- Prolonged talent development cycles: The analysis and interpretation of condition monitoring data, particularly spectrum analysis and vibration analyst-level professional capabilities, require lengthy development periods. Whether in Sinopec or PetroChina, sufficient technical forces cannot yet be said to have been accumulated.
- Management system execution drift: “When you set up so many management processes, are more processes necessarily better? At the execution end, we frequently see that many management measures have been taken, but actual results may deviate from the original intent.”
- The gap between technological innovation and implementation: Zhang Jie pointed out that condition monitoring has incorporated many new technologies, including AI, but how innovative solutions are introduced and implemented involves significant challenges, compounded by responsibility definition issues.
- Weak data governance foundations: Many enterprises have not yet achieved adequate data management, with shortcomings in data accumulation and organization, and slow progress in informatization and digitalization. “If your data chain is incomplete, where does your artificial intelligence manifest?” he questioned pointedly.
Zhang Jie concluded: “Often we cannot fully rely on these products and technologies — there is still a long road ahead. Returning to the relationship between major overhaul and condition monitoring — my view is that we can proceed step by step with exploration and accumulation on Class B and C equipment, but for critical equipment that affects the safe and stable operation of the plant, this bottom line must be maintained. I do not recommend using monitoring data to challenge its mandatory overhaul cycle.”
3.4 Extended Maintenance Evaluation Mechanism in Fine Chemicals
Meng Fanyu approached the topic from the perspective of fine chemical industry practice, introducing the extended maintenance evaluation mechanism under the equipment integrity management system. “For equipment extended maintenance, this is actually one of the elements within our equipment integrity management system.” He explained in detail that for critical equipment, the team establishes initial data archives after new installation or maintenance, and continuously benchmarks the data during operation.
If the equipment reaches its maintenance cycle but operating data shows no anomalous changes, the equipment management department leads a joint assessment with production, safety, and other departments to evaluate the operating condition. Only when all parties unanimously agree is an extended maintenance plan executed.
However, he specifically emphasized that for “only-child” type critical equipment (single-unit, non-redundant installations), extended maintenance is not recommended — these should be synchronized with plant major overhauls. An unplanned shutdown of such equipment would cause enormous losses. For equipment with backup units that can be switched in, extended evaluation may be permitted provided risks are controlled. This approach embodies a pragmatic management philosophy that seeks a balance point between risk and benefit.
(To be continued)
Source: Equipment Horizon (WeChat Official Account) | www.dlseals.com equivalent industry source