Logo
Condition Monitoring

Your Vibration Route Runs Monthly. Some Faults Develop in a Week.

A periodic route is a sampling strategy, and the interval is a bet that no machine fails faster than the gap between visits. On one 200 kW fan that bet was one week against two.

September 8, 202610 min readOnline Vibration Monitoring
A fault developing over about one week between vibration rounds spaced four weeks apart, reaching overhaul before the next scheduled reading

Most vibration programs run on an interval. Monthly on the important machines, quarterly on the rest, an engineer with a data collector walking a route that was set years ago and has worked ever since.

That approach is sound, and it catches most of what it is asked to catch. But it carries an assumption that almost nobody writes down: that no machine on the route will develop a fault faster than the gap between two visits. On most assets that assumption holds comfortably. On a few it does not, and those are the machines that produce the failures nobody saw coming.

The business problem

A periodic route is a sampling strategy, and the interval is the bet.

If a fault takes one week to develop and the machine is measured every four, the program is not measuring that machine. It is hoping the fault waits. Nothing is wrong with the instrument or the analyst. The interval is wrong for that specific asset.

Faults do not develop on a convenient schedule

The reason interval-based monitoring works is that most mechanical faults are slow. The reason it occasionally fails badly is that not all of them are.

Days, on the worst ones

A heavily loaded, high-speed bearing can go from the first detectable rise to needing an overhaul inside a week. On the fan described below, the whole deterioration took about that long.

Months, on most of them

Slow-speed shafts, lightly loaded gearboxes and many misalignment or looseness problems develop gradually enough that a quarterly round sees them coming.

Speed changes the answer

Load, speed, temperature and the defect type all move the timeline. The same bearing in two duties does not fail on the same schedule.

Access changes it too

If a machine sits in a hazardous area, behind a guard or at height, the reading is not taken when it should be. It is taken when someone can safely get to it.

This is why the right question is not how often you should measure vibration in general. It is how fast the worst realistic fault can develop on each machine, and whether your interval is shorter than that number. Most maintenance teams know their inspection frequency precisely. Very few have been asked to justify it against fault development time, because it is not a question that routine practice puts in front of them.

What continuous measurement adds beyond more readings

The obvious benefit of an always-on system is that it sees the machine between visits. The more useful benefit is that a capable system does the diagnosis, not just the detection.

It names the defect, not just the level

Envelope analysis, also called demodulation, pulls the repetitive impacts of a bearing defect out of the general noise. Matched against the calculated bearing frequencies, it identifies which element is damaged rather than reporting that vibration is up.

It holds accuracy on variable speed

A machine that changes speed smears any fixed-frequency analysis. A real-time tachometer feeding order tracking keeps the frequency mathematics locked to shaft speed, which is what makes the diagnosis trustworthy on a drive that never sits still.

Current systems also run that analysis at the machine rather than shipping raw data somewhere to be processed. Edge computing matters here for a practical reason rather than a fashionable one: the system keeps working when the network does not, and it can capture a transient at full resolution the moment it happens instead of on the next scheduled upload. Results then move into the plant systems through standard industrial protocols, so condition data sits alongside process data rather than in an isolated tool.

A fan that deteriorated in about a week

A metal works plant ran a 200 kW dust extraction fan at roughly 1,280 rpm on variable speed, with a permanently installed monitoring system on it. The customer is not named in the record.

The system recorded a rapid rise in acceleration on the drive-end fan bearing. Envelope analysis produced frequencies that lined up with the calculated ball-pass frequency for the outer race, and the time waveform showed repeated impacts spaced at that same interval. A real-time tachometer kept the speed reference accurate, which on a variable-speed fan is the difference between a clean match and a smeared one. The diagnosis was specific: an outer-race bearing defect, with the evidence to support it.

Two facts made the result useful. The deterioration took about one week from first detectable rise. And routine data on that machine was roughly two weeks from the next visit.

The value was scheduling, not rescue

A down-day was already imminent, so the fan was overhauled during a stop that was happening anyway. It did not fail, and it was not a dramatic save. Early warning is worth exactly what you can do with it, and here it converted an unplanned repair into a line item on a planned shutdown.

The plant put its own exposure at more than €20,000 per hour if that fan had gone down during production. That figure is the customer's estimate of what the stoppage would have been worth to them, not a measured saving and not a number we are claiming on anyone's behalf. The monitoring on this machine was a Twave T8, running the analysis on the device.

Which machines actually justify it

This is the part that matters commercially, and the candid answer is that most machines do not. Continuous monitoring is not an upgrade every asset deserves. It is the right answer for a specific and usually small group, and four questions identify them.

1

How fast can this machine's worst fault develop?

Not the average fault. The fastest one this asset is capable of. If that number is shorter than your inspection interval, the interval is a gamble on that machine.

2

What does an hour of it being down actually cost?

Production value, not repair cost. A machine that stops a line is a different decision from one that stops a task.

3

Is there a redundant unit?

A standby that can be switched in changes the urgency completely, and usually the answer.

4

Can somebody safely reach it while it runs?

If the honest answer is not really, the readings are already less frequent than the schedule claims, whatever the paperwork says.

Continuous monitoring belongs on the machines where the answers are fast, expensive, no, and not really. Everything else stays on a route, and should. A walk-around program covering a few hundred machines is an efficient use of an analyst's time, and replacing it wholesale with hardware would be an expensive way to make a reliability program worse.

The practical route into this is not a plant-wide rollout. It is to take the list of machines that have caused unplanned downtime in the last two years, run those four questions across them, and see how short the resulting list is. In most plants it is shorter than people expect.

What this means for a plant in Egypt or Saudi Arabia

The reasoning does not change by geography, but the asset list does, and the machines where this argument bites hardest are common across both markets. Cement plants run kiln induced-draught fans, mill fans and separator fans that are central to output and awkward to reach. Steel and metal works run dust and furnace extraction fans in exactly the duty described above. Petrochemical and fertilizer plants run process and exhaust fans alongside critical compressors that cannot be stopped on demand. Power stations run balance-of-plant fans and pumps that rarely get the attention the main train receives.

Two local conditions push the same way. High ambient temperatures and heavy dust loading raise the duty on this kind of equipment, and both make some machines genuinely harder to approach safely while running, which is the fourth question on the list above.

Where to start

  1. Write down your current intervals. Which machines are measured, how often, and when the schedule was last reviewed against anything other than convenience.
  2. Put a fault-speed estimate next to each critical asset. How quickly could the worst credible failure mode develop, given the load, speed and duty. This is an engineering judgement, not a lookup, and it is worth making explicit.
  3. Compare the two columns. Wherever fault speed is shorter than the interval, you have found a machine your program is not actually covering, whatever the completion rate on the route says.
  4. Decide per asset, not per plant. Some of those machines need continuous monitoring. Others need a shorter interval, better access, or a redundant unit. The measurement strategy is the output of the decision, not the start of it.

NATCOM covers both ends of that range, from portable route-based programs through to permanently installed online systems, which is the reason we can argue for asset selection rather than for one approach. The vibration monitoring solutions page sets out what we supply and support.

Source note

Fan case from the Twave T8 case study book. One machine, not a typical result. The hourly figure is the plant's own estimate, not a measured saving.

Frequently asked questions

What is online vibration monitoring?

Permanently installed sensors measure a machine continuously and a local analyser processes the data on site, rather than an engineer visiting with a portable instrument on a schedule. Modern systems run the analysis at the machine and pass results to the plant network or a cloud platform, so trends update between visits rather than at them.

Is online monitoring better than a route-based program?

Neither is better in general terms. They answer different questions. A route is a sampling strategy that works whenever a fault develops more slowly than the gap between visits, and it covers a large population of machines economically. Continuous monitoring earns its place on the specific assets where a fault can develop faster than the interval, where downtime is expensive, or where nobody can safely take a reading while the machine runs.

How often should vibration readings be taken?

The interval should be set by the fastest credible failure mode on that particular asset, not by the calendar. Most programs inherit monthly or quarterly rounds without anyone checking that figure against how quickly the machines involved can actually deteriorate. That check is the useful exercise, and it usually shows that most assets are correctly served and a small number are not.

Can continuous monitoring diagnose a fault or only raise an alarm?

A capable system diagnoses. Envelope analysis separates the repeating impacts of a bearing defect from background noise, and comparing those against the calculated bearing frequencies identifies the damaged element. Order tracking driven by a live tachometer keeps that analysis valid on variable-speed machines. The output is a named defect with evidence behind it, not just a level that has risen.

Which machines justify continuous monitoring?

Four questions settle it: how fast the worst fault on that asset can develop, what an hour of downtime costs, whether a redundant unit exists, and whether anyone can safely take a reading while it runs. Continuous monitoring is for the machines where the answers are fast, expensive, no, and not really. Everything else is well served by a route.

Does this apply to plants in Egypt and Saudi Arabia?

The reasoning is the same anywhere, and the asset types where it bites hardest are common across both markets: cement kiln and mill fans, dust and furnace extraction on steel and metal works, process and exhaust fans in petrochemical and fertilizer plants, and balance-of-plant fans and pumps in power generation. High ambient temperatures and dust loading also make some machines harder to reach safely, which pushes the access question in the same direction.

The interval is a decision, even when nobody makes it

Every monitoring program already has an idea on how fast its machines can fail. On most assets that judgement is right, which is why route-based programs have earned their place.

The exercise worth doing is finding the machines where that judgement is wrong.

Eng. Khairy Arsanios
Eng. Khairy ArsaniosManaging Director, NATCOM
TagsOnline Vibration MonitoringCondition MonitoringRoute-Based VibrationBearing FaultsPredictive Maintenance

Is your monitoring interval shorter than your fastest fault?

Bring the machine list, the current inspection intervals and the unplanned downtime history to NATCOM's reliability team, and we will work through which assets the route actually covers.