All articles
OT + AI12 min

Anomaly Detection on PLC Time-Series Data: Why Most False Positives Happen

Most false positives in industrial anomaly detection are not model failures. The model correctly found a statistical deviation that an engineer would have recognised instantly as a normal change of state. Here is how to fix that.

Most false positives in plant anomaly detection are not the model being wrong. The model found a genuine statistical deviation, and an engineer looking at the same chart would have said "that is a grade change" in about two seconds. The gap is context, not accuracy: the model was never told what operating state the equipment was in, so it cannot distinguish a fault from an intended transition. Fix the context problem and a large share of the false alerts disappear without touching the algorithm.

This is the single most common reason these projects get switched off six weeks after go-live. The pilot looked excellent on historical data, the production system produced forty alerts a day, and operators learned to ignore it. I have watched that sequence more than once, and it is almost never a modelling problem.

What "anomaly" means, and why that is the trouble

An anomaly, statistically, is an observation that is unlikely under the model you fitted to the data. That definition is precise and it is the one the algorithm implements.

An anomaly, operationally, is something happening that should not be happening. That definition is the one the operator cares about.

These overlap, but nowhere near as much as people assume. Starting a line is statistically extraordinary and operationally routine. A bearing degrading over four months is statistically unremarkable at any single moment and operationally the thing you most wanted to catch. Every false positive and every miss lives in the gap between those two definitions, and closing that gap is the actual engineering work.

The real causes, in the order I would check them

Operating state is ignored. This is the big one. A process is not one thing: it has modes. Running, starting, stopping, idling, cleaning, changed over to a different product, running at reduced rate because upstream is down. Each mode has its own normal. Fit one model across all of them and it learns a blurred average that describes none of them, then flags every transition. The fix is not a better algorithm, it is joining the process data to the machine state and modelling each regime separately.

Setpoint changes are treated as process changes. An operator moves a setpoint, the process follows, and the model reports that the process deviated. It did, exactly as intended. Any tag that tracks a setpoint should be modelled as the error between them rather than as the raw value, otherwise you are detecting the operator.

The training window is contaminated. "Normal" is usually defined as some historical period, and that period contains the faults nobody logged, the two weeks the transmitter was drifting, and the month the line ran at reduced rate. Train on that and the model learns those conditions as normal, which suppresses real detections and shifts the baseline. Curating the training window against the maintenance record is unglamorous and it is the highest-value hour you will spend.

Compression artefacts are read as signal. Historians filter twice before archiving, by exception at the interface and by compression on the archive, and what comes back may be reconstructed by interpolation rather than measured. A tag with a wide deadband produces flat stretches followed by steps, and a model reading those steps sees change points that never occurred in the process. This is the same trap that breaks retrieval, and I went through the mechanics in why RAG on historian data fails. For detection work, pull recorded values rather than interpolated ones, and know the deadband per tag before you trust a result.

Quality codes are discarded. A failed thermocouple whose last value is being held looks like a perfectly stable process. Not only does the fault go undetected, the frozen value drags the baseline and makes the following genuine excursion look smaller than it was. Carry quality through as a first-class field and exclude bad-quality windows from both training and scoring.

Tags are resampled carelessly. A pressure loop at one second and a lab result once a batch cannot be aligned by pairing rows. Forward-filling the slow tag invents a signal that holds steady and then jumps, and a multivariate model will happily learn correlations that are artefacts of the resampling rather than properties of the process. Resample explicitly, record the method, and be suspicious of any correlation that appears only at one resampling choice.

Ambient and seasonal effects are unmodelled. Cooling water temperature in July is not cooling water temperature in January. A model trained in one season flags the next one. Either include the ambient driver as an input or retrain on a window long enough to contain the cycle.

Everything is univariate. Single-tag detection is easy to build and produces the most noise, because any individual tag moves for many reasons. The informative signal is usually in the relationship: discharge temperature rising is ordinary, discharge temperature rising while flow and load are steady is not. Multivariate methods cost more to build and cut false positives sharply, because holding the other variables constant is exactly what an engineer does mentally.

The threshold was chosen without an alert budget. A sensitivity that produces forty alerts a day guarantees the system gets ignored, whatever its accuracy. I spent a long stretch of my career on alarm management and rationalisation, and the lesson transfers directly: an alert that nobody acts on is worse than no alert, because it trains people to dismiss the channel. Decide first how many alerts per day a human can genuinely investigate, then set the threshold to fit that budget, then work on precision inside it.

SymptomUsual root causeFix
Alerts cluster at startups and shutdownsOne model across all operating modesJoin to machine state, model each regime separately
Alerts follow operator actionsRaw value modelled instead of setpoint errorModel the deviation from setpoint
Real faults are missed, baseline looks wideTraining window contains unlogged faultsCurate the window against maintenance records
Step changes that no one witnessedCompression or interpolated retrievalQuery recorded values; check deadband per tag
A tag is suspiciously stable during a faultQuality codes dropped at ingestCarry quality through; exclude bad windows
Correlations that engineers do not recogniseResampling artefacts across mismatched ratesExplicit resampling; test stability across methods
Alerts reappear every seasonAmbient drivers not modelledAdd ambient inputs or extend the training window
High volume of individually plausible alertsUnivariate detection, threshold with no budgetMultivariate model, threshold set to an alert budget

What to build instead

The pattern that works is less about the detector and more about what surrounds it.

Make state a first-class input. Before any modelling, get a reliable signal of what the equipment was doing at each moment. If it does not exist, deriving it from motor current, line speed and product code is usually possible and is worth the effort, because every downstream decision depends on it.

Model per regime, and say which regime you are in. An alert that reads "abnormal during steady running" is actionable. An alert that reads "abnormal" is not.

Require persistence. A single sample outside the envelope is noise. Requiring a deviation to hold for a defined period, or to accumulate a defined area outside the band, removes a large fraction of nuisance alerts at almost no cost to genuine detection, because real faults persist and noise does not.

Rank, do not classify. Binary anomaly or not anomaly forces a threshold decision on a continuous quantity. Producing a ranked list, most unusual first, lets the alert budget be a queue depth rather than a cutoff, and it degrades gracefully when the process changes.

Give every alert its evidence. Which tags contributed, over what window, with what values, and what the same period looked like historically. An alert that can be checked in thirty seconds gets checked. One that cannot gets dismissed, and after enough dismissals the system is dead regardless of how good it is.

Close the loop. Every alert needs a disposition: real, explained, nuisance. Without that feedback you cannot measure precision, you cannot retrain, and you have no way to tell whether the system improved. Most industrial deployments skip this and then have no answer when someone asks whether it is working.

Measure it honestly

Accuracy is meaningless here. Anomalies are rare, so a detector that never fires scores above 99 percent and is useless.

The metrics that matter are precision at your alert budget (of the alerts a human actually investigated, what fraction were real), detection lead time (how long before the failure or the trip did the system flag it, because catching it four hours out is worth far more than four minutes), and missed events reconstructed from the maintenance record after the fact.

That last one requires going back through work orders and asking whether the system flagged each real failure. It is tedious and it is the only way to know your miss rate. The maintenance record is the ground truth you already own, which is part of why it is such a valuable text corpus in its own right, as I covered in RAG on maintenance logs and work orders.

Where the data layer comes in

Nearly every root cause in the list above is a data problem rather than a modelling one: state not available, quality dropped, timestamps unreliable, sampling rates mismatched, context stripped at ingest. That is not a coincidence.

A plant that publishes structured, contextualised, real-time data with quality and source timestamps intact makes most of these problems small. A plant that exposes numbered tags through bespoke connections makes every one of them expensive, and you pay that cost again for every new use case. This is the practical return on the architecture I described in OPC UA, MQTT and the Unified Namespace, it is why the pipeline decisions in OT data to cloud architecture determine what detection is even possible, and it is the reason the storage decision in historian vs time-series database vs data lake matters more than it appears to at the time you make it.

The same dependency runs the other way too. Once detection is producing trustworthy events, those events become context for everything else the plant asks of its data, which is the retrieval problem I worked through in industrial RAG.

If you are being asked to build detection on a plant where none of that exists, build the state signal and the quality handling first anyway. It will feel like a detour and it is the shortest path.

The bottom line

The reason industrial anomaly detection gets abandoned is almost never that the algorithm was too weak. It is that the system did not know what the plant was doing, so it reported normal transitions as faults, produced more alerts than anyone could investigate, and taught its users to ignore it within a month.

Give the model operating state, clean its training window against the maintenance record, keep quality codes, model relationships rather than single tags, require persistence, rank instead of classify, and set the threshold from an alert budget rather than from a statistic. Attach evidence to every alert and record a disposition for every one. None of that is exotic, and together it is the difference between a detector that engineers trust and one that gets muted in week six.

Written by Usman Nasir — control systems engineer, Stockholm.