A predictive maintenance AI model can learn a plant’s recurring problem so well that it starts treating the problem as normal.
Consider a hypothetical control valve that has been sticking for months. Operators compensate, the process continues, and the historian records the resulting oscillation. Train an anomaly detector on that period without understanding the valve’s condition, and the model may accept degraded behavior as its baseline. The dashboard can look reassuring while the underlying defect remains.
This is where instrumentation judgment matters. A historian contains measurements of plant behavior, shaped by sensors, configuration, and operating decisions. Calling that history ‘normal operation’ requires evidence.
In April 2026, NIST launched work on a Trustworthy AI in Critical Infrastructure Profile, including applications involving operational technology and industrial control systems. The profile remains under development. For plant leaders considering AI, it is a timely reason to make the conditions for trusting a prediction explicit. NIST project
Predictive maintenance AI needs an instrument history
Before accepting a model’s baseline, I would want to know what changed during the training period. Was the transmitter recalibrated? Did its measurement range change? Was a control loop left in manual? Did maintenance replace the sensor while retaining the same tag name?
Those events can change the meaning of a trend. A step change may reflect a corrected measurement rather than deteriorating equipment. A flat signal may indicate a stale value. A stable process variable may conceal the effort operators are making to hold it there.
The practical starting point is a small group of consequential assets. Reconcile their instrument records, operating modes, maintenance findings, and historian tags. Mark known defects and interventions. Give the analytics team a history they can interpret before asking them to predict its future.

Test the decision the prediction will support
NIST’s AI Risk Management Framework emphasizes evaluation under conditions representative of intended use. In a plant, that means asking where a model has actually been evaluated: steady production, reduced rates, changing feed, or restart. Performance in one condition does not establish performance in all of them. NIST AI RMF
For a maintenance advisory pilot, I would ask the team to define the decision first. Is an alert supposed to trigger an inspection, a review of operating conditions, or a planned repair? Each requires different evidence and a different amount of useful warning.
Then evaluate the model against records it has not already learned, keeping later information out of earlier predictions. Count missed events and false alerts, and assess whether the warning arrived early enough for the intended action. Compare that result with the plant’s existing monitoring practice. A model that improves a statistical score but adds no useful decision may add little operational value.
Make uncertainty visible to the people using it
An engineer reviewing an alert should be able to see the supporting measurements, the operating condition, and important gaps in the evidence. If a key sensor becomes unreliable or the plant moves outside evaluated conditions, the advisory should indicate that its output needs additional scrutiny.
Keep responsibility for action clear. An advisory pilot should have defined human review and escalation, with any proposed change to control or protection functions assessed through the plant’s engineering and change-management processes.
For me, the useful question is whether the system helps someone make a better maintenance decision with enough time to act. That is a result a plant can evaluate. It begins with knowing what the measurements actually mean.

