All articles
OT + AI13 min

AI in Safety-Critical Systems: What IEC-World Engineers Require

A safety system has to be provably correct, and a language model cannot be proven correct in the way IEC 61508 means. That does not rule AI out of the plant. It rules it out of one specific layer, and the boundary is drawable.

AI does not belong inside a safety instrumented function, and it is unlikely to get there soon. Not because the models are not good enough, but because functional safety standards demand a kind of evidence that a probabilistic system cannot produce: deterministic behaviour, a bounded and enumerable set of failure modes, and a verification argument that covers the whole input space. A language model offers none of those. What it does offer is real value one layer out, in the advisory space, and the engineering problem worth solving is where exactly to draw that line and how to enforce it.

That is the whole argument, and I want to make it carefully, because the discussion usually collapses into one of two useless positions. Either AI is dismissed outright by people who have read the standards, or it is waved through by people who have not. Both are wrong, and the second one is dangerous.

What "safety-critical" means in IEC terms

The phrase gets used loosely, so it is worth being precise about what it means to an engineer who works to these standards.

IEC 61508 is the base functional safety standard for electrical, electronic and programmable electronic systems, and IEC 61511 is its process-industry application. The central idea is the safety instrumented function: a specific action that takes the process to a safe state when a specific hazardous condition occurs. Close the valve when the pressure exceeds the limit. Trip the compressor when the vibration exceeds the threshold. Each function is assigned a Safety Integrity Level, SIL 1 through SIL 4, which is a target for how reliably it must perform, expressed as an average probability of failure on demand.

Two structural points follow, and both matter for this discussion.

The safety instrumented system is separate from the basic process control system. The control layer optimises the process; the safety layer protects against it. Keeping them independent is the reason a fault in one does not defeat the other, and it is not a suggestion, it is the architecture.

And the whole approach rests on being able to make an argument in advance about how the function behaves. You perform hazard analysis, allocate a target, design to meet it, verify the design against the specification, validate the installed system, and then prove-test it periodically for its whole operating life. The evidence is the deliverable.

The four requirements AI cannot currently meet

Set aside model quality entirely. Assume a model that is extremely good. It still fails against these four, and each failure is structural rather than a matter of more training.

Determinism. The same input must produce the same output, every time, for the life of the system. A safety function that behaves differently on the same demand cannot be assigned a failure probability, because the thing you are measuring is not stable. Language models are sampled, and even at temperature zero you are relying on implementation details that are not contractual. Worse, the model can be updated underneath you, and an update is a change to safety-related software, which triggers the whole verification cycle again.

An enumerable failure mode set. Safety engineering works by listing how a thing can fail and designing for each case. A valve fails open, fails closed, or sticks. A transmitter fails high, fails low, or freezes. You can write them down. The failure modes of a language model are not enumerable. "Produces a fluent, confident, wrong answer" is not a failure mode you can allocate a rate to, and it is the characteristic failure.

Verification coverage. You have to demonstrate that the implemented logic matches the specification across the input space. For safety logic this is achievable because the logic is small and the inputs are bounded. For a model with an effectively unbounded input space and no readable internal logic, there is no coverage argument to make. Testing a sample of inputs and finding them acceptable is evidence about the sample.

Proof testing. Safety functions are tested at defined intervals to confirm they still work, and the interval is part of the SIL calculation. There is no equivalent procedure that confirms a model still behaves as verified, because there is no specification it is being confirmed against.

IEC-world requirementWhat it demandsWhere a language model stands
Deterministic responseSame input, same output, for the system's lifeSampled output; model updates change behaviour
Enumerable failure modesA list you can allocate failure rates toNot enumerable; the typical failure is a confident wrong answer
Verification coverageImplemented logic provably matches specificationNo readable logic and an unbounded input space
Proof testingPeriodic test confirms continued correct operationNo specification to test against
Change controlAny change re-triggers verification and validationModels are updated by the provider, often silently
Independence from control layerSafety separate from process controlAchievable, and this is the one that points to the answer

Notice the last row. It is the only requirement in the list that AI can satisfy cleanly, and it is the one that tells you where AI goes.

So where does AI actually belong

The independence principle already exists in every plant I have worked in. The safety layer is separate from the control layer, deliberately. The useful move is to treat AI as a third layer, separated from both by the same logic.

That layer is advisory. It reads, it analyses, it explains, it recommends. It does not write to a controller, it does not change a setpoint, and it does not participate in a trip. The boundary is not a policy or a code review convention, it is enforced by what the system is physically able to reach: the AI layer has read access to the data and no write path into control or safety at all.

This is less of a limitation than it sounds, because most of the value in industrial AI is on the read side anyway. Explaining why a compressor discharge temperature rose last Tuesday, finding the three historical events that resemble the one happening now, surfacing the maintenance history that is relevant to an alarm, drafting the first version of a functional specification. None of those require write access, and all of them are work that currently takes an engineer hours.

The architecture I would argue for is simple to state: AI proposes, certified logic disposes. The model can recommend a setpoint. A deterministic, verified, range-checked function decides whether that setpoint is applied, and it applies its own limits regardless of what the model said. If the model recommends something outside the envelope, the envelope wins and the event is logged. The safety of the system then rests entirely on the deterministic component, which is the thing you can actually verify, and the model's contribution is bounded by construction.

This is the same containment instinct behind giving an agent real autonomy inside a constrained space, which I worked through in how the AI agent loop works. The loop is powerful precisely because the boundaries around it are not negotiable by the thing inside them.

Human in the loop is not a control unless you design it as one

"There will be a human reviewing it" is offered constantly as the mitigation that makes AI acceptable in a critical context, and in the form it is usually offered, it is not a control at all.

A human who is shown a fluent recommendation with no supporting evidence, under time pressure, dozens of times per shift, will approve nearly all of them. That is not a failure of the individual, it is the well-documented behaviour of humans in monitoring roles, and it is the same phenomenon as alarm flooding: presented with too many signals of uniform apparent importance, people stop discriminating between them. Anyone who has run an alarm rationalisation programme has watched this happen.

For a human to function as a real control, three things have to be true.

The reviewer must be able to see the evidence, not just the conclusion. Which tags, which time window, which values, which documents. An answer traceable back to a tag and a timestamp can be checked in seconds; a fluent paragraph cannot be checked at all.

The system must state its uncertainty honestly, including when data was interpolated, when quality codes were bad, and when it found nothing relevant. A system that never says "I do not know" trains the reviewer to stop reading.

And the review must be rare enough to be real. If the operator is approving forty AI recommendations a shift, the review is theatre. Route the routine cases through deterministic logic and reserve human judgement for the cases that genuinely need it.

The audit trail is not optional, and it is not a log file

In regulated manufacturing this is where AI projects quietly die, and it is worth being concrete because the requirements are specific.

In GMP environments, records must satisfy ALCOA principles: attributable, legible, contemporaneous, original, accurate. Applied to an AI-assisted decision, that means you have to be able to answer, potentially years later during an inspection, what the system recommended, what evidence it used, which model version produced it, who reviewed it, and what was actually done. "The AI suggested it" is not a record.

In eighteen years I have watched plenty of good engineering fail validation not because it was wrong but because it could not be evidenced. An AI layer has to be designed for that from the start: version the model and the prompt, store the retrieved evidence alongside the output, capture the reviewer's identity and decision, and make the whole record immutable and queryable. Retrofitting this is close to impossible, because the evidence you needed was never captured.

There is a related trap worth naming. A system that accumulates knowledge across sessions is subject to the same discipline: what it retained, when, and on what basis all become part of the record. I went through the mechanics of that in long-term memory for AI agents, and the governance conclusion there applies with much more force here.

What I would actually build

Concretely, for a plant that wants to use AI without touching the safety case:

Start with a read-only data path. The AI layer subscribes to plant data through the structured hub rather than connecting into control systems, which is one of the strongest practical arguments for the architecture in OPC UA, MQTT and the Unified Namespace: the subscription is outbound, one-directional, and there is no write path to remove later because there never was one.

Build the evidence-first retrieval underneath it, so every answer cites tags, timestamps, values and documents. This is the query-first pattern rather than the naive one, for the reasons in industrial RAG.

Put a deterministic envelope around anything that could ever influence an action, and make that envelope the only thing that writes. Verify the envelope to the standard the action requires. Do not verify the model, because you cannot.

Capture the record at the point of decision, not afterwards.

And write down, in one page, what the system may never do. Not as a prompt instruction, which is a request rather than a constraint, but as an architectural statement about what it cannot reach. If the only thing preventing an AI system from writing to a controller is an instruction in its prompt, it is not constrained, it is asked nicely.

The bottom line

The IEC-world objection to AI in safety functions is not conservatism and it is not a failure to understand the technology. It is that functional safety is a discipline built on producing evidence in advance, and a probabilistic system with an unbounded input space cannot produce that evidence. That is a structural mismatch, not a maturity gap, and engineers who work to these standards are right to hold the line.

They are also right that this leaves an enormous amount of useful work available. Almost everything an engineer actually wants from AI on a plant is analysis, explanation, retrieval and drafting, and none of it requires write access to anything. Draw the boundary at the write path, enforce it architecturally rather than procedurally, make every answer carry its evidence, and capture the record as you go. Do that and you get the value without touching the safety case, which is the only version of this that survives an audit.

Written by Usman Nasir — control systems engineer, Stockholm.