The situation
Language and vision models can now do useful planning and perception work in industrial settings: turning an operator’s instruction into a task sequence, spotting defects on a line, scheduling maintenance from sensor history, coordinating supply-chain exceptions. The difficulty is not getting a demo to work. It is connecting a probabilistic model to equipment that can hurt people or damage product, in a way your engineers and safety function can accept.
The most common mistake is architectural: the model ends up, by accident, as part of the control loop. It sends commands directly, or its output is trusted without a deterministic check, or nobody has written down what happens when the network drops mid-task. The second most common mistake is the opposite: the project stalls because nobody can say where the model’s authority ends, so the safety function cannot approve anything.
What the work involves
Separate planning from control from safety. I treat the system as layers with different responsibilities. The AI layer, in the cloud or at the edge, interprets, perceives and proposes. A deterministic layer validates every proposal against limits: workspace bounds, speeds, forces, sequence rules, interlocks. The controller executes. The safety system, which the AI layer cannot touch, stops things when they go wrong. Each layer’s job, inputs, latency budget and failure behaviour is written down.
Decide what runs where. Constrained connectivity is normal in plants and warehouses. Perception and short-horizon decisions usually need to run on edge hardware close to the equipment; heavier planning and fleet analytics can run centrally. The map states what each component needs to keep working when the link fails, and what the equipment does then.
Validate in simulation first. Planners and perception models are tested against a simulator, digital twin or recorded operating data before anything moves. Failure cases are generated deliberately: occluded parts, unexpected objects, out-of-range instructions, stale data. Only behaviour already seen in simulation goes to a supervised hardware trial.
Hand safety dependencies to the people who own safety. The AI layer relies on things it cannot guarantee itself: that guarding is in place, that controller limits are configured, that an emergency stop is independent of software. Those go into a register for your safety owner to confirm or act on.
At Orangewood Labs I led RoboGPT, AutoInspect and an EdgeML platform for industrial robotics, and at Manufactured I built AI agents for supply-chain operations. The layered approach also follows my published work on constraining what AI agents can do in production.
The signature deliverable, illustrated
Illustrative extract from a responsibility map, not taken from a client:
| Layer | Runs on | Decides | Must keep working offline? | On failure |
|---|---|---|---|---|
| Task planner (language model) | Central server | Task sequence from operator request | No | Task not started; operator informed |
| Defect detector (vision model) | Edge device at the cell | Pass, reject or refer to inspector | Yes | Parts referred to manual inspection |
| Action validator | Edge device | Whether a proposed move is inside limits | Yes | Move refused and logged |
| Robot controller | Existing controller | Motion execution | Yes | Existing controller behaviour |
| Safety system | Existing safety-rated hardware | Stop | Yes | Unchanged; outside AI scope |
How acceptance is judged
The scope sets validation criteria in simulation (task success, refusal of out-of-limit proposals, behaviour under each listed failure case) and the conditions for a hardware trial. Your engineering lead accepts the build against those criteria; your safety owner, separately, decides whether a hardware trial may proceed. I do not combine those two decisions.
Ownership and handover
Your engineering team owns the code, models and edge packaging, with version pinning, monitoring and a rollback path. Your safety function owns the dependency register. The runbook covers model updates, revalidation in simulation and how to add a new task to the action interface.
When to choose something else
For private model hosting without physical systems, see private and local LLM deployment. If inference speed or cost on existing hardware is the problem, see AI infrastructure and inference. Other sectors, including defence and dual-use, are on the industries overview.