An AI agent on a plant floor should be allowed to do exactly what the person whose job it does is allowed to do, and nothing wider. The interesting question about AI agents in manufacturing is not how capable the model is, it is which actions the agent can take on its own, which ones stop at a human, and whether every one of them leaves a record somebody could hand to an auditor without a week of preparation. Write those three things down and autonomy becomes an ordinary operations decision. Skip them and you have software taking actions nobody authorised. I spent an afternoon in a design review with a quality manager who had already decided she did not trust any of this. We were working through what an agent could do with a batch on hold. I said it could clear the hold once the retest came back clean. She said: on paper I clear that hold, and if I clear it wrong my name is on the deviation and I answer for it in the audit. Who answers when your software clears it? It was the right question and I did not have a clean answer that day. We changed the design instead. The agent gathers the retest, assembles the evidence, and stages the release. She still presses the button. The three tests an action has to pass Three tests, and an action has to pass all three before it runs without a person in the loop. Is it reversible? Re-sequencing tomorrow’s run is reversible. Releasing a lot to a customer is not. Reversibility is the cheapest thing you can check for, and most pilots never check it. Does the authority already exist? Somebody in the building is allowed to do this today, under a named role, with a defined scope. An agent inherits that role. It should never hold a permission no human holds, and it should never act outside the facility, line, or shift its role covers. Does it record itself? The action, the trigger, the role it acted under, the time, and the state of the operation when it happened. Written as it happens, without anyone remembering to write it. Anything that fails one of the three goes into a queue and waits for a person. That is not a limitation to design around. It is the design. Decide who is allowed to do what before AI agents in manufacturing act alone Most operations manage authority badly, and the agent conversation is the first thing that makes it obvious. Shared credentials on a floor terminal. A permissions spreadsheet that is one reorganisation behind. A contractor who finished in March and still has an account in October. Everyone knows who is really allowed to approve a release, and none of it is written anywhere a system can read. You cannot delegate authority you have not defined. So the first work is not the model. Deciding what an agent may do starts as a list of which role can take which action: who can approve a deviation, who can change a spec, who can raise a purchase order over a threshold, who can override a hold, and in which facility, on which line, during which shift. Boring work. It also happens to be the same map an auditor asks for, so the work is not wasted either way. Once that map exists, the agent question is easy to answer. An agent acts under a role. The role has a boundary. The agent stops at the boundary and asks. Why every agent action needs a record Demos rarely show the record, because the record is not interesting until something goes wrong. Then it is the only thing that matters. An investigation into an agent’s decision should take the same time as an investigation into a person’s, and produce the same kind of evidence. What happened, when, under whose authority, what the agent saw, what it decided, what it did next, and who signed off. If your answer to any of that is “we can probably reconstruct it from the logs,” you do not have a record. You have logs that somebody will have to interpret, and interpretations get argued. The reason to insist on this now rather than later is that a record cannot be added retroactively. The day you need it is the day it either exists or does not. Write down who holds each decision before you set up the agent Here is something you can do this week without buying anything. Take one decision you would genuinely like an agent to make. One. Then fill in six columns on a single page: The decision. Stated as an action, not an area. “Re-sequence the afternoon run when a line goes down,” not “scheduling.” Who holds it today. A named role, and the scope of that role. What the agent may do alone. The subset that passes all three tests. What stops at a person. The rest, and which role that person holds. What gets recorded. The fields, and where they land. How you reverse it. In minutes, by whom, with what side effects. Do it for one decision and you will learn more about your readiness than any vendor evaluation will tell you. In every review I have sat in, the hard column is the second one. Nobody can say cleanly who holds the authority today, because it has always lived in a person rather than a system. That gap is the actual project. The agent is the easy part afterwards. The plants that will get real value out of this are not the ones with the best models. They are the ones that can already say, on paper, who is allowed to do what. That work is unglamorous and it is sitting in front of you right now, whether or not you ever buy a piece of software. The alternative is finding out during an audit that a decision was made by something with no name, no role, and no signature.