Most AI pilots in manufacturing stop somewhere between the demo and the plant floor. It is almost never the model that killed it. It is that nobody named the decision the model was supposed to change, the person who would make that decision differently, or the number that would move if it worked. The test is plain: if this ran every day for a month with the vendor gone, would anything on the floor be different? If you cannot answer that in a sentence, you do not have a pilot yet. I sat in a room once while a vendor predicted a bearing failure eleven days out. The chart was beautiful. Everyone around the table nodded. The maintenance planner in the corner, the one who would actually have to move the schedule, asked one question: where does this show up in my week? Nobody had an answer for him. The pilot got funded anyway. Eleven months later he was still building the schedule the way he always had, on a spreadsheet he kept himself, and the model was still right about bearings nobody had time to pull. Why AI in manufacturing demos impress the room and change nothing on the floor A demo is built for the room. It runs on an extract somebody spent two weeks cleaning, from one line, on one failure mode, with no changeover in the window and no night crew in the data. There is no shift handover in a demo. Nobody in a demo has to type an extra field into a terminal with gloves on. No sensor drops out. The plant in the demo is the plant on its best day, and every plant I have run had maybe four of those a year. The deeper problem is what the pilot was scoped to prove. Most get scored on accuracy. The model beat its accuracy target, so the pilot went down as a success. But accuracy describes the model. Value describes what changed on the floor. Between the two sits a planner, a supervisor, a scheduler, a quality lead, and every one of them already has a way of doing the thing you are about to change. If the pilot never asked them to do it differently, of course nothing moved. It was never pointed at them. What a demo leaves out The same things are missing from every demo I have sat through: Where the data came from, and how many hands touched it before it went in. Who has to enter something new for this to keep working after month one. What happens when a sensor drops out, or a line runs a product the model has not seen. Who gets the alert at two in the morning, and what they are authorised to do about it. What the operator sees when the model is wrong, and whether they can tell you it was wrong. Which screen this lives on, out of the ones that person already opens every shift. None of those are model problems. All of them are production problems, and production is where the cost actually is. A model that is right and unused has produced nothing. A model that only runs in a demo environment has produced nothing. Value shows up at the point of use or it does not show up. The production test to run before you fund a pilot A pilot is a model running with the vendor in the room, on data someone prepared, scored on accuracy. A production deployment runs on the shift schedule, on live data, and is measured by a number the plant already reports. This takes an hour and one page. Before any money moves, write down five things and make somebody sign their name next to each. The decision. Which specific call changes? Not “better visibility into downtime.” Something like: the planner pulls the bearing during Thursday’s planned stop instead of waiting for it to fail on a Sunday. The person. Name them. A person, not a role. If you cannot name who makes the call differently, the call is not going to change. The moment. Where in that person’s shift does the output land, and in which system that they already open? A new login is one more thing to open, and every one of those means fewer people use it. The number and the baseline. What moves, measured how, starting from what. Write down today’s number, measured the same way you intend to measure it after. Most of the arguments I have sat through about AI results are actually arguments about a baseline nobody wrote down. The date it runs without the vendor. Put a real date on the day the demo team leaves and the thing still runs on a Tuesday. If nobody will commit to a date, everybody in the room already knows the answer. Add a sixth if you want to be thorough: what happens when it is wrong, and who is allowed to say so. A pilot with no path for an operator to say “that alert is garbage” will be quietly ignored by week three, and you will hear about it in month nine. If you cannot fill in all five on one page, you are not evaluating a pilot. You are funding a demo. What it costs when a pilot never reaches production The money is the small part. Set against what the line it was supposed to help is worth, a stalled pilot is a small number. The real cost is what it does to the floor. Operators have long memories for this. Every project that arrives with a kickoff and leaves without a trace teaches the crew that the smart move is to wait it out. Do it three times and the fourth project, the good one, walks into a plant that has already decided not to believe it. You spend the first year of a real deployment earning back the trust the earlier pilots used up. That is the part I would push back hardest on. You are not just spending money on pilots. You are spending the crew’s trust, and trust on a plant floor is a lot harder to get back than budget. The maintenance planner with the spreadsheet is still the test, in every plant I walk into. He is not resisting technology. He is protecting a schedule that keeps his week from coming apart, and he has watched enough pilots come and go to know that protecting it is the rational move. He will change how he works the day something proves it will still be there next quarter. Until then, every beautiful chart is just another thing that showed up once.