Sovereign AI means the data, the compute, and what a model learns from your operation all stay under your control and under the law of the country you operate in. Most of the argument stops at hosting, which is the easy half: procurement can check where the servers sit and a contract can hold it. The harder half is who owns what the model learned: the corrections and judgement your people feed it over months. It is yours only when it exists as a file you can read without the vendor’s software; otherwise it belongs to whoever holds the model. A database moves in a weekend. What a model learned from your team does not move at all unless you hold a copy of it. Picture a maintenance planner six months into a pilot, sitting down every morning with a coffee to correct a model. It flags a sensor fault on a conveyor drive. He marks it wrong, again, because that alarm is never a sensor. It is a conduit that rubs on a guard rail and shorts when the line runs warm. He does that for months. By spring the model has stopped calling it a sensor fault. It has learned how this plant actually behaves. Then the pilot ends and a different vendor wins the work. The plant gets its data back: tags, timestamps, alarm history, all of it, in clean CSVs. What does not come back is the six months of corrections. Somebody has to teach the next model about the conduit from scratch, except the planner has moved to days by then, and nobody wrote it down, because the system was the place it was being written down. What sovereign AI usually means Sovereign AI describes systems where the data, the compute and the rules governing both stay under the control of the organisation running the operation and the jurisdiction it sits in. In practice, most of the conversation is about three things: where the data is stored, where it is processed, and which government could compel access to it. Those are good questions. They are also answerable in an afternoon, which is part of why they dominate. The gap is that all three are about where information is stored. None of them are about what happens to information as it passes through a model and comes out the other side as something new. Where the corrections your team makes end up When your team uses an AI system for a year, they produce two different things. The first is data. Readings, records, transactions, images, notes. This is the part every AI data ownership clause negotiates over, and it is the part you can almost always get back, because it is a file and files are easy to hand over. The second is the correction layer. Every time somebody overrides a recommendation, relabels an event, fixes a bad classification, adjusts a threshold, or tells the system that this alarm on this asset in this season means that, they add a piece of judgement that was not written down anywhere before. A charge nurse does the same thing when she relabels a bed-flow alert that fires every Friday because of a discharge round, not a capacity problem. That is what the model learns from you. It is the difference between a model that knows manufacturing in general and a model that knows your line, your ward, your enrolment cycle, your audit pattern. That learning can end up in three very different places. It can live in a file you own and can read without the vendor’s software. It can live inside a model you rent, which means you can use it and cannot inspect it or take it. Or it can be folded into a shared model that every other customer of that vendor benefits from, including the two competitors down the road. The three look identical in a demo. They are wildly different when you leave. Why the corrections cost more to replace than the data Your raw data is replaceable in the sense that it keeps being produced. Next quarter generates more of it. The corrections do not work that way. They are the output of specific people applying judgement that took them years to build, and once those people move on, that judgement is not reproducible at any price. This is the part of sovereign AI that gets skipped, and it gets skipped because data is easy to talk about and learning is not. Data sits somewhere you can point to. What a model learned from your team sits nowhere you can point to, so it does not show up on an architecture diagram, so it does not show up in the contract, so nobody owns it until the day somebody needs it and finds out who does. Six questions to ask before you teach a vendor’s model Put these to any AI vendor, in writing, before your people start correcting anything. The answers take a vendor about ten minutes to give. What does the system learn from us, specifically? Not “it improves over time.” Name the artifacts: labels, corrections, thresholds, rules, embeddings, fine-tuned weights. Where is that learning written down? If the answer is “in the model,” ask whether it also exists as something a person can open and read. Can we export it, and in what format? Ask for the file names and the schema. A vendor who has thought about this can answer immediately. Does what we teach it improve the version other customers get? There is a defensible yes and a defensible no here. What is not defensible is not knowing. If we leave, what exactly comes with us? Get the list. Then ask what stays behind, which is the more revealing question. Can we run the export today? This is the one that matters. If you have never run the export, you do not know whether it works. Run the last one this week on a system you already use. Ask for a copy of everything your team has taught it, then look at what arrives. If what comes back is your raw data with none of your corrections, you now know exactly where you stand, two years before renewal instead of two weeks before it. What you should get back when you leave a vendor A clean exit is boring. You get your records, and you also get the correction history in a format that reads without the vendor’s software: what was flagged, what a human said instead, when, and on what asset. Somebody at the new vendor can load that on day one and the model starts where the old one stopped. Anything less means starting over. The plant, the hospital, the shop pays again in staff time to teach a new system what the last one already knew, and pays a second time in the months where the recommendations are not trustworthy yet and people quietly stop reading them. Most sovereignty arguments stop at the data centre. What happens to what your people teach the system matters more. Every shift, they are teaching it something, and that teaching is the most valuable thing your operation produces, whether or not anyone writes it down. The only real question is who owns it when they stop. This week: Ask your current AI vendor for a copy of everything your team has taught the system, then look at what actually arrives.