You do not need a data lake to get SCADA data off the plant floor. You need one line, one gateway, a protocol your equipment already speaks, and about forty tags chosen because they answer a question somebody is actually asking on Monday morning. It turns into a year-long project the moment you try to collect everything, from everywhere, before anyone has agreed what the data is for. Start narrow, prove it in a week, then widen. The last time I did this properly it started in an electrical room with a folding chair, a laptop and a controls technician who had run that line for nine years. He pulled up the tag list. Names like PLC1_DB44_INT12. I asked what that one was. He said, without looking, that it is the infeed conveyor, that a one means it has stopped, and that if it sits at one for more than about ninety seconds the packaging end will run short before the shift ends. None of that was written down anywhere. It was in him, and the export file we were about to take off that machine contained exactly none of it. What is SCADA data? SCADA data is the stream of timestamped tag values a supervisory control system reads from PLCs and other controllers: machine states, counts, temperatures, pressures, and alarm events. PLC data starts as registers. A SCADA system polls those registers, gives them tag names, draws them on a screen and raises alarms. A historian keeps them for years. That is the whole stack, and what comes out of it is a long list of readings with timestamps attached. Readings are true and, on their own, useless. “Tag 4021 went to 1 at 06:14:22” tells you nothing until somebody says which machine that is, what a one means, and what happens downstream when it stays there. Every project that stalls in month four stalls on that gap. The data moved fine. What it meant stayed with the people on the floor. Which protocol should you use to get SCADA data off the floor? Take them in this order and stop at the first one your equipment can do. OPC UA if you already have it. Modern control systems and a lot of newer equipment speak it. It authenticates, it encrypts, and the tags describe themselves, which saves you a mapping spreadsheet nobody will maintain. MQTT with Sparkplug B if you can add a gateway. Devices publish changes outward to a broker instead of being polled. The plant opens the connection, so the firewall rule is outbound only. And devices report by exception, so a tag that has not moved sends nothing. Modbus TCP, EtherNet/IP or serial for older gear. Poll it through an off-the-shelf gateway. This is fine. It is boring and it works. Do not write the driver yourself, because you will own it forever and the person who wrote it will leave. Avoid OPC DA across a firewall. It runs on DCOM, it needs ports nobody wants open, and it has cost more weeks of good engineers’ time than anything else on this page. One opinion, held firmly: publish beats poll whenever you have the choice. Polling means an outside system opens a connection into the plant. Publishing means the plant opens the connection and decides what goes out. What to collect first, and what to leave Pick one line. Pick one question the plant already argues about. Then collect only what answers it. Machine state, timestamped at every change rather than averaged over five minutes Counts at the boundaries the plant already counts: infeed, good, reject The two or three process variables that decide whether product passes Alarms as events, with code, text, start time and clear time Product or order identity, plus changeover start and end The downtime reason code, if operators enter one Leave the rest. Specifically, leave anything sampled at 100 milliseconds that you cannot justify in one sentence, leave values the HMI calculates from other values (take the inputs and recompute later), leave screen navigation and scan diagnostics, and leave any tag whose name means nothing to a person who did not build the line. Two habits save months. Timestamp at the source, because a timestamp applied when the record lands tells you when your network was busy, not when the machine stopped. And name at the source, so the mapping from PLC1_DB44_INT12 to “infeed conveyor stopped” travels with the data instead of living in a spreadsheet on somebody’s desktop. How do you connect without opening the plant network? I own the network, so this is the part I get asked about first. The pattern that holds up under review is dull on purpose. The collector sits in a DMZ, a separate network zone between the plant and the outside, not on the control network. The connection opens outward from the plant, and nothing initiates a session inward. The account it uses on the control side is read-only, unique to this job, and not the engineering workstation login. Credentials are certificates where the protocol supports them. The path has no ability to write to a PLC, and if someone asks for write-back later, that is a separate project with separate approval and a different set of people in the room. Everything gets logged somewhere a human will actually look. I am describing data flows here. The applicable standards and your own security lead govern what you are allowed to do, and that is a conversation to have before the gateway arrives, not after. Ask any vendor these six questions: Which direction does the connection open? What account does it use on the control network, and what can that account do? Where does the data land, and who holds the keys to it? Can we read the full tag mapping ourselves, in a format we can export? What happens when the link drops, and does the gateway store the data locally and send it once the link is back? What can it write, and how do we turn writing off permanently? A vendor who cannot answer four and six in plain language has not thought about your plant. Why the data still will not answer your question You will get the tags flowing in a week, and then you will hit the real work, which is writing down what each tag means and how the things it describes relate to each other. A stream of readings does not know that this tag belongs to this machine, that this machine sits on this line, that this line ran this order for this customer, or that ninety seconds of stopped infeed becomes a short shipment on Friday. Those relationships are how your operation actually works, and no protocol carries them. Until those relationships are written down somewhere a machine can read, any analysis you put on top, including anything with AI in the name, is working with numbers it cannot connect to anything. The reason to start narrow and start soon has nothing to do with technology. It is the technician in the electrical room. He can read those tag names the way you read your own handwriting, and he is the reason your data has any meaning at all right now. When he goes, the tags stay and the meaning leaves with him. Moving SCADA data off the floor is the easy half. Doing it while the people who can explain it are still in the building is the half with a deadline you do not set. This week: Pick one line and one question this week, write down the thirty to fifty tags that would answer it, and see how many already have a name a stranger could read.