Sorting workloads by disconnection tolerance is the architectural decision
Everything else follows from it, and it has to be done with operations in the room rather than inferred by IT.
Architectures that assume connectivity fail in exactly the places manufacturers need them most. This paper examines the design decision and the failure nobody plans for.
A cloud analytics design that works in a demonstration meets its first real test in a plant with a congested link, a basement with no signal, and a shift that continues regardless of whether the network is available.
This paper argues that the design question is not how to guarantee connectivity but what happens when there is none and for how long, examines the reconnection failure that most designs carry undetected, and proposes a boundary between reading operational data and writing to anything that controls a machine.
If you read nothing else, read these. The analysis that follows sets out the evidence for each.
Everything else follows from it, and it has to be done with operations in the room rather than inferred by IT.
The link returns, hours of buffered data flow in, and something double-counts or arrives out of order. The failure is silent until a monthly report is wrong.
The outage that matters is the unusual one, and a buffer that overflows loses data quietly.
Reading machine data to inform planning is straightforward. Writing to anything that controls a machine carries safety implications and deserves an explicit boundary.
Some things must continue through a disconnection: production execution, safety systems, quality holds, anything that stops a line. Others can wait: analytics, reporting, most integration, anything whose value is measured in days rather than minutes.
Sorting workloads into those two categories, explicitly and with operations present, is the architectural decision. It determines what runs locally, what runs in the cloud, and what the store-and-forward path between them has to guarantee.
Done by IT alone, this exercise reliably places things in the wrong category, because the question of what genuinely stops a line is operational knowledge rather than technical.
This is the part that produces production incidents months after go-live.
The link returns after an outage. Several hours of buffered messages flow into the cloud. Something double-counts, or arrives out of order, or overwrites a value that was already correct. Nobody notices, because everything appears to be working. The discovery comes weeks later when a monthly report disagrees with the floor, by which point nobody can reconstruct which period was affected.
The design response is idempotent message handling with a stable deduplication key, a defined window for accepting late arrival, and a quarantine for anything outside it. And then deliberate testing: disconnect for a realistic period, let the buffer fill, reconnect, and verify counts reconcile exactly. Then do it again with deliberately out-of-order delivery.
Reading machine data to inform planning, costing and analytics is straightforward and valuable. Writing back to anything that controls a machine is a different risk class entirely.
We treat that boundary as explicit and written down, and we are deliberately conservative about it. Any proposal that quietly crosses it deserves a direct question about who has assessed the safety implications and under what standard.
This is not a technical limitation. It is a position about where a technology consultancy's competence ends and a controls engineering discipline begins, and being clear about it is part of what a client is buying.
Every paper in this series ends with a framework you can run internally. We would rather you used it and reached your own conclusion than took ours on trust.
Five decisions, in order.
Workloads by disconnection tolerance, with operations in the room.
Edge buffers for the worst realistic outage plus margin, not the typical one.
A stable key applied at ingestion, so reconnection cannot double-count.
Disconnect deliberately, reconnect, and reconcile exactly. Then test out-of-order delivery.
Write the read-only boundary into the design document and the contract.
The same argument lands differently across an executive team. These are the three versions worth separating.
Microsoft's own documentation for the product behaviour described above. We would rather you verified the basis than accepted our summary of it.
On these references: each entry names a Microsoft Learn article or documentation area by title, because deep links change while titles are stable. Searching the title on learn.microsoft.com will reach the current version. Where we have cited a figure or a product behaviour, it is Microsoft's statement rather than ours; where we have given a number of our own it is labelled as such in the text.
We will assess one plant's connectivity profile and workload mix, and produce an edge-to-cloud design with the disconnection and reconnection behaviour specified and tested.
Deferral is a decision with a running cost. This paper quantifies where that cost accumulates and offers a framework for deciding whether to defer again.
Ask a manufacturing executive what a security incident would cost and the answer involves stolen designs. Ask what a week of stopped production would cost and the number is immediate and much larger.
The parallel quarter is the whole method. Switching without one produces an accurate number nobody trusts.
Describe the situation in your own words.