JJC SystemsBook a Consultation
Azure · Manufacturing

The edge-to-cloud architecture decision

Architectures that assume connectivity fail in exactly the places manufacturers need them most. This paper examines the design decision and the failure nobody plans for.

PublishedApril 12, 2026
Length15 pages · 16 min read
SectorManufacturing
PlatformAzure
Service areaStrategy & Transformation
Abstract

The edge-to-cloud architecture decision

Summary

A cloud analytics design that works in a demonstration meets its first real test in a plant with a congested link, a basement with no signal, and a shift that continues regardless of whether the network is available.

This paper argues that the design question is not how to guarantee connectivity but what happens when there is none and for how long, examines the reconnection failure that most designs carry undetected, and proposes a boundary between reading operational data and writing to anything that controls a machine.

Key findings

Four things this paper argues

If you read nothing else, read these. The analysis that follows sets out the evidence for each.

01

Sorting workloads by disconnection tolerance is the architectural decision

Everything else follows from it, and it has to be done with operations in the room rather than inferred by IT.

02

The reconnection path is where most designs fail in production

The link returns, hours of buffered data flow in, and something double-counts or arrives out of order. The failure is silent until a monthly report is wrong.

03

Buffers should be sized for the worst realistic outage, not the typical one

The outage that matters is the unusual one, and a buffer that overflows loses data quietly.

04

Reading and writing are different risk classes and should be treated as such

Reading machine data to inform planning is straightforward. Writing to anything that controls a machine carries safety implications and deserves an explicit boundary.

Analysis

The argument in full

The sorting exercise

Some things must continue through a disconnection: production execution, safety systems, quality holds, anything that stops a line. Others can wait: analytics, reporting, most integration, anything whose value is measured in days rather than minutes.

Sorting workloads into those two categories, explicitly and with operations present, is the architectural decision. It determines what runs locally, what runs in the cloud, and what the store-and-forward path between them has to guarantee.

Done by IT alone, this exercise reliably places things in the wrong category, because the question of what genuinely stops a line is operational knowledge rather than technical.

The reconnection failure

This is the part that produces production incidents months after go-live.

The link returns after an outage. Several hours of buffered messages flow into the cloud. Something double-counts, or arrives out of order, or overwrites a value that was already correct. Nobody notices, because everything appears to be working. The discovery comes weeks later when a monthly report disagrees with the floor, by which point nobody can reconstruct which period was affected.

The design response is idempotent message handling with a stable deduplication key, a defined window for accepting late arrival, and a quarantine for anything outside it. And then deliberate testing: disconnect for a realistic period, let the buffer fill, reconnect, and verify counts reconcile exactly. Then do it again with deliberately out-of-order delivery.

  • A deduplication key stable across retransmission — device, sequence and event timestamp
  • Deduplication applied at ingestion rather than downstream
  • A defined late-arrival window with a monitored quarantine beyond it
  • Buffer depth alerting, because silent accumulation is the dangerous failure
  • Reconnection tested deliberately, including with out-of-order delivery

The boundary worth stating

Reading machine data to inform planning, costing and analytics is straightforward and valuable. Writing back to anything that controls a machine is a different risk class entirely.

We treat that boundary as explicit and written down, and we are deliberately conservative about it. Any proposal that quietly crosses it deserves a direct question about who has assessed the safety implications and under what standard.

This is not a technical limitation. It is a position about where a technology consultancy's competence ends and a controls engineering discipline begins, and being clear about it is part of what a client is buying.

Framework

Something you can apply without us

Every paper in this series ends with a framework you can run internally. We would rather you used it and reached your own conclusion than took ours on trust.

Framework

Designing for intermittent connectivity

Five decisions, in order.

1

Sort

Workloads by disconnection tolerance, with operations in the room.

2

Size

Edge buffers for the worst realistic outage plus margin, not the typical one.

3

Deduplicate

A stable key applied at ingestion, so reconnection cannot double-count.

4

Test

Disconnect deliberately, reconnect, and reconcile exactly. Then test out-of-order delivery.

5

Bound

Write the read-only boundary into the design document and the contract.

Implications

What this means, depending on your seat

The same argument lands differently across an executive team. These are the three versions worth separating.

For the plant manager

For the CIO

For the CFO

References

Where to check this for yourself

Microsoft's own documentation for the product behaviour described above. We would rather you verified the basis than accepted our summary of it.

01
Azure IoT Edge documentation
02
Microsoft Fabric Real-Time Intelligence
03
Azure Event Hubs
04
Idempotent message processing patterns
05
Azure Well-Architected Framework: Reliability

On these references: each entry names a Microsoft Learn article or documentation area by title, because deep links change while titles are stable. Searching the title on learn.microsoft.com will reach the current version. Where we have cited a figure or a product behaviour, it is Microsoft's statement rather than ours; where we have given a number of our own it is labelled as such in the text.

Recognise the situation?

We will assess one plant's connectivity profile and workload mix, and produce an edge-to-cloud design with the disconnection and reconnection behaviour specified and tested.

Discuss this paper Run the related checklist We reply to every message within one business day.
Keep reading

Related papers

defender
manufacturing14 pages

Operational continuity as a security objective

Ask a manufacturing executive what a security incident would cost and the answer involves stolen designs. Ask what a week of stopped production would cost and the number is immediate and much larger.

April 26, 2026 · 14 pages · 15 min readRead
Get In Touch

Tell us what you're trying to fix

Describe the situation in your own words.

Please enter your first name.
Please enter your last name.
Please enter a valid email address.
Please enter your company name.
Please choose an option.
Please add a short description.

We reply to every message within one business day.