← Insights

"Saved" is not "received"

The app says "Saved." The device never got the update. Nothing failed, and both screens look fine.

A pattern I see a lot in connected devices: a cloud app saves a record, the device receives a small MQTT message with just the record ID, then fetches the full record over REST.

It's a sound pattern. MQTT is the doorbell. REST is the source of truth.

The mistake is treating the doorbell as the guarantee. On AWS IoT Core, it can't be.

Where the doorbell fails

  • QoS 0 provides no delivery guarantee. The device may never receive the ID.
  • QoS 1 can deliver a message more than once. The fetch and apply operation has to be safe to repeat.
  • Message order isn't guaranteed. Two updates can arrive reversed.
  • Persistent sessions have limits. Messages queued while a device is offline can eventually expire, leaving the device unaware that it missed something.

A live connection only proves the link is up. Keepalive and Last Will don't prove that a particular record arrived or was applied.

So the guarantee can't live in the transport. It has to live in the protocol you build on top.

What the protocol needs

  • The notification is a hint, never the sole trigger for a safety-relevant action.
  • The fetch returns current state, and the device checks that state before acting. An ID that arrives late may point to a record that has since changed or been cancelled.
  • Records carry a version. The device ignores anything older than what it already holds.
  • Applying a record is idempotent.
  • The device confirms after it applies the update, so "saved," "received," and "applied" are different states in the backend.
  • The device reconciles on a timer, asking REST for anything changed since its last successful sync. This catches missed notifications, duplicates, reordering, and expired sessions.

The timer is a clinical decision

That timer is the important decision. It is not just a tuning knob. In a medical device, it can become part of the risk control. It sets the maximum period a device can remain out of date without detection. Someone has to define that number, and the rationale should come from the risk analysis.

The device UI should also show when it last successfully synced, so the person using it can judge how fresh the data is.

Test the failures on purpose

In SaMD, each of these mechanisms can become a software risk control: a requirement, traceable to a hazard, and verified. That means testing the failures deliberately:

  • Dropped notifications
  • Duplicate notifications
  • Reordered notifications
  • Delayed notifications
  • A device offline beyond the session window
  • Stale or cancelled records arriving late

If you own a platform like this: when a notification is lost, what catches it, and where is that written down?

Seen this differently?

Questions, corrections and counterexamples are welcome.

Related

Keep reading