← Blog

React Native offline queue: why 409 is not a saved confirmation

8 min read
Watch the walkthrough on YouTube

The companion video becomes viewable when its YouTube release goes public.

An offline write has two acknowledgments: the phone persisted the intent, and the server accepted the operation. Treating either a local pending badge or an ambiguous server response as the second acknowledgment can lose work.

This is Day 19 of Ship Native / Production fixes. It follows the existing workQueue implementation in HM-tasks2 and the project's backend correspondence. The example key K1 and timestamps are teaching data. Source was inspected; no native app, live credentials or backend was executed.

Persist the action, reuse its key, and remove it only when the server outcome is known.

The rule is a design goal. The source fixes an important 409 mistake but still has delivery boundaries worth reviewing.

The documented bug

The backend reply dated 2026-08-18 distinguishes a completed replay from an original request still running:

Situation under this backend contract Result Queue consequence
Original request has finished Original status plus Idempotent-Replay: true Inspect the original outcome; only a successful outcome confirms acceptance
Original request is still running 409 Retain intent and retry using the original key

A follow-up in BACKEND-ASKS-ROUND-2.md records that the old offline queue interpreted 409 as already recorded and removed the action. If the first request later failed, nothing remained to retry. The current classifier changes that decision to bounded retry.

The project documents this history. We did not reproduce a production incident. Older closeout documents still describe the obsolete remove-on-409 behavior. The current source and later correction govern this lesson.

A 409 on another endpoint can mean a different conflict. Copy the policy only after verifying that endpoint's contract.

Persist the action before presenting local durability

The queued record holds a stable id, action kind and payload, occurredAt, createdAt, and attempts. storage.ts writes the array as JSON in a dedicated MMKV instance:

export function enqueue(action: IQueuedAction): void {
  const queue = loadQueue();
  queue.push(action);
  saveQueue(queue);
}

That creates a persistence path across app launches. It does not establish that every storage write succeeds or that every restored payload is valid. loadQueue returns an empty array when JSON parsing fails, and the cast does not validate its shape.

More directly, enqueueAction catches and logs persistence failures but still returns an ID and starts a flush. The task wrapper can then mark that task pending. A production interface should expose local storage failure before claiming the action was saved locally. This lesson documents that gap; it does not patch the source application.

Keep the states visible: saving locally, pending sync, syncing, accepted by the server, and failed with a recovery action. A local success message should identify what actually succeeded.

Reuse one key for the same intent

The request builder uses the action's persisted ID:

const headers = { 'Idempotency-Key': action.id };
const occurred_at = action.occurredAt;

Imagine the server commits K1, but the response never reaches the phone. The client cannot infer whether the write occurred. Repeating the same intent with K1 lets this backend return its recorded outcome. Assigning a fresh key turns the retry into another operation and bypasses that deduplication boundary.

Different user intents still need different IDs. A key is not a global loading flag. Server-side scope, retention, payload matching and atomic execution determine what protection it supplies. The header alone is not proof of exactly-once delivery.

Keep the original event time as well. Syncing an earlier event later should not silently rewrite its time as the reconnect time. The project comments describe backend support for occurred_at; we inspected that statement. We did not verify it against a live server in this session. Clock accuracy, timezone handling and acceptable event age remain integration concerns.

Serialize a snapshot, then understand its limits

The manager guards concurrent flush loops:

if (isFlushing) return;
isFlushing = true;
try {
  const queue = loadQueue().sort((a, b) =>
    a.createdAt.localeCompare(b.createdAt),
  );
  // Send the loaded actions sequentially.
} finally {
  isFlushing = false;
}

The guard is acquired before awaiting network work. Enqueue, network events and foreground events can ask to flush, but only one loop runs in this JavaScript runtime. For retryable outcomes under the cap, the loop breaks at the current action so later snapshot entries do not overtake it.

The loop processes the loaded snapshot. It does not continuously drain new entries. An action enqueued during the active loop is absent from its loaded array, and its flush call returns while isFlushing is true. It may remain pending until another trigger. Reloading the queue until eligible work is exhausted, or retaining a drain-again request, are possible improvements; neither is implemented here.

The startup integration subscribes to NetInfo and foreground events. It uses isConnected as a trigger. NetInfo distinguishes network connectivity from internet reachability; neither proves that this application server is currently accepting writes. The actual request outcome remains decisive. See the NetInfo API documentation.

These listeners also do not make the module a background worker that runs while the app is terminated.

Read the actual retry policy

The current classifier and manager together produce these decisions:

Outcome Classifier Manager behavior
No response Retry Retain; stop without incrementing the response attempt count
2xx Remove Reconcile and remove
409, 429, 5xx Retry Increment count; stop under cap; at five attempts remove, show failure and continue
Other status, including 401 and 422 Drop Remove; 401 skips the local error toast

The cap applies to all response-bearing retry cases, even though some comments and log messages mention only server errors. A classifier comment mentions backoff for 429, but the manager has no retry timer or Retry-After handling. Calling a branch retryable does not implement a retry schedule.

At the cap, removal is a terminal failure policy. It does not mean the server accepted the action, and it does not guarantee eventual delivery. For valuable user work, retaining a failed entry with explicit recovery may be preferable to deleting the intent. That is a recommendation. The current implementation deletes capped entries.

Expired authentication exposes a second boundary

claimQueueFor(userId) preserves the stored queue for the same owner and clears it when a different known owner claims it. The explicit sign-out helper clears queued work and UI mirrors. The token-expiry logout path intentionally avoids that bulk clear.

However, that does not mean every action survives an expired session. classifyFlushResult still drops 401, and manager.test.ts explicitly expects the individual unauthorized action to be removed. The authentication interceptor rejects the error after logging out. Bulk preservation and per-action removal coexist.

A preserve-and-pause authentication policy needs coordinated classifier, manager and session handling. Account switching during an active flush also needs generation or cancellation checks: clearing the stored queue cannot recall a request already sent, and an older snapshot may still hold later actions. The content identifies these limits without claiming they are solved.

Reconcile the interface with server state

Startup hydrates pending task badges from stored actions. Success or terminal removal clears the corresponding pending state, then invalidates related query caches so a refetch can reconcile optimistic state or roll it back.

A cleared pending badge is not automatically a success acknowledgment. The manager also clears it on terminal failure. Multiple queued actions for one task deserve separate review because the UI store tracks task ID and kind rather than a distinct badge for each queued action.

Reproducible regression evidence

The delivery includes a dependency-injected teaching runner with in-memory storage and a source-derived classifier. Native/UI details are omitted. Run:

node --test day-19/source/queue-policy.test.mjs

Seven isolated checks passed: 409 followed by success retains the same key/time, ambiguous network failure preserves intent, snapshot ordering blocks later work on retry, overlapping flushes do not duplicate the loop, 429 reaches the shared cap, 401 removes the current action, and enqueue during a snapshot requires a later trigger.

The deliberately obsolete decision can be selected with:

SHIP_NATIVE_QUEUE_LEGACY=1 node --test day-19/source/queue-policy.test.mjs

That run has one expected failure: the action disappears on 409, exposing the regression. Six other checks pass. Saved TAP outputs accompany the package.

The original project Jest tests were inspected but not run: its dependencies are not installed. These seven teaching tests establish isolated policy behavior, including known limitations. They do not establish native durability, device UI behavior, authentication integration or server deduplication.

Device and server checks still needed

  1. Persist an action, terminate the app, and restart under the same owner. Confirm its ID and payload survive.
  2. Commit a write on the server, lose the response, then replay the same key. Inspect actual server records and the returned outcome.
  3. Hold the original operation pending and return the contract's 409. Confirm no saved acknowledgment or premature removal appears.
  4. Expire auth during a flush and switch accounts during an active send. Confirm pending intent cannot cross owners.
  5. Fail persistence, corrupt stored data and enqueue during an active snapshot. Confirm the interface explains recovery.
  6. Rate limit and reach the response cap. Confirm waiting and terminal failure follow the desired product policy.

The documented correction makes a meaningful full debugging lesson. The remaining gaps explain why a durable queue is a protocol between storage, request identity, server outcomes, ownership and the interface, rather than just an array of requests.

Comments

No account needed.

  1. Loading comments…