Skip to content

Operating in production

The read and write primitives are fast, but a few runtime and deployment properties around them decide whether a live app feels fast. None are toolkit bugs — they are how serverless edges, browser HTTP, and CDN-shaped caching behave — but every one of them was hit dogfooding the demo, so they are collected here with their fixes. If writes or sync feel slow in a real browser but your server benchmarks are fast, the cause is almost always on this page.

Convergence cadence: event-driven, with the interval as a fallback

Section titled “Convergence cadence: event-driven, with the interval as a fallback”

When you pass an autoSync trigger to createSyncClient, the client drives a flush → reconcile pass. The pass is event-driven: the client calls requestPass() the moment a mutation is enqueued, so a local write flushes to the server immediately — it does not wait for the trigger’s next interval tick. The interval (createBrowserConvergenceTrigger({ intervalMs }), default 1.5s) is therefore only a fallback for retries, recovery, and cross-tab wake-ups.

Because the happy path is event-driven, you should make that fallback interval long. A short interval is the dominant idle cost: every local query carries ~50ms of WASM overhead and serializes on the one worker thread, and an unconditional reconcile each tick fires the store’s live-query notifications, re-running every mounted query. The toolkit already idle-skips an empty reconcile, but the cheapest idle board is still a rare interval — the board demo runs intervalMs: 15_000, taking idle CPU from ~70% of a core to ~2% with no change to convergence latency (latency is bounded by the read-path echo, not the interval).

In worker mode the worker owns this loop: defineSyncWorker’s convergenceIntervalMs is the same fallback interval and already defaults to 15s. Writes still flush on enqueue (the write RPC requests a pass) and tabs forward online/visibilitychange as wake signals, so the same rule holds — the interval is a retry/recovery sweep, not the write path.

Local write latency: the durability preference (relaxed by default)

Section titled “Local write latency: the durability preference (relaxed by default)”

durability is declared once, on the registry — SyncRegistryDefinition.storage.durability ("relaxed" | "strict", default "relaxed"). It is not a minting-surface, worker-entry, or attach-site option: whether losing the last not-yet-flushed action is acceptable is decided by what the data IS, so one declaration binds every open of every store minted from that registry — no tab can ever disagree with another. The physical behavior depends on the capability-selected backend:

  • IndexedDB: strict flushes the whole datadir synchronously at the end of every query, setting a ~100–200ms optimistic-write floor. Relaxed returns before that snapshot and schedules it asynchronously.
  • OPFS-repacked: pgwasm always awaits the store’s sync. Relaxed asserts VFS health and runs any due deferred repack without an ordinary physical flush; strict flushes arena data before metadata. Initialization, repack activation, and open-state close keep strict ordering in both modes.

The resolved mode is stamped on the boot pgwasm.create rail line. A capability fallback from opfs to idbfs keeps the registry-declared durability unchanged.

What you trade. On idb, writes since the last completed snapshot are at risk only if the browser terminates before both their journal rows reach the write API and the scheduled snapshot lands. On OPFS-repacked, relaxed recovery returns the longest valid stable metadata-log prefix, so an unflushed suffix may be absent; a returned strict boundary is stable under the browser-termination model. Synced tables are server-recoverable by construction. Your own local-only tables have no such copy.

On a serverless Edge platform a worker is suspended when idle and evicted after longer idle, so the first write after a quiet period pays a cold start while steady-state writes are instant. Measured on the self-hosted Supabase edge-runtime: a warm write applies in ~20ms; a write to a worker suspended ~15s pays ~0.45s (a Postgres reconnect on resume); a write to a worker whose module cache is cold pays ~5.8s (a fresh isolate re-imports the whole bundle). Drag the first card after the board sits idle and that cold worker is the entire delay — not the sync rail.

This is a property of the serverless deployment target, not pgxsinkit: the same functions on a long-lived Bun or Deno process (one warm process, a pooled connection) or on a managed warm pool have no cold start. Two mitigations if you stay serverless:

  • Keep the worker warm with a periodic cheap request. The cheapest request that still reaches the worker is a no-op write — an empty {"mutations":[]} POST, rejected at request validation before any DB work. A small sidecar pinging it every ~8s keeps writes at ~20ms after idle.
  • Set the worker’s wall-clock timeout above your longest held-open long-poll. A durable-streams long-poll is held until data arrives or the hold expires, so a stock 60s worker budget recycles live subscriptions constantly; the board demo runs its read functions at EDGE_WORKER_TIMEOUT_MS: "600000" (ten minutes). See Deploying the server.

Caching: the control plane is private, the edge is the cacheable surface

Section titled “Caching: the control plane is private, the edge is the cacheable surface”

The read path splits into two ingress points with opposite caching postures, and getting them the wrong way round is the mistake this section exists to prevent.

The control plane is never cacheable. /sync/v1/subscribe, /sync/v1/refresh, /sync/v1/release and /sync/v1/barrier answer per-subject questions and mint stream tokens. Force cache-control: no-store on it. A cached subscribe answer is a capability served to the wrong subject; a cached barrier answer can license an alignment against an engine that has already lost membership effects (which is why the barrier’s own cache window defaults to zero).

The edge is the only surface a CDN may share, and it can only do that if it is a separate origin — the cache key is the URL, so a shared tier mounted beside the control plane can never be fronted without also caching the private one.

The stream token must be excluded from the cache key. This is a correctness condition for the shared tier, not a tuning detail: include it and every subscriber gets a unique key, which destroys exactly the sharing the tier exists for.

Catch-up reads are where the bytes are, and durable-streams already marks them cacheable — a catch-up response carries cache-control: public, max-age=60, stale-while-revalidate=300 and an etag. Live long-polls are not cacheable and are not marked so. Respect the origin’s cache-control verbatim rather than overriding TTLs in either direction, key on the full URL minus the token, and set proxy/CDN read timeouts above the long-poll hold so a completed poll is not cut off. Cacheability is a property of the shared tier only: a private-tier shape’s stream is per-subject by construction and no cache can share it.

Revocation latency is deployment-shaped, and you have to state which shape yours is. Behind a third-party CDN serving hits, it is bounded by the stream token’s TTL (5 minutes by default). Behind a cache sitting behind the edge’s own verifier, it is entitlement-propagation latency instead. Private-tier grants are re-authorized on every re-mint too — the control plane recompiles the shape against the subject’s current claims and revokes the grant if it no longer compiles to the shape that was granted — so a claims change (a role removed, a tenant switched) is bounded by the same TTL as an entitlement change, without the subject having to re-authenticate for it to take effect.

The re-mint only narrows, though. A scope a subject gains after subscribing — a membership acquired mid-session, or a shape that was refused at subscribe — is granted at the next subscribe: a boot, or the automatic restart that follows a stream ending or a grant being revoked. It does not arrive within the TTL on its own. Widening an existing private-tier grant is the exception, by way of the fingerprint: a claims change that compiles to a different predicate is revoked at the re-mint, and the restart re-subscribes with the new claims.

Engine shape claims: a closed subscription gives them back

Section titled “Engine shape claims: a closed subscription gives them back”

Every grant a subscribe hands out is a named claim on an engine shape, and a live claim is what keeps that shape active — native reads terminate on durable-streams, so a claim is the only thing telling the engine a reader exists. Closing a subscription releases those claims (POST /sync/v1/release), which is what lets an idle shape go dormant and eventually be evicted promptly.

A claim is also a lease. It counts as live only while it is renewed within the engine’s CIRCUITS_SHAPE_IDLE_SECS window, and the control plane renews every still-authorized grant on the token re-mint. So claims of a client that crashed or was unloaded before its release left the browser are reclaimed by the engine within one lease window rather than held until a restart — nothing renews them once the session is gone. Because each claim is named, the release is idempotent: repeating one is the same act once, and cannot take another subscriber’s claim.

That pairing has a configuration requirement, and the control plane enforces it: if the engine’s lease window is shorter than twice the stream-token TTL, subscribe and refresh answer 503 {"error":"sync engine lease window shorter than the token TTL"} rather than serving sessions whose shapes would lapse after a single missed refresh. Raise CIRCUITS_SHAPE_IDLE_SECS, lower ttlSeconds, or set the idle window to 0 (which disables dormancy, and with it leases).

The browser connection budget (serve the gateway over HTTP/2)

Section titled “The browser connection budget (serve the gateway over HTTP/2)”

@pgxsinkit/client’s reader holds one live long-poll connection open per synced stream. A client subscribing to six streams keeps six connections continuously busy — and on the shared tier a subject holding K scopes of one shape holds K streams, not one — while browsers cap HTTP/1.1 at ~6 connections per origin. So over plain HTTP those long-polls consume every slot and the write request (same origin) gets Stalled in the browser’s connection queue for a whole long-poll cycle before it is even dispatched. Serve the gateway over HTTP/2 (or HTTP/3), which multiplexes every request over one connection, so the cap never binds. This is mandatory rather than advisory: durable-streams serves one stream per request and the Rust server is HTTP/1.1-only with no TLS. It only bites a local stack served over plain http://; any production ingress (Cloud Supabase, an istio/Envoy gateway, a TLS reverse proxy) already speaks HTTP/2. Full detail and the symptom check are in Deploying the server.

Read-path apply failures hold, they don’t diverge

Section titled “Read-path apply failures hold, they don’t diverge”

If a local commit into the store keeps failing — a bad local migration, a storage/quota error, a corrupt local store — the read path retries with backoff and then latches into a degraded phase instead of advancing past the change it could not apply. It holds the read cache at the last good commit, so the client never silently diverges from the server. The failure is surfaced through the onSyncError callback you pass to createSyncClient, and the runtime’s status reports the degraded phase.

How it recovers depends on why it degraded:

  • A subscribe failure — a control plane that is unreachable or keeps erroring — takes the same recoverable degraded phase, with lastError naming it, and subscribe keeps retrying with backoff behind it. It never masks a more actionable status: an auth failure wins over it, and a commit failure (below) is never overwritten by it.
  • A read-stream error (a dropped stream connection) clears automatically on the next successful batch. The reader retries 429, every 5xx and network failures with backoff. A 401/403 gets one re-mint and one retry, so an expired stream token is replaced instead of presenting as a stall, and a token refused again ends the read, as every other 4xx does, rather than being swallowed. A connection that dies by hanging rather than failing (a pulled cable mid-long-poll) produces neither a failure nor a batch, so a runtime claiming ready is additionally held to a read-silence window (readSilenceMs, default 45s — a healthy stream’s long-poll always cycles well inside it): silence past the window drops ready to the same self-recovering degraded. That is the phase to key an offline/“connection needed” surface off; navigator.onLine is not a substitute. Recovery is wired to delivered traffic — any batch arriving clears it and re-arms the window.
  • An auth failure on the read path is its own phase, not a degraded. A persistent 401/403 from the control plane puts the runtime in auth-needed while subscribe keeps retrying for a fresh token, so the app can prompt for re-login; it clears the moment a batch lands. This is the phase behind an expired session, and it is deliberately distinguishable from an unreachable deployment.
  • A commit failure is sticky — fetching can keep succeeding while applies fail — so it clears only on the next commit that succeeds, after you fix the underlying cause, or on a client restart. There is no separate reset call for this; recoverSending rebuilds the write journal, not the read frontier.
  • A degraded engine is terminal, not something to wait out. If the engine abandons a membership flip batch after exhausting its retries, those effects are lost and it latches itself degraded (answering 503). The client refuses to align rather than holding, and reports the refusal through onSyncError. Recovery is an operator restart, after which clients re-subscribe and re-snapshot.

If your app appends events (the model is in The event lane), the client gives you two surfaces to operate it with, and neither is optional reading before you ship one. The knobs on the other side of them — batch caps, the fallback interval, backoff bounds, and the per-stream fairness cap — are the client’s events option, documented under Tuning the flush.

The drain signal is client.onOutboxStatus(({ empty }) => …) — the empty ↔ non-empty transitions, with the current state delivered on subscribe (await client.outboxStatus() is the one-shot pull). “Empty” means nothing is awaiting a server verdict. Use it to invalidate a view that composes pending events with down-synced aggregates. It carries no count deliberately — a count that updates only on transitions is stale by construction — so query the Outbox table (getOutboxTable(registry)) if you want one.

If that composed view dips between the ack and your consumer’s fold, that is the gap the acked ledger closes: events.ackedRetentionMs keeps an acked row, stamped, instead of deleting it at the ack. A retained row is not pending — it does not hold the drain signal, is never re-sent, and never blocks a destroy() — so a composition reads acked_at_us IS NULL for “still owed” and treats the rest as its grace window.

The verdicts are client.onEventLaneReport(cb): per flush pass, the terminal non-acked verdicts, the deferred ones, and the lane’s batch-level backoff transitions. Subscribe for the app’s lifetime, not per screen — the subscription is ephemeral, and once a terminal row is deleted the Outbox cannot answer for it. (With nothing subscribed the library logs each report at warn level rather than dropping it.)

Read them this way:

  • refused — your eventGate declined it. Expected, not an error.
  • rejected — a schema-invalid or oversized payload for a known stream. The library validates at append, so in practice this means a non-library caller or a broken deployment: treat it as a bug.
  • deferred — the server does not (yet) know that stream. This is not a failure and not terminal: it is ordinary rollout skew, the rows stay in the Outbox, and they drain when the server deploy lands. A burst right after a client release is the deploy order; a burst that never clears means deployment skew — the server’s registry does not declare that stream (it is the registry, and only the registry, that decides this verdict). Never “clean up” the Outbox in response to it.

A lane stuck in batch-level backoff is a different diagnosis, and the report’s backoff transitions (plus a persistent 503 with Retry-After) are its signal. The server knows the stream but could not enqueue the batch, and it fails the batch whole — nothing is enqueued and no per-event verdicts are issued. The first thing to check is the queue itself: registering a stream requires generating and applying the --events migration, and without it the endpoint enqueues onto a pgmq queue that does not exist. (Then: the database’s reachability from the ingress, and the server log line the 503 always writes.)

The optional write-side audit log (operationsLog: { enabled: true }, off by default) persists every mutation — table, kind, and payload — including the raw body of mutations that failed validation. Those payloads contain whatever your users typed. So treat the operations_log table as sensitive: restrict access and set a retention policy, and leave it disabled in environments where that content should not sit at rest.

Its mutation_id column is text, not uuid — deliberately. This is the server-side tier of a two-tier invariant: pgxsinkit’s public write surface (the HTTP route’s request/ack schemas and the client’s mutation_id UUID journal) is UUID-only by contract, but the generated apply function and this log accept an opaque text id for one narrow case — a direct, server-side caller that derives child envelopes with composite ids (${parentMutationId}:<tag>:<n>) and invokes the apply function itself, never crossing the HTTP route or the client journal. A non-UUID id can never reach the UUID-typed public surface.

Debugging latency: globalThis.__pgxsinkitDebug

Section titled “Debugging latency: globalThis.__pgxsinkitDebug”

@pgxsinkit/client ships opt-in, timestamped instrumentation that traces a write through every phase — exactly what localises a “writes are slow” problem to a single hop. It is off by default and adds nothing to a normal run; enable it from the console or before the client boots:

globalThis.__pgxsinkitDebug = true; // then reproduce; filter the console to "pgxsinkit" + enable Verbose

Each line is stamped with a monotonic millisecond clock, so you read the gaps between phases directly:

  • mutation staged {mutationId, table} (the write’s origin — correlate by id with the sent/acked lines)
  • convergence pass requested → convergence flush → convergence reconcile (with durations)
  • board-write auth token resolved {ms} (a stalling per-request getSession() shows up here)
  • board-write responded {status, ms} (a cold edge worker, or a browser connection stall, shows up here)
  • live-query register / retained / dedup-hit / evicted (the local read model’s subscription bookkeeping) and live query updated → re-render (the final UI hop, from @pgxsinkit/react)
  • boot pgwasm.create → boot local schema → boot journal recovery → boot store-version reconcile → boot sync start → boot client ready (the boot phases, for attributing a slow first paint to one of them)
  • boot pgwasm.create phase (on an OPFS store: module-loaded, directory-ready, store-opened, pgwasm-ready), so a create that never returns is attributable to one step
  • boot pgwasm build warm (only when the host times its pre-warm on the rail, as the board demo does — see next section)

The read path itself does not emit rail lines. Its transport, @pgxsinkit/client’s own long-poll reader, is not instrumented on this rail. Read it at the two boundaries instead: boot client ready (fired when the first initial sync completes) and the runtime’s status phase, which names syncing, ready, degraded (with lastError) and auth-needed. For per-request detail, inspect the network layer directly — and in worker mode that means inspecting the worker, not the page (see below).

The structured BootReport — measure before you optimize

Section titled “The structured BootReport — measure before you optimize”

The rail lines above are for a human mid-debug: you eyeball the gaps as they scroll. For numbers a machine can keep — a dashboard series, a CI budget gate, an honest before/after — every boot also builds a structured, versioned BootReport, independently of the rail, so it exists whether or not __pgxsinkitDebug is on. Boot performance regressed repeatedly because it was optimized on guesses (the suspected bottleneck was rarely the real one); the report is the evidence to start from instead (ADR-0034). Read it by push, by pull, or both:

const client = await createSyncClient({
registry,
controlPlaneUrl,
streamBaseUrl,
batchWriteUrl,
onBootReport: (report) => {
// fires exactly once, at boot completion — ship it to a dashboard, or assert a CI budget
metrics.timing("boot.total", report.totalMs);
},
});
const report = await client.bootReport(); // the most recent completed boot, or null before the first sync

report.totalMs is boot start → every eager group caught up; phases decomposes the local work (store create, schema exec, journal recovery, store-version reconcile, sync start, catch-up) and groups[] breaks out the per-consistency-group boot catch-up (rows, requests, fetchMs, applyMs, start/ready offsets).

localReadReadyMs and writeReadyMs mark the staged-boot crossings (ADR-0041; additive, reportVersion stays 1). localReadReadyMs is boot start → cached reads are safe (store open, schema compatible, reconcile done — zero network); it is the moment attachSyncClient / createSyncClient resolve. writeReadyMs is boot start → the write runtime + boot recovery finished (enqueue is safe). Both are null when the boot rejected before reaching that stage. The gap from localReadReadyMs to totalMs is the whole-sync catch-up a cached paint no longer waits on.

storeKind names how the store presented at boot — "restored" (seeded from a backup via restoreFrom), "fresh" (a caller-proven schemaless spare — the same signal as the freshStore boolean, which stays alongside it), or "warm" (an existing persisted store, the common case). The warmBoot group carries the two warm-boot fast paths, both live.

Durable-schema replay. Boot hashes the registry-generated durable SQL and compares it with the fingerprint stored in the store. On a match the whole durable replay is skipped — schemaSkipped: true, schemaFingerprintMatch: true. On a mismatch or a store with no stored fingerprint (fresh, rebuilt) the durable schema is replayed and the new fingerprint stamped, and both flags read false. The ephemeral schema is recreated on every boot regardless (TEMP relations die with the old engine), so it is never part of the skip. A boot that adopts a caller-supplied pgwasmInstance runs no schema stage at all and leaves both flags at their conservative false.

Journal recovery. The boot-time sending → pending recovery pass is driven by a durable recovery marker that a clean settle clears:

  • Marker clear (the common warm boot — the previous run proved no sending row remains): the per-table pass is skipped entirely. journalRecoverySkipped: true, journalRecoveryRequired: false, journalTablesVisited: 0, journalRowsRecovered: 0.
  • Marker set (a crash may have left committed sending rows): the per-table updates plus the self-verifying marker clear run in one transaction. journalRecoverySkipped: false, journalRecoveryRequired: true, journalTablesVisited is the registry’s writable-journal count, and journalRowsRecovered is the real count of rows lifted sending → pending.
  • Marker absent (never initialised), a caller-supplied pgwasmInstance (pgxsinkit never touches the marker table it does not own), or a restore boot (which ignores the marker and quarantines what it recovers): one conservative unconditional pass. journalRecoverySkipped: false, journalRecoveryRequired: true, journalTablesVisited is the writable-journal count, and journalRowsRecovered is null — that pass is uncounted, so null means “not measured”, never “zero”.

Read fetchMs/applyMs as concurrent segments, not a network bill. Groups catch up in parallel on a single-threaded WASM host, so a group’s fetchMs (its settle→next-delivery wall) absorbs the OTHER groups’ apply transactions and main-thread work landing between its deliveries — it is an upper bound on network wait (“time this group spent not applying”), not pure network cost. applyMs likewise includes waiting behind a sibling group’s transaction on the single shared connection, so concurrent groups’ applyMs can overlap. Do not sum the per-group segments into a partition of totalMs. (A related reading note: phases.syncStartMs is structurally 0 when the boot is ready inside the sync-start call itself — zero eager groups, or instant catch-up.)

The provision block is what your login-dwell amortized. When it is non-null, the store was adopted from a pre-provisioned spare (the worker-mode / eager-create pattern below): provision.initdbMs is the store create cost that ran off-thread before this boot, and provision.provisionedMsBeforeBoot is how long that store sat ready before the boot claimed it — the spare’s amortized initdb, made visible. On such a boot phases.pgwasmCreateMs is null, because the create cost is reported in provision instead.

The report is reportVersion: 1 — a contract number a consumer can branch on (additive fields keep it; a breaking reshape bumps it). It is a plain structured-clone-safe object, so in worker mode it crosses the bridge unchanged; see Worker mode for the push-at-finalize vs pull-for-late-tabs semantics (onBootReport fires only for a tab attached when the boot finalizes; a later tab reads the same boot through bootReport()).

A cold store create spends ~2.5s fetching and compiling the Postgres WASM (plus the initdb WASM and the filesystem bundle) before it can open a store — and that cost otherwise lands after sign-in, on the critical path to first paint. createSyncClient takes a build option, the Postgres build the store runs on (cBuild by default). Pass createCBuild({ assets }) from @pgxsinkit/pgwasm-c, where assets is a promise of the already-fetched/compiled files (CBuildAssets: { postgresWasmModule, initdbWasmModule, fsBundle }, whose URLs cBuildArtefacts holds), and the build uses them instead of its own lazy load. Kick the fetch+compile off on an earlier screen (a login/identity picker) and hand the still-pending promise in, and the WASM cost hides behind user think-time. It is pure best-effort: a rejected warm falls back to the build’s own asset load — the warm never fails the boot. The client waits for an unfinished warm before it starts the boot pgwasm.create stamp, so phases.pgwasmCreateMs measures the create alone. (The board demo wires this from its login route and times the warm itself on a boot pgwasm build warm rail line; see apps/board/src/board/pgwasm-warm.ts.) The build must be the registry’s declared storage.build, or the boot fails with StorageBuildMismatchError before any store is touched.

In worker mode the engine loads the build’s own assets — deliberately; do not hand the tab’s compiled modules to the worker. The tab’s warm still serves the engine, by priming the same-origin HTTP cache the worker fetches from. Handing the engine a pre-compiled WebAssembly.Module benched net-negative: it forces compile-to-completion before instantiate, forfeiting the pipelining the build gets from its own streaming load, and the engine realm has no overlap window longer than the placement/handshake gap, so the compile only competes for CPU at worker spawn. (A worker entry can still pass defineSyncWorker({ build }), e.g. createCBuild({ assets }) over assets warmed inside the worker; it is checked against storage.build the same way.)

Pre-warming hides only the WASM fetch+compile — the create still spends ~1.9s on initdb and opening the store, and that cannot start until the store id is known (typically the signed-in user). To hide that cost too, create the store eagerly under a generated id on the first screen and bind it at auth. createPgwasmClient(storePath, { build }) runs the identical create the client does internally (pgwasm’s live extension — the only create-time one — plus the build’s warm and the boot pgwasm.create stamp) and returns a schemaless, deliberately engine-less instance; hand the still-pending promise to createSyncClient’s precreatedPgwasm option. Unlike pgwasmInstance (which assumes the caller applied the schema), precreatedPgwasm still lets the client run schema exec, prepare hooks, journal recovery, and store-version reconcile — so the eager create buys only initdb, and the role/registry-derived schema stays post-auth. A rejected precreatedPgwasm is caught and falls back to the storePath create path (on the same build), so the pattern is a pure accelerator, never a boot dependency. Both adopted instances are checked against storage.build through their own pg.build. Bind eager stores to users with a small localStorage registry (userId→storeId plus one unbound “spare”): create a spare on the login screen, claim it at sign-in, and GC any store that is neither mapped nor the spare. (Board demo: apps/board/src/board/store-registry.ts.)

The server side has the matching rail: createSyncServer({ logTimings: true }) (default off) emits one compact [pgxsinkit-timing] JSON line per request — the mutation route with preTxMs/txOpenMs/authMs/applyMs/totalMs (txOpenMs is the driver’s lazy connect + BEGIN, where a serverless worker’s connection cost hides). The read path emits no timing line — the control plane is a short per-subscribe call and the edge is a byte proxy, so their cost shows up as routing latency rather than as phases. Client-observed minus server totalMs isolates routing + network. On serverless hosts, mind the geometry: workers run near the caller while the database lives in one region, so a chatty write pays the cross-region round trip per statement — pin the DB-bound write function to the database’s region (Supabase: the x-region header, carried by the client’s writeRequestHeaders option) so the long hop is paid once per request instead. Pin only DB-bound functions: the control plane’s upstream is the engine and the edge’s is durable-streams, neither of which is the database, and the edge in particular wants to sit near the caller so catch-up bytes take the short hop. So keep the region header in writeRequestHeaders, never in the shared requestHeaders (which reads also send), and leave the read functions following the caller.

Worker mode: reading the rail off the main thread

Section titled “Worker mode: reading the rail off the main thread”

In a browser app you will usually attach through a SharedWorker rather than run on the calling thread — defineSyncWorker in a worker entry, attachSyncClient in the tab. A capability probe at boot decides the engine’s home: real Safari hosts the OPFS engine in that SharedWorker, while Chromium and Firefox elect a dedicated engine worker behind it. See Worker mode for the SharedWorker factory, relocation, and storage lifecycle. The local store, the stream subscriptions, and convergence stay off the main thread in either home.

  • The debug rail is forwarded and origin-tagged. A SharedWorker’s own console is invisible to the page (only chrome://inspect reaches it), so the worker forwards every rail line to each attached tab, stamped with the worker’s monotonic clock and re-printed as [pgxsinkit·w <ms>ms] … — gated by that tab’s own globalThis.__pgxsinkitDebug. Set the flag on the tab as usual; the write/read/boot phases read the same, just origin-tagged. Without the forwarding a worker-mode app would go dark, so this is on whenever the tab’s debug flag is. The front half of boot runs on the first attach, before any tab is listening, so the worker buffers those pre-attach rail lines in a bounded ring (last 500) and replays them, [replay]-marked, to the first attaching tab — so even the boot’s opening phases reach it (ADR-0034). The worker’s network traffic is invisible the same way: the subscribe calls and the stream long-polls never appear in the page’s Network panel, so a page showing no read traffic at all is normal — inspect the worker itself (chrome://inspect/#workers) for the real requests, status codes, and errors such as CORS rejections. That is also where a missing Access-Control-Expose-Headers on the edge shows itself, as a hot loop of stream requests that never advance an offset.
  • The spare store is a pre-spawned worker. The spare-store pattern from Pre-warming the Postgres build becomes a schemaless worker spawned at the login screen; claiming it binds the store id (tab-side localStorage) and attaches with config + token, so the initdb cost is already paid by the time sign-in completes. Boot-rail stamps trace it: boot spare store ensured, boot mapped store prewarm, and boot store claimed. Boot itself is sequential — the local phases run, then catch-up — so the win here is the amortized create, not an overlap.
  • ready is unchanged; per-group readiness is available. client.ready still gates on every eager group. For progressive paint, await client.groupReady(tableKey) or read status.groups — no contract change.

Initial catch-up and the convergence barrier

Section titled “Initial catch-up and the convergence barrier”

A consistency group syncs its shapes together, and the strength of that guarantee depends on when you ask. At boot and catch-up it is absolute: every shape drains, the engine’s convergence barrier is read after they all report drained, and everything held lands in one transaction — a boot never presents a half-applied cross-table state. Live, each delivered response commits on arrival, so a server transaction that touched two of the group’s tables can have one half applied while the other is still in flight; the window is the two streams’ response inter-arrival, normally milliseconds. Closing it needs a cross-stream fence the engine does not emit yet, so it is recorded as a known limitation (ADR-0056 decision 5) rather than claimed away. Two mechanisms hold the rest together, and both are worth knowing when a boot feels slow or a group appears stuck.

Dedup is per stream. Each shape’s frontier is its durable-streams offset — the same value the client persists as its resume token, so the position it dedups against and the position it resumes from cannot disagree. Offsets are meaningful only within one stream, so nothing compares them across shapes.

The steady-state commit gate is a predicate over reports, not a comparison of positions: the group commits when every shape’s most recent response asserted stream-up-to-date — so a group never commits while any member is still draining backfill, and once caught up each delivery applies as one transaction. There is no commit floor and no stale-watermark hazard to compensate for: durable-streams carries up-to-date as a response header on a live request, and a long-poll timeout returns 204 with that header set — so a quiet shape re-asserts freshness every poll cycle instead of replaying a watermark captured earlier.

Boot alignment additionally consults the engine’s convergence barrier, once. When every shape in the group has reported up-to-date at least once, the client reads GET /sync/v1/barrier — the control plane’s authenticated proxy for the engine’s convergence state — and commits what it is holding only if the barrier is satisfied. The argument is a happens-before: each stream reported drained, and a barrier read after those reports asserts the engine has nothing further to propagate.

Two terms of that barrier matter operationally, and they behave in opposite ways:

  • pendingFlips > 0 means wait. These are deferred membership flip batches — move-in and move-out — that the engine has computed but not yet propagated. The client does not align; it retries with backoff and stays on the pre-alignment gate. The term is read engine-globally, so another table’s pending work can delay an alignment but can never falsely satisfy one. Correct but slower is the deliberate trade; without it, a boot could report converged while a computed revocation was still undelivered.
  • flipFailures > 0 is terminal — never wait on it. The engine abandons a flip batch only after exhausting its propagation retries, and an abandoned batch keeps its pendingFlips count held, so waiting on it is waiting forever. The engine says so instead: it counts the abandoned batch in flipFailures and latches itself degraded, answering 503. A group reading a non-zero flipFailures refuses terminally and reports through onSyncError that the store cannot be proven consistent. Recovery is an operator restart of the engine, after which clients re-subscribe and re-snapshot — a deliberate act, not something a client can sit out.

A barrier the client cannot read is a third case and a benign one: the group stays on the pre-alignment gate and tries again on the next delivery. Boot therefore acquires a control-plane dependency — an unreachable barrier endpoint delays alignment rather than breaking it.

Resets are discovered, not announced. There is no must-refetch message on the wire. At every subscribe the client compares the stream it persisted for a shape with the one the control plane has just granted; when they differ, the stored offset addresses a different stream and means nothing there, so that shape rewinds to the start of its stream, clears its rows in the same transaction the new snapshot lands in, and re-arms alignment. Every reset path — an evicted shape, a deleted (404) or soft-deleted (410) stream, a closed stream — funnels through that one check, because each ends in a re-subscribe and a re-subscribe is what produces a new stream. Two conditions are deliberately not resets: a 403 that survives a token re-mint is a revocation (truncate that scope and unsubscribe), and a 503 from a degraded engine is the terminal refusal above. The re-subscribe is automatic — a read that ends re-subscribes its whole consistency group with backoff, and the client reports a stream-degraded status until the next delivered batch clears it. The same subscribe comparison also catches a revocation that landed while the client was offline: a scope the control plane no longer grants is cleared — the scope’s rows for the shared tier, the whole table for the private tier — before the group reports ready, so a boot never presents rows the subject may no longer read.

See ADR-0056 for the full rationale.