Operating in production
The read and write primitives are fast, but a few runtime and deployment properties around them decide whether a live app feels fast. None are toolkit bugs — they are how serverless edges, browser HTTP, and CDN-shaped caching behave — but every one of them was hit dogfooding the demo, so they are collected here with their fixes. If writes or sync feel slow in a real browser but your server benchmarks are fast, the cause is almost always on this page.
Convergence cadence: event-driven, with the interval as a fallback
Section titled “Convergence cadence: event-driven, with the interval as a fallback”When you pass an autoSync trigger to createSyncClient, the client drives a flush → reconcile pass.
The pass is event-driven: the client calls requestPass() the moment a mutation is enqueued, so a
local write flushes to the server immediately — it does not wait for the trigger’s next interval
tick. The interval (createBrowserConvergenceTrigger({ intervalMs }), default 1.5s) is therefore
only a fallback for retries, recovery, and cross-tab wake-ups.
Because the happy path is event-driven, you should make that fallback interval long. A short
interval is the dominant idle cost: every local query carries ~50ms of WASM overhead and serializes on
the one worker thread, and an unconditional reconcile each tick fires the store’s live-query notifications,
re-running every mounted query. The toolkit already idle-skips an empty reconcile, but the cheapest idle
board is still a rare interval — the board demo runs intervalMs: 15_000, taking idle CPU from ~70% of a
core to ~2% with no change to convergence latency (latency is bounded by the read-path echo, not the
interval).
In worker mode the worker owns this loop: defineSyncWorker’s
convergenceIntervalMs is the same fallback interval and already defaults to 15s. Writes still flush
on enqueue (the write RPC requests a pass) and tabs forward online/visibilitychange as wake signals,
so the same rule holds — the interval is a retry/recovery sweep, not the write path.
Local write latency: the durability preference (relaxed by default)
Section titled “Local write latency: the durability preference (relaxed by default)”durability is declared once, on the registry — SyncRegistryDefinition.storage.durability
("relaxed" | "strict", default "relaxed"). It is not a minting-surface, worker-entry, or attach-site
option: whether losing the last not-yet-flushed action is acceptable is decided by what the data IS, so
one declaration binds every open of every store minted from that registry — no tab can ever disagree with
another. The physical behavior depends on the capability-selected backend:
- IndexedDB: strict flushes the whole datadir synchronously at the end of every query, setting a ~100–200ms optimistic-write floor. Relaxed returns before that snapshot and schedules it asynchronously.
- OPFS-repacked: pgwasm always awaits the store’s sync. Relaxed asserts VFS health and runs any due deferred repack without an ordinary physical flush; strict flushes arena data before metadata. Initialization, repack activation, and open-state close keep strict ordering in both modes.
The resolved mode is stamped on the boot pgwasm.create rail line. A capability fallback from opfs to
idbfs keeps the registry-declared durability unchanged.
What you trade. On idb, writes since the last completed snapshot are at risk only if the browser terminates before both their journal rows reach the write API and the scheduled snapshot lands. On OPFS-repacked, relaxed recovery returns the longest valid stable metadata-log prefix, so an unflushed suffix may be absent; a returned strict boundary is stable under the browser-termination model. Synced tables are server-recoverable by construction. Your own local-only tables have no such copy.
Edge serverless cold starts
Section titled “Edge serverless cold starts”On a serverless Edge platform a worker is suspended when idle and evicted after longer idle, so the first write after a quiet period pays a cold start while steady-state writes are instant. Measured on the self-hosted Supabase edge-runtime: a warm write applies in ~20ms; a write to a worker suspended ~15s pays ~0.45s (a Postgres reconnect on resume); a write to a worker whose module cache is cold pays ~5.8s (a fresh isolate re-imports the whole bundle). Drag the first card after the board sits idle and that cold worker is the entire delay — not the sync rail.
This is a property of the serverless deployment target, not pgxsinkit: the same functions on a long-lived Bun or Deno process (one warm process, a pooled connection) or on a managed warm pool have no cold start. Two mitigations if you stay serverless:
- Keep the worker warm with a periodic cheap request. The cheapest request that still reaches the
worker is a no-op write — an empty
{"mutations":[]}POST, rejected at request validation before any DB work. A small sidecar pinging it every ~8s keeps writes at ~20ms after idle. - Set the worker’s wall-clock timeout above your longest held-open long-poll. A durable-streams
long-poll is held until data arrives or the hold expires, so a stock 60s worker budget recycles live
subscriptions constantly; the board demo runs its read functions at
EDGE_WORKER_TIMEOUT_MS: "600000"(ten minutes). See Deploying the server.
Caching: the control plane is private, the edge is the cacheable surface
Section titled “Caching: the control plane is private, the edge is the cacheable surface”The read path splits into two ingress points with opposite caching postures, and getting them the wrong way round is the mistake this section exists to prevent.
The control plane is never cacheable. /sync/v1/subscribe, /sync/v1/refresh,
/sync/v1/release and /sync/v1/barrier answer per-subject questions and mint stream tokens. Force
cache-control: no-store on it. A cached subscribe answer is a capability served to the wrong subject; a cached
barrier answer can license an alignment against an engine that has already lost membership effects
(which is why the barrier’s own cache window defaults to zero).
The edge is the only surface a CDN may share, and it can only do that if it is a separate origin — the cache key is the URL, so a shared tier mounted beside the control plane can never be fronted without also caching the private one.
The stream token must be excluded from the cache key. This is a correctness condition for the shared tier, not a tuning detail: include it and every subscriber gets a unique key, which destroys exactly the sharing the tier exists for.
Catch-up reads are where the bytes are, and durable-streams already marks them cacheable — a catch-up
response carries cache-control: public, max-age=60, stale-while-revalidate=300 and an etag. Live
long-polls are not cacheable and are not marked so. Respect the origin’s cache-control verbatim
rather than overriding TTLs in either direction, key on the full URL minus the token, and set
proxy/CDN read timeouts above the long-poll hold so a completed poll is not cut off. Cacheability
is a property of the shared tier only: a private-tier shape’s stream is per-subject by
construction and no cache can share it.
Revocation latency is deployment-shaped, and you have to state which shape yours is. Behind a third-party CDN serving hits, it is bounded by the stream token’s TTL (5 minutes by default). Behind a cache sitting behind the edge’s own verifier, it is entitlement-propagation latency instead. Private-tier grants are re-authorized on every re-mint too — the control plane recompiles the shape against the subject’s current claims and revokes the grant if it no longer compiles to the shape that was granted — so a claims change (a role removed, a tenant switched) is bounded by the same TTL as an entitlement change, without the subject having to re-authenticate for it to take effect.
The re-mint only narrows, though. A scope a subject gains after subscribing — a membership acquired mid-session, or a shape that was refused at subscribe — is granted at the next subscribe: a boot, or the automatic restart that follows a stream ending or a grant being revoked. It does not arrive within the TTL on its own. Widening an existing private-tier grant is the exception, by way of the fingerprint: a claims change that compiles to a different predicate is revoked at the re-mint, and the restart re-subscribes with the new claims.
Engine shape claims: a closed subscription gives them back
Section titled “Engine shape claims: a closed subscription gives them back”Every grant a subscribe hands out is a named claim on an engine shape, and a live claim is what
keeps that shape active — native reads terminate on durable-streams, so a claim is the only thing
telling the engine a reader exists. Closing a subscription releases those claims (POST /sync/v1/release), which is what lets an idle shape go dormant and eventually be evicted promptly.
A claim is also a lease. It counts as live only while it is renewed within the engine’s
CIRCUITS_SHAPE_IDLE_SECS window, and the control plane renews every still-authorized grant
on the token re-mint. So claims of a client that crashed or was unloaded before its release left
the browser are reclaimed by the engine within one lease window rather than held until a restart —
nothing renews them once the session is gone. Because each claim is named, the release is
idempotent: repeating one is the same act once, and cannot take another subscriber’s claim.
That pairing has a configuration requirement, and the control plane enforces it: if the engine’s lease
window is shorter than twice the stream-token TTL, subscribe and refresh answer
503 {"error":"sync engine lease window shorter than the token TTL"} rather than serving sessions whose
shapes would lapse after a single missed refresh. Raise CIRCUITS_SHAPE_IDLE_SECS, lower
ttlSeconds, or set the idle window to 0 (which disables dormancy, and with it leases).
The browser connection budget (serve the gateway over HTTP/2)
Section titled “The browser connection budget (serve the gateway over HTTP/2)”@pgxsinkit/client’s reader holds one live long-poll connection open per synced stream. A client
subscribing to six streams keeps six connections continuously busy — and on the shared tier a subject
holding K scopes of one shape holds K streams, not one — while browsers cap HTTP/1.1 at ~6
connections per origin. So over plain HTTP those long-polls consume every slot and the write
request (same origin) gets Stalled in the browser’s connection queue for a whole long-poll cycle
before it is even dispatched. Serve the gateway over HTTP/2 (or HTTP/3), which multiplexes every
request over one connection, so the cap never binds. This is mandatory rather than advisory:
durable-streams serves one stream per request and the Rust server is HTTP/1.1-only with no TLS. It only
bites a local stack served over plain http://; any production ingress (Cloud Supabase, an
istio/Envoy gateway, a TLS reverse proxy) already speaks HTTP/2. Full detail and the symptom check are
in Deploying the server.
Read-path apply failures hold, they don’t diverge
Section titled “Read-path apply failures hold, they don’t diverge”If a local commit into the store keeps failing — a bad local migration, a storage/quota error, a corrupt
local store — the read path retries with backoff and then latches into a degraded phase instead of
advancing past the change it could not apply. It holds the read cache at the last good commit, so the
client never silently diverges from the server. The failure is surfaced through the onSyncError
callback you pass to createSyncClient, and the runtime’s status reports the degraded phase.
How it recovers depends on why it degraded:
- A subscribe failure — a control plane that is unreachable or keeps erroring — takes the same
recoverable
degradedphase, withlastErrornaming it, and subscribe keeps retrying with backoff behind it. It never masks a more actionable status: an auth failure wins over it, and a commit failure (below) is never overwritten by it. - A read-stream error (a dropped stream connection) clears automatically on the next successful
batch. The reader retries 429, every 5xx and network failures with backoff. A 401/403
gets one re-mint and one retry, so an expired stream token is replaced instead of presenting as a
stall, and a token refused again ends the read, as every other 4xx does, rather than being swallowed.
A connection that dies by hanging rather than failing (a pulled cable mid-long-poll) produces
neither a failure nor a batch, so a runtime claiming
readyis additionally held to a read-silence window (readSilenceMs, default 45s — a healthy stream’s long-poll always cycles well inside it): silence past the window dropsreadyto the same self-recoveringdegraded. That is the phase to key an offline/“connection needed” surface off;navigator.onLineis not a substitute. Recovery is wired to delivered traffic — any batch arriving clears it and re-arms the window. - An auth failure on the read path is its own phase, not a
degraded. A persistent 401/403 from the control plane puts the runtime inauth-neededwhile subscribe keeps retrying for a fresh token, so the app can prompt for re-login; it clears the moment a batch lands. This is the phase behind an expired session, and it is deliberately distinguishable from an unreachable deployment. - A commit failure is sticky — fetching can keep succeeding while applies fail — so it clears only
on the next commit that succeeds, after you fix the underlying cause, or on a client restart.
There is no separate reset call for this;
recoverSendingrebuilds the write journal, not the read frontier. - A degraded engine is terminal, not something to wait out. If the engine abandons a membership
flip batch after exhausting its retries, those effects are lost and it latches itself degraded
(answering
503). The client refuses to align rather than holding, and reports the refusal throughonSyncError. Recovery is an operator restart, after which clients re-subscribe and re-snapshot.
Watching the event lane’s Outbox
Section titled “Watching the event lane’s Outbox”If your app appends events (the model is in The event lane), the client gives you
two surfaces to operate it with, and neither is optional reading before you ship one. The knobs on the other
side of them — batch caps, the fallback interval, backoff bounds, and the per-stream fairness cap — are the
client’s events option, documented under
Tuning the flush.
The drain signal is client.onOutboxStatus(({ empty }) => …) — the empty ↔ non-empty transitions, with
the current state delivered on subscribe (await client.outboxStatus() is the one-shot pull). “Empty” means
nothing is awaiting a server verdict. Use it to invalidate a view that composes pending events with
down-synced aggregates. It carries no count deliberately — a count that updates only on transitions is stale
by construction — so query the Outbox table (getOutboxTable(registry)) if you want one.
If that composed view dips between the ack and your consumer’s fold, that is the gap
the acked ledger closes:
events.ackedRetentionMs keeps an acked row, stamped, instead of deleting it at the ack. A retained row is
not pending — it does not hold the drain signal, is never re-sent, and never blocks a destroy() — so a
composition reads acked_at_us IS NULL for “still owed” and treats the rest as its grace window.
The verdicts are client.onEventLaneReport(cb): per flush pass, the terminal non-acked verdicts, the
deferred ones, and the lane’s batch-level backoff transitions. Subscribe for the app’s lifetime, not
per screen — the subscription is ephemeral, and once a terminal row is deleted the Outbox cannot answer for
it. (With nothing subscribed the library logs each report at warn level rather than dropping it.)
Read them this way:
refused— youreventGatedeclined it. Expected, not an error.rejected— a schema-invalid or oversized payload for a known stream. The library validates at append, so in practice this means a non-library caller or a broken deployment: treat it as a bug.deferred— the server does not (yet) know that stream. This is not a failure and not terminal: it is ordinary rollout skew, the rows stay in the Outbox, and they drain when the server deploy lands. A burst right after a client release is the deploy order; a burst that never clears means deployment skew — the server’s registry does not declare that stream (it is the registry, and only the registry, that decides this verdict). Never “clean up” the Outbox in response to it.
A lane stuck in batch-level backoff is a different diagnosis, and the report’s backoff transitions
(plus a persistent 503 with Retry-After) are its signal. The server knows the stream but could not
enqueue the batch, and it fails the batch whole — nothing is enqueued and no per-event verdicts are
issued. The first thing to check is the queue itself: registering a stream requires generating and applying
the --events migration, and without it the endpoint enqueues onto a pgmq queue that does not exist.
(Then: the database’s reachability from the ingress, and the server log line the 503 always writes.)
operations_log records user content
Section titled “operations_log records user content”The optional write-side audit log (operationsLog: { enabled: true }, off by default) persists every
mutation — table, kind, and payload — including the raw body of mutations that failed validation.
Those payloads contain whatever your users typed. So treat the operations_log table as sensitive:
restrict access and set a retention policy, and leave it disabled in environments where that content
should not sit at rest.
Its mutation_id column is text, not uuid — deliberately. This is the server-side tier of a
two-tier invariant: pgxsinkit’s public write surface (the HTTP route’s request/ack schemas and the
client’s mutation_id UUID journal) is UUID-only by contract, but the generated apply function and
this log accept an opaque text id for one narrow case — a direct, server-side caller that derives
child envelopes with composite ids (${parentMutationId}:<tag>:<n>) and invokes the apply function
itself, never crossing the HTTP route or the client journal. A non-UUID id can never reach the
UUID-typed public surface.
Debugging latency: globalThis.__pgxsinkitDebug
Section titled “Debugging latency: globalThis.__pgxsinkitDebug”@pgxsinkit/client ships opt-in, timestamped instrumentation that traces a write through every phase —
exactly what localises a “writes are slow” problem to a single hop. It is off by default and adds
nothing to a normal run; enable it from the console or before the client boots:
globalThis.__pgxsinkitDebug = true; // then reproduce; filter the console to "pgxsinkit" + enable VerboseEach line is stamped with a monotonic millisecond clock, so you read the gaps between phases directly:
mutation staged {mutationId, table}(the write’s origin — correlate by id with the sent/acked lines)convergence pass requested→convergence flush→convergence reconcile(with durations)board-write auth token resolved {ms}(a stalling per-requestgetSession()shows up here)board-write responded {status, ms}(a cold edge worker, or a browser connection stall, shows up here)live-query register/retained/dedup-hit/evicted(the local read model’s subscription bookkeeping) andlive query updated → re-render(the final UI hop, from@pgxsinkit/react)boot pgwasm.create→boot local schema→boot journal recovery→boot store-version reconcile→boot sync start→boot client ready(the boot phases, for attributing a slow first paint to one of them)boot pgwasm.create phase(on an OPFS store:module-loaded,directory-ready,store-opened,pgwasm-ready), so a create that never returns is attributable to one stepboot pgwasm build warm(only when the host times its pre-warm on the rail, as the board demo does — see next section)
The read path itself does not emit rail lines. Its transport, @pgxsinkit/client’s own long-poll
reader, is not instrumented on this rail. Read it at the two boundaries instead: boot client ready (fired when
the first initial sync completes) and the runtime’s status phase, which names syncing, ready,
degraded (with lastError) and auth-needed. For per-request detail, inspect the network layer
directly — and in worker mode that means inspecting the worker, not the page (see below).
The structured BootReport — measure before you optimize
Section titled “The structured BootReport — measure before you optimize”The rail lines above are for a human mid-debug: you eyeball the gaps as they scroll. For numbers a
machine can keep — a dashboard series, a CI budget gate, an honest before/after — every boot also builds
a structured, versioned BootReport, independently of the rail, so it exists whether or not
__pgxsinkitDebug is on. Boot performance regressed repeatedly because it was optimized on guesses (the
suspected bottleneck was rarely the real one); the report is the evidence to start from instead (ADR-0034).
Read it by push, by pull, or both:
const client = await createSyncClient({ registry, controlPlaneUrl, streamBaseUrl, batchWriteUrl, onBootReport: (report) => { // fires exactly once, at boot completion — ship it to a dashboard, or assert a CI budget metrics.timing("boot.total", report.totalMs); },});
const report = await client.bootReport(); // the most recent completed boot, or null before the first syncreport.totalMs is boot start → every eager group caught up; phases decomposes the local work (store
create, schema exec, journal recovery, store-version reconcile, sync start, catch-up) and groups[] breaks
out the per-consistency-group boot catch-up (rows, requests, fetchMs, applyMs, start/ready offsets).
localReadReadyMs and writeReadyMs mark the staged-boot crossings (ADR-0041; additive, reportVersion
stays 1). localReadReadyMs is boot start → cached reads are safe (store open, schema compatible, reconcile
done — zero network); it is the moment attachSyncClient / createSyncClient resolve. writeReadyMs is boot
start → the write runtime + boot recovery finished (enqueue is safe). Both are null when the boot rejected
before reaching that stage. The gap from localReadReadyMs to totalMs is the whole-sync catch-up a cached
paint no longer waits on.
storeKind names how the store presented at boot — "restored" (seeded from a backup via
restoreFrom), "fresh" (a caller-proven schemaless spare — the same signal as the freshStore boolean,
which stays alongside it), or "warm" (an existing persisted store, the common case). The warmBoot group
carries the two warm-boot fast paths, both live.
Durable-schema replay. Boot hashes the registry-generated durable SQL and compares it with the
fingerprint stored in the store. On a match the whole durable replay is skipped —
schemaSkipped: true, schemaFingerprintMatch: true. On a mismatch or a store with no stored fingerprint
(fresh, rebuilt) the durable schema is replayed and the new fingerprint stamped, and both flags read false.
The ephemeral schema is recreated on every boot regardless (TEMP relations die with the old engine), so it
is never part of the skip. A boot that adopts a caller-supplied pgwasmInstance runs no schema stage at all
and leaves both flags at their conservative false.
Journal recovery. The boot-time sending → pending recovery pass is driven by a durable recovery marker
that a clean settle clears:
- Marker clear (the common warm boot — the previous run proved no
sendingrow remains): the per-table pass is skipped entirely.journalRecoverySkipped: true,journalRecoveryRequired: false,journalTablesVisited: 0,journalRowsRecovered: 0. - Marker set (a crash may have left committed
sendingrows): the per-table updates plus the self-verifying marker clear run in one transaction.journalRecoverySkipped: false,journalRecoveryRequired: true,journalTablesVisitedis the registry’s writable-journal count, andjournalRowsRecoveredis the real count of rows liftedsending → pending. - Marker absent (never initialised), a caller-supplied
pgwasmInstance(pgxsinkit never touches the marker table it does not own), or a restore boot (which ignores the marker and quarantines what it recovers): one conservative unconditional pass.journalRecoverySkipped: false,journalRecoveryRequired: true,journalTablesVisitedis the writable-journal count, andjournalRowsRecoveredisnull— that pass is uncounted, sonullmeans “not measured”, never “zero”.
Read fetchMs/applyMs as concurrent segments, not a network bill. Groups catch up in parallel on
a single-threaded WASM host, so a group’s fetchMs (its settle→next-delivery wall) absorbs the OTHER
groups’ apply transactions and main-thread work landing between its deliveries — it is an upper bound on
network wait (“time this group spent not applying”), not pure network cost. applyMs likewise
includes waiting behind a sibling group’s transaction on the single shared connection, so concurrent
groups’ applyMs can overlap. Do not sum the per-group segments into a partition of totalMs. (A related
reading note: phases.syncStartMs is structurally 0 when the boot is ready inside the sync-start call
itself — zero eager groups, or instant catch-up.)
The provision block is what your login-dwell amortized. When it is non-null, the store was adopted
from a pre-provisioned spare (the worker-mode / eager-create pattern below): provision.initdbMs is the
store create cost that ran off-thread before this boot, and provision.provisionedMsBeforeBoot is how
long that store sat ready before the boot claimed it — the spare’s amortized initdb, made visible. On
such a boot phases.pgwasmCreateMs is null, because the create cost is reported in provision instead.
The report is reportVersion: 1 — a contract number a consumer can branch on (additive fields keep it; a
breaking reshape bumps it). It is a plain structured-clone-safe object, so in worker mode it crosses the
bridge unchanged; see Worker mode for the push-at-finalize vs pull-for-late-tabs
semantics (onBootReport fires only for a tab attached when the boot finalizes; a later tab reads the same
boot through bootReport()).
Pre-warming the Postgres build
Section titled “Pre-warming the Postgres build”A cold store create spends ~2.5s fetching and compiling the Postgres WASM (plus the initdb WASM and the
filesystem bundle) before it can open a store — and that cost otherwise lands after sign-in, on the
critical path to first paint. createSyncClient takes a build option, the
Postgres build the store runs on (cBuild by default). Pass
createCBuild({ assets }) from @pgxsinkit/pgwasm-c, where assets is a promise of the
already-fetched/compiled files (CBuildAssets: { postgresWasmModule, initdbWasmModule, fsBundle },
whose URLs cBuildArtefacts holds), and the build uses them instead of its own lazy load. Kick the
fetch+compile off on an earlier screen (a login/identity picker) and hand the still-pending promise
in, and the WASM cost hides behind user think-time. It is pure best-effort: a rejected warm falls back to
the build’s own asset load — the warm never fails the boot. The client waits for an unfinished warm
before it starts the boot pgwasm.create stamp, so phases.pgwasmCreateMs measures the create alone.
(The board demo wires this from its login route and times the warm itself on a
boot pgwasm build warm rail line; see apps/board/src/board/pgwasm-warm.ts.) The build must be the
registry’s declared storage.build, or the boot fails with StorageBuildMismatchError before any store
is touched.
In worker mode the engine loads the build’s own assets — deliberately; do not hand the tab’s compiled
modules to the worker. The tab’s warm still serves the engine, by priming the same-origin HTTP cache the worker fetches
from. Handing the engine a pre-compiled WebAssembly.Module benched net-negative: it forces
compile-to-completion before instantiate, forfeiting the pipelining the build gets from its own streaming
load, and the engine realm has no overlap window longer than the placement/handshake gap, so the compile
only competes for CPU at worker spawn. (A worker entry can still pass defineSyncWorker({ build }), e.g.
createCBuild({ assets }) over assets warmed inside the worker; it is checked against storage.build the
same way.)
Pre-warming hides only the WASM fetch+compile — the create still spends ~1.9s on initdb and opening
the store, and that cannot start until the store id is known (typically the signed-in user). To hide that
cost too, create the store eagerly under a generated id on the first screen and bind it at auth.
createPgwasmClient(storePath, { build }) runs the identical create the client does internally
(pgwasm’s live extension — the only create-time one — plus the build’s warm and the
boot pgwasm.create stamp) and returns a schemaless, deliberately engine-less instance; hand the still-pending promise to createSyncClient’s precreatedPgwasm option.
Unlike pgwasmInstance (which assumes the caller applied the schema), precreatedPgwasm still lets the
client run schema exec, prepare hooks, journal recovery, and store-version reconcile — so the eager create
buys only initdb, and the role/registry-derived schema stays post-auth. A rejected precreatedPgwasm is
caught and falls back to the storePath create path (on the same build), so the pattern is a
pure accelerator, never a boot dependency. Both adopted instances are checked against storage.build
through their own pg.build. Bind eager stores to users with a small localStorage registry
(userId→storeId plus one unbound “spare”): create a spare on the login screen, claim it at sign-in, and GC
any store that is neither mapped nor the spare. (Board demo: apps/board/src/board/store-registry.ts.)
The server side has the matching rail: createSyncServer({ logTimings: true }) (default off) emits one
compact [pgxsinkit-timing] JSON line per request — the mutation route with
preTxMs/txOpenMs/authMs/applyMs/totalMs (txOpenMs is the driver’s lazy connect + BEGIN, where
a serverless worker’s connection cost hides). The read path emits no timing line — the control plane
is a short per-subscribe call and the edge is a byte proxy, so their cost shows up as routing latency
rather than as phases. Client-observed minus server totalMs isolates routing + network. On serverless
hosts, mind the geometry: workers run near the caller while the database lives in one region, so a
chatty write pays the cross-region round trip per statement — pin the DB-bound write function to the
database’s region (Supabase: the x-region header, carried by the client’s writeRequestHeaders
option) so the long hop is paid once per request instead. Pin only DB-bound functions: the control
plane’s upstream is the engine and the edge’s is durable-streams, neither of which is the database, and
the edge in particular wants to sit near the caller so catch-up bytes take the short hop. So keep
the region header in writeRequestHeaders, never in the shared requestHeaders (which reads also
send), and leave the read functions following the caller.
Worker mode: reading the rail off the main thread
Section titled “Worker mode: reading the rail off the main thread”In a browser app you will usually attach through a SharedWorker rather than run on the calling thread —
defineSyncWorker in a worker entry, attachSyncClient in the tab. A capability probe at boot decides
the engine’s home: real Safari hosts the OPFS engine in that SharedWorker, while Chromium and Firefox
elect a dedicated engine worker behind it. See Worker mode for the SharedWorker
factory, relocation, and storage lifecycle. The local store, the stream subscriptions, and convergence stay off
the main thread in either home.
- The debug rail is forwarded and origin-tagged. A SharedWorker’s own
consoleis invisible to the page (onlychrome://inspectreaches it), so the worker forwards every rail line to each attached tab, stamped with the worker’s monotonic clock and re-printed as[pgxsinkit·w <ms>ms] …— gated by that tab’s ownglobalThis.__pgxsinkitDebug. Set the flag on the tab as usual; the write/read/boot phases read the same, just origin-tagged. Without the forwarding a worker-mode app would go dark, so this is on whenever the tab’s debug flag is. The front half of boot runs on the first attach, before any tab is listening, so the worker buffers those pre-attach rail lines in a bounded ring (last 500) and replays them,[replay]-marked, to the first attaching tab — so even the boot’s opening phases reach it (ADR-0034). The worker’s network traffic is invisible the same way: the subscribe calls and the stream long-polls never appear in the page’s Network panel, so a page showing no read traffic at all is normal — inspect the worker itself (chrome://inspect/#workers) for the real requests, status codes, and errors such as CORS rejections. That is also where a missingAccess-Control-Expose-Headerson the edge shows itself, as a hot loop of stream requests that never advance an offset. - The spare store is a pre-spawned worker. The spare-store pattern from
Pre-warming the Postgres build becomes a schemaless worker spawned at
the login screen; claiming it binds the store id (tab-side
localStorage) and attaches with config + token, so theinitdbcost is already paid by the time sign-in completes. Boot-rail stamps trace it:boot spare store ensured,boot mapped store prewarm, andboot store claimed. Boot itself is sequential — the local phases run, then catch-up — so the win here is the amortized create, not an overlap. readyis unchanged; per-group readiness is available.client.readystill gates on every eager group. For progressive paint,await client.groupReady(tableKey)or readstatus.groups— no contract change.
Initial catch-up and the convergence barrier
Section titled “Initial catch-up and the convergence barrier”A consistency group syncs its shapes together, and the strength of that guarantee depends on when you ask. At boot and catch-up it is absolute: every shape drains, the engine’s convergence barrier is read after they all report drained, and everything held lands in one transaction — a boot never presents a half-applied cross-table state. Live, each delivered response commits on arrival, so a server transaction that touched two of the group’s tables can have one half applied while the other is still in flight; the window is the two streams’ response inter-arrival, normally milliseconds. Closing it needs a cross-stream fence the engine does not emit yet, so it is recorded as a known limitation (ADR-0056 decision 5) rather than claimed away. Two mechanisms hold the rest together, and both are worth knowing when a boot feels slow or a group appears stuck.
Dedup is per stream. Each shape’s frontier is its durable-streams offset — the same value the client persists as its resume token, so the position it dedups against and the position it resumes from cannot disagree. Offsets are meaningful only within one stream, so nothing compares them across shapes.
The steady-state commit gate is a predicate over reports, not a comparison of positions: the group
commits when every shape’s most recent response asserted stream-up-to-date — so a group never
commits while any member is still draining backfill, and once caught up each delivery applies as one
transaction. There is no commit floor and no stale-watermark hazard to compensate for: durable-streams
carries up-to-date as a response header on a live request, and a long-poll timeout returns 204 with
that header set — so a quiet shape re-asserts freshness every poll cycle instead of replaying a watermark
captured earlier.
Boot alignment additionally consults the engine’s convergence barrier, once. When every shape in the
group has reported up-to-date at least once, the client reads
GET /sync/v1/barrier — the control plane’s authenticated proxy for the engine’s convergence state —
and commits what it is holding only if the barrier is satisfied. The argument is a happens-before: each
stream reported drained, and a barrier read after those reports asserts the engine has nothing further
to propagate.
Two terms of that barrier matter operationally, and they behave in opposite ways:
pendingFlips > 0means wait. These are deferred membership flip batches — move-in and move-out — that the engine has computed but not yet propagated. The client does not align; it retries with backoff and stays on the pre-alignment gate. The term is read engine-globally, so another table’s pending work can delay an alignment but can never falsely satisfy one. Correct but slower is the deliberate trade; without it, a boot could report converged while a computed revocation was still undelivered.flipFailures > 0is terminal — never wait on it. The engine abandons a flip batch only after exhausting its propagation retries, and an abandoned batch keeps itspendingFlipscount held, so waiting on it is waiting forever. The engine says so instead: it counts the abandoned batch inflipFailuresand latches itself degraded, answering503. A group reading a non-zeroflipFailuresrefuses terminally and reports throughonSyncErrorthat the store cannot be proven consistent. Recovery is an operator restart of the engine, after which clients re-subscribe and re-snapshot — a deliberate act, not something a client can sit out.
A barrier the client cannot read is a third case and a benign one: the group stays on the pre-alignment gate and tries again on the next delivery. Boot therefore acquires a control-plane dependency — an unreachable barrier endpoint delays alignment rather than breaking it.
Resets are discovered, not announced. There is no must-refetch message on the wire. At every
subscribe the client compares the stream it persisted for a shape with the one the control plane has
just granted; when they differ, the stored offset addresses a different stream and means nothing there,
so that shape rewinds to the start of its stream, clears its rows in the same transaction the new
snapshot lands in, and re-arms alignment. Every reset path — an evicted shape, a deleted (404) or
soft-deleted (410) stream, a closed stream — funnels through that one check, because each ends in a
re-subscribe and a re-subscribe is what produces a new stream. Two conditions are deliberately not
resets: a 403 that survives a token re-mint is a revocation (truncate that scope and unsubscribe),
and a 503 from a degraded engine is the terminal refusal above. The re-subscribe is automatic — a read
that ends re-subscribes its whole consistency group with backoff, and the client reports a
stream-degraded status until the next delivered batch clears it. The same subscribe comparison also
catches a revocation that landed while the client was offline: a scope the control plane no longer
grants is cleared — the scope’s rows for the shared tier, the whole table for the private tier — before
the group reports ready, so a boot never presents rows the subject may no longer read.
See ADR-0056 for the full rationale.