# Cross-cluster sync

**Experimental:** requires an explicitly enabled build based on current development source. It is not part of public APT installation or the ordinary build's compatibility contract. Obtain a matched server, tool, and SDK release before evaluation.

## Synchronize writable clusters through S3

Each group has a stable group ID included in the bucket path. Each enrolled cluster writes original local mutations to its own origin streams and ingests changes from other enrolled origins. Imported changes are not republished as original writes, preventing endless forwarding.

Changes are batched by age or size; one-minute batches are a starting policy. This is asynchronous exchange, not low-latency synchronous replication across regions. Multiple clusters can continue accepting local writes, subject to their own health and configured storage limits.

## Resolve concurrent writes

The receiver uses the implemented record-version conflict ordering rather than arrival order alone. Do not assume independent column edits in two clusters merge field by field. Concurrent changes can overwrite a row/value according to that ordering. Keep clocks synchronized and validate conflict behavior for the intended application before deployment.

Deletes and expiration carry version/expiry information so delayed imports do not simply recreate older values. Group enrollment, origin identity, sequence continuity, and object verification are part of acceptance; do not manually edit stream objects to bypass a gap.

## Scheduled sync checkpoints

Optional scheduled capture publishes immutable per-origin checkpoints and assembles complete recovery sets for enrolled target clusters. A new cluster can load a selected complete set and then ingest the tail rather than replay all historical changes. Scheduling is disabled by default and has explicit disk, upload, work, spool-size, and record-count limits.

Checkpoint work is paced below foreground operations, but consumes disk, memory, and CPU. Inspect `smkv-ctl checkpoints` and the dashboard Activity view. An incomplete origin collection is not a complete recovery set.

The experimental `smkv-sync-checkpoint` tool supports bootstrap from an explicit `--checkpoint-set` key into fresh data directories. Configuration and enrollment must match the target identity. Request the matching experimental configuration/runbook before running this workflow; ordinary backup import is a different command and format.

## What has been verified

An isolated DigitalOcean Spaces drill used 10,000 records with 1 KiB values, three source nodes, RF2, a second writable origin, and a fresh target. It verified publisher restart after a committed upload, complete-set publication, fresh bootstrap, and both record copies, including deletion, expiration, and competing writes.

At approximately 800 reads and 200 writes per second, the steady checkpoint phase had zero request errors and read/write p99 around 5/16 ms. The injected crash produced temporary request errors. This small, same-region, cache-hot workload is not a general performance guarantee.

## Remaining boundaries

Tail verification took about 240 seconds; replication, scheduling, and observation delays still need investigation. Bootstrap under continuous foreground traffic, larger datasets, longer S3 outages, and cross-region behavior need further qualification. Automatic retention cleanup is not implemented; do not delete batches or checkpoints based on age alone.

Keep [manual full backups](backup.html) independent. Sync propagates changes, including mistakes; it is not a substitute for recovery archives and tested restoration.
