simple magicDATA

SMKV / Operate

Restart, upgrade, and resize

Applies to: current development clusters unless stated otherwise. Perform maintenance on one participant at a time, with verified recovery between steps.

Restart safely

For packaged data services:

sudo systemctl stop smkv
sudo systemctl start smkv
sudo journalctl -u smkv -n 100 --no-pager

SIGTERM and Ctrl-C drain accepted work and sync partitions. Allow adequate shutdown time; SIGKILL or a service stop timeout bypasses graceful completion. Preserve the entire data directory, including cluster identity and incarnation metadata. A live directory is not a file-copy backup.

Before stopping a primary, confirm a ready replica and understand the selected failover mode. Explicit-ownership clusters do not automatically promote replicas. Quorum mode can promote a validated replica but may still lose recently acknowledged asynchronous writes. Losing two of three authority voters prevents quorum-confirmed operations, including reads.

Qualify upgrades

Record installed package/build identities, current map, compatibility declarations, and a verified backup. New development tools expose build-info and release-info; older binaries may not. Check their help instead of assuming support.

A backward-compatible release can be rolled through a working cluster after qualification. Update authority voters one at a time first, preserving their state, then data nodes one at a time with readiness and replica checks. Plan transient request failures. Packaging does not imply a zero-interruption upgrade.

For incompatible storage or protocol changes, create a new cluster and use a supported migration path with explicit validation. Do not open upgraded files using older binaries merely because a package downgrade succeeds. Keep rollback binaries and data-format constraints in the change plan.

Add a data node

Current quorum clusters support operator-driven join and drain. This is not autoscaling. Upgrade all voters and data nodes to a compatible resizing build first. Use the joining-node template, replacing its addresses and budgets.

smkv-ctl resize-plan --join joining-node.json --output next-map.json \
  --authority 10.0.0.11:7390 --authority 10.0.0.12:7390 --authority 10.0.0.13:7390

Planning reads current budgets and creates a map without changing ownership. Save the returned source map as current-map.json. Start the new node with a fresh directory, the original partition count, matching budgets, --cluster-map current-map.json --join-map next-map.json, its new node ID, and all three authority seeds.

smkv-ctl resize --map next-map.json --timeout-seconds 3600 \
  --authority 10.0.0.11:7390 --authority 10.0.0.12:7390 --authority 10.0.0.13:7390
smkv-ctl resize-status \
  --authority 10.0.0.11:7390 --authority 10.0.0.12:7390 --authority 10.0.0.13:7390

Copy preparation serves traffic; cutover briefly fences writes and verifies both destination copies. Reserve headroom for incoming and retained outgoing data. Partition counts are balanced, not bytes or hot-key traffic.

Drain or recover an interruption

Use resize-plan --drain NODE_ID --output drain-map.json with the same authorities, then resize to that exact map. Keep the departing data process running until completion. A co-located voter remains part of the fixed authority quorum and must stay running after the data node drains.

A timeout is not rollback. Inspect resize-status, repair unreachable participants, and resume the same target map. There is no cancel operation. Never delete pending metadata, start a conflicting transition, or reuse a retired identity with stale files. Automatic promotion is suspended while a resize is pending.

Draining requires reachable participants; recovering a failed node is a different failover/replacement workflow. Seek a recovery plan before removing directories or quorum metadata. Retained outgoing partitions require explicit quorum-authorized pruning after validation; resizing does not delete cloud servers or erase retired disks.