# Troubleshoot common problems

Begin with the affected operation, exact error, installed build, recent changes, and observation time. Collect metadata and counters without recording tokens or customer values.

## Cannot connect

For the packaged service, inspect `systemctl status smkv` and `journalctl -u smkv`. Confirm the bind address, firewall, and port. A loopback bind is reachable only on that host. Management discovery uses 7381 in these guides; client data uses 7379. A standalone node does not provide cluster management discovery.

If the seed works but routed requests fail, check every advertised endpoint from the application host. Private addresses must be reachable through the intended network or tunnel. For secured transport, confirm both sides use the same mode and certificates match endpoint SANs. Do not disable certificate verification to hide a mismatch.

## Writes return Busy or Capacity

Busy means admission rejected the request before execution. Back off within a bounded deadline and examine queue pressure. Capacity is a known rejection requiring more headroom or lower retained usage. On the development build:

```sh
smkv-ctl --seed 10.0.0.11:7381 capacity --json
```

Inspect resource sample age, rejection reasons, logical quotas, available disk/inodes, memory, and maintenance reserves. Stale/missing samples reject writes. A normal physical-capacity state does not mean every partition has remaining entry/log quota. Replica pressure can stall catch-up. Do not turn off protection as the first response.

## Reads are stale or missing

Confirm the key/table, TTL, read owner, and cluster identity. `GetReplica` explicitly allows lag. Expired keys are logically absent even before disk reclamation. For writes followed by reads, await the write first; concurrent requests across pooled connections are unordered.

Inspect partition readiness and replication observations. Small cursor differences during traffic are not a consistent snapshot; use a quiescent verification when exact agreement matters. Persistent lag merits resource, transport, and error-log checks.

## A mutation timed out

The server may have executed the mutation before the response was lost. In the current Go SDK, check `errors.Is(err, smkv.ErrOutcomeUnknown)`. Do not blindly replay it. Reconcile application state or use an application-level idempotency design appropriate to the operation.

Increasing a timeout does not repair unavailable quorum or insufficient capacity. For confirmed writes, ensure the client deadline exceeds the server confirmation budget plus network time.

## Quorum cluster unavailable

Check all three authority processes, their connectivity, and durable state. Losing quorum prevents successful client operations; adding a random voter or deleting its state is not a supported repair. Inspect pending resize operations before attempting another ownership change.

Keep original node IDs and complete directories during ordinary restarts. Replacement nodes require explicit enrollment under a new identity. Never copy only an incarnation file onto an empty store.

## Dashboard shows unavailable metrics

Configure each node's metrics endpoint using `--metrics NODE_ID=IP:PORT` and ensure the backend can reach it. Browser refresh does not force backend collection. Check sample ages; a blank or interrupted graph is not zero throughput. The dashboard `/healthz` reports web-process liveness, not database readiness.

## A scan stops or returns no rows

An empty page can carry a continuation cursor. Continue with that cursor and the same query. Busy can be retried after a short delay; restart-required errors mean the cursor is invalid. Foreground traffic has priority and can starve scans. Avoid unbounded retries and use backup exports for recovery archives.

## Share a useful support report

Include package/build versions, redacted configuration, cluster identity and epoch, affected partitions, recent maintenance, error counts, and relevant log timestamps. Add `smkv-ctl health --json` and capacity output where supported. Send initial inquiries to [info@simplemagic.com](mailto:info@simplemagic.com); agree on a secure channel before transferring sensitive diagnostics.
