# CLI, dashboard, and metrics

**Applies to:** baseline cluster observation plus newer development diagnostics. Capacity, scans, compatibility, and sync-checkpoint views need matching newer tools and servers.

## Inspect from the command line

Install tools on an operator host that can reach every advertised management address. Standalone nodes do not expose the sharded-cluster management API.

```sh
smkv-ctl --seed 10.0.0.11:7381 info
smkv-ctl --seed 10.0.0.11:7381 partitions
smkv-ctl --seed 10.0.0.11:7381 health --json
smkv-ctl --seed 10.0.0.11:7381 health --watch 5
```

Inspection is read-only. Degraded health returns exit code 2; missing observations are reported as unavailable. Samples across nodes are not atomic: small transient replica lag during writes is not by itself evidence of failure. Investigate sustained lag, missing owners, fenced writes, and capacity rejections.

Primary entries and replica copies are counted separately. Indexed entries may include expired records pending cleanup. Retained log size includes old versions and tombstones. Index memory is an estimate, not measured process RSS.

## Start the dashboard

Run the backend on a management host with cluster connectivity:

```sh
smkv-web --seed 10.0.0.11:7381 \
  --metrics n1=10.0.0.11:9108 \
  --metrics n2=10.0.0.12:9108 \
  --metrics n3=10.0.0.13:9108
```

It listens on `127.0.0.1:8080` by default. For a remote management host:

```sh
ssh -N -L 8080:127.0.0.1:8080 operator@MANAGEMENT_HOST
```

Open `http://localhost:8080`. A non-loopback bind requires `SMKV_WEB_TOKEN` with at least 32 printable ASCII characters, plus an SSH tunnel or trusted HTTPS proxy. Database transport credentials are configured separately with `--security-config`.

Overview shows live read/write QPS; Activity shows request metrics. Choose a browser refresh interval using the refresh selector. Browser refresh reads cached observations, not a new cluster-wide query. Collection intervals remain independently configured. Graph gaps represent unavailable data; refreshing or reopening the page does not provide durable historical storage. Actionable health badges open further findings.

## Expose Prometheus metrics

Start each server with a private `--metrics-bind IP:9108`. Example Prometheus job:

```yaml
scrape_configs:
  - job_name: smkv
    scrape_interval: 10s
    static_configs:
      - targets: ['10.0.0.11:9108', '10.0.0.12:9108', '10.0.0.13:9108']
```

The HTTP `/metrics` listener reports request outcomes and latency, worker queues, storage, replication, and readiness. It does not inherit database transport TLS/authentication. Limit access to your monitoring network.

Use counter rates for throughput and error rates. Server processing latency is not the same as end-to-end client latency, which also includes network, queuing, and client behavior. Use Prometheus or another monitoring system for persistent history and alerting.

## Daily checks

Check capacity headroom, readiness, sustained replica lag, request errors, and p99 latency against your own service objectives. After a restart or topology change, verify all partition owners and both copies before another maintenance step. For a capacity rejection, use `smkv-ctl capacity` and the dashboard Nodes details. Experimental sync checkpoint status appears under `smkv-ctl checkpoints` and Activity.
