Skip to content

/pg-diagnose — PostgreSQL HA diagnosis

Target: $ARGUMENTS

Invoke the postgres-ha-ops skill and walk the diagnosis decision tree in order. Read-only.

  1. Protocol probe, not L4. Send an SSLRequest (8 bytes: len 8 + code 80877103) to the target :5432. S/N = alive → problem is app/network side. EOF/close = no healthy backend → continue. Never accept a TCP-port-open check as proof the database works.
  2. HAProxy stats: curl -s -u <user>:<pass> http://<lb>:7000/;csv → which pgsql-* are DOWN and since when (L4CON=nothing listening, L7STS=Patroni says not-leader/replica).
  3. On a DB node (SSH): patronictl -c /etc/patroni/patroni.yml list — leader present? replicas streaming? empty cluster or start failed?
  4. Root cause: journalctl -u patroni -n 50 — mvcc: database space exceeded → Runbook A (etcd NOSPACE); requested timeline N is not a child → Runbook B (divergent replica).
  5. Port fingerprint the DB nodes: 5432 / 6432 / 8008 (RST vs filtered).

Remember: under Patroni, systemctl status postgresql = inactive/disabled is normal, and a refused :8008 from outside does not mean Patroni is dead — check patronictl/journalctl.

Report findings with the observed evidence, then propose the matching runbook — do not execute any state-changing step without an explicit go-ahead.