Skip to content

Kubernetes Platform — Production Installation Guide

Reusable engineering playbook for a self-managed, on-prem, HA Kubernetes platform: 3 control-plane

  • 6–9 workers, Cilium CNI, Longhorn storage, LoadBalancer VIPs, cert-manager, ArgoCD GitOps, Vault/OpenBao secrets, Kyverno policy + tenant isolation, VictoriaMetrics observability, and Keycloak SSO for the cluster and every management UI. Structured as a minimal BASE cluster, then a CHECKLIST of platform layers, then operations & DR. Grounded in the team’s real clusters (docs/K8S-U2/, docs/CG-Stage/) and hardened against the incidents recorded there.

Companion (the “why”, for bids): [[Kubernetes-Platform-Tender-Response]]. Sibling playbooks: [[PostgreSQL-HA-Production-Install-Guide]] · [[Database-Tender-NFR-Responses]]. Last revised: 2026-07-25.

⚠️ Version currency. Versions below reflect the 2026-07 research pass (Kubernetes v1.36.2, Cilium 1.19.6, Longhorn 1.12.0, ArgoCD 3.4.5, Kyverno 1.18.2, VictoriaMetrics 1.148.0, Keycloak 26.7.0, CloudNativePG 1.30.x). Re-validate every version’s support window / EOL at install time — several will have moved. The team’s clusters currently run k8s v1.30, Calico 3.31.2, Longhorn 1.9.2, ArgoCD 3.2.1, Vault 1.20.4, Keycloak 26.3.3.


Three parts:

  • BASE (§0–3): the minimum for a working, HA, networked cluster. Each step is verified end-to-end with observed evidence before the next — the discipline the whole workspace runs on.
  • CHECKLIST (§4–10): the platform layers. Each is independently installable and has its own post-install checklist.
  • OPERATE (§11–13): day-2, disaster recovery, and the safety rules.

⚠️ Lesson boxes are failure modes we have actually hit on the real clusters — they are why a step exists. Two rules govern everything (from CLAUDE.md):

  1. Monitoring + alerting live from day 0. A client database sat dead for 12 days unnoticed because there was no metrics/alerting (CG-006). Build §9 early, not last.
  2. No bulk cluster-wide mutations. Storage/reclaim changes go one volume at a time; order of operations is backups → Retain → then GitOps; snapshot before any state-changing op.

De-risking note (from the critique). The team wanted Cilium but could not configure it. This guide gets the one setting they were missing right (§2), and deliberately sequences the risky pieces: Cilium L2 announcements first, BGP later; diagnose the pre-existing Longhorn/iSCSI errors before expanding Longhorn (§4); introduce kube-proxy replacement and LB advertisement as separate, verified steps, not simultaneously.


0. Prerequisites and node preparation (all nodes)

Section titled “0. Prerequisites and node preparation (all nodes)”
Terminal window
# Topology: 3 control-plane + 6-9 workers on one L2/L3 fabric, static IPs, DNS, NTP.
# Reserve TWO VIPs in DNS now: the API VIP and the ingress VIP.
# 0.1 kernel + sysctl (Cilium needs kernel >= 5.10; cgroup v2)
uname -r # must be >= 5.10 on EVERY node
cat >/etc/modules-load.d/k8s.conf <<'EOF'
overlay
br_netfilter
EOF
modprobe overlay && modprobe br_netfilter
cat >/etc/sysctl.d/k8s.conf <<'EOF'
net.ipv4.ip_forward = 1
net.bridge.bridge-nf-call-iptables = 1
net.bridge.bridge-nf-call-ip6tables = 1
EOF
sysctl --system
swapoff -a && sed -i '/ swap / s/^/#/' /etc/fstab # swap OFF
# 0.2 time sync — TLS, OIDC and etcd all fail on clock skew (critique). Monitor the offset later.
apt-get install -y chrony && systemctl enable --now chrony && chronyc tracking
# 0.3 containerd 2.x + runc, systemd cgroup driver on cgroup v2
# (containerd 2.x plugin path is io.containerd.cri.v1.runtime)
apt-get install -y containerd runc
containerd config default >/etc/containerd/config.toml
sed -i 's/SystemdCgroup = false/SystemdCgroup = true/' /etc/containerd/config.toml
systemctl restart containerd && crictl info | grep -i systemdCgroup
# 0.4 DEDICATED Longhorn data disk — never the OS/containerd root (see §4 Lesson)
# mkfs.ext4 /dev/sdb ; mount at /var/lib/longhorn ; add to /etc/fstab
apt-get install -y open-iscsi nfs-common && systemctl enable --now iscsid
# multipath must NOT grab longhorn devices:
grep -q longhorn /etc/multipath.conf || cat >>/etc/multipath.conf <<'EOF'
blacklist { devnode "^sd[a-z0-9]+" } # or specifically exclude /dev/longhorn/*
EOF

Checklist §0

  • Kernel ≥ 5.10 and cgroup v2 on every node · swap off · sysctls persisted
  • containerd 2.x with SystemdCgroup=true, verified via crictl info
  • A dedicated disk mounted for Longhorn (not the OS/containerd root); iscsid running
  • chrony synced; API VIP and ingress VIP reserved in DNS before any init

Installer decision (record it): RKE2 (upstream v1.36, embedded etcd, static-pod control plane, built-in scheduled etcd snapshots, CIS-hardened by default) is the pragmatic default for lowest day-2 toil; Talos where an immutable, SSH-less appliance OS is acceptable; kubeadm where you need an existing golden image or maximal control — this is what both team clusters already use, so the kubeadm path is shown here.

Terminal window
# 1.1 API VIP FIRST — set as controlPlaneEndpoint so it is baked into every cert SAN and kubeconfig.
# Retrofitting later means regenerating all certs.
# kube-vip static pod (ARP/leader-election) on each control-plane node, OR HAProxy+Keepalived VRRP.
# ⚠️ (critique) kube-vip ARP mode: L2 failover only, leader election needs the API reachable
# (chicken-and-egg) — deploy the static pod BEFORE the CNI, and prefer BGP where the network allows.
export VIP=10.x.x.10 KVVERSION=v1.2.x
ctr image pull ghcr.io/kube-vip/kube-vip:$KVVERSION
# generate /etc/kubernetes/manifests/kube-vip.yaml (arp, controlplane, services=false), probing /readyz
# 1.2 init first control-plane
kubeadm init --control-plane-endpoint "$VIP:6443" \
--upload-certs --skip-phases=addon/kube-proxy \ # ← no kube-proxy: Cilium replaces it (§2)
--pod-network-cidr=10.244.0.0/16
# add the VIP + any DNS names to apiServer.certSANs in the kubeadm config
# 1.3 join the other 2 control-plane nodes (--control-plane --certificate-key ...),
# then join the 6-9 workers (kubeadm join $VIP:6443 --token ...)
# 1.4 etcd health + HARDENING
etcdctl endpoint health --cluster && etcdctl endpoint status --cluster -w table # leader, sane DB size
# add to the etcd static-pod manifest / kubeadm extraArgs:
# --auto-compaction-mode=periodic --auto-compaction-retention=1h
# --quota-backend-bytes=8589934592 (8 GiB; default 2 GiB is a cliff — see the CG-001 Lesson)
# schedule off-cluster snapshots + rehearse a restore BEFORE production cutover
# 1.5 ⚠️ (critique) SECURITY FUNDAMENTALS the base MUST include
# a) API-server audit logging → ship to VictoriaLogs (who-did-what; WB table stakes)
# --audit-policy-file=/etc/kubernetes/audit-policy.yaml --audit-log-path=/var/log/kube-audit.log
# b) etcd secret encryption-at-rest (secrets are plaintext in etcd by default)
# --encryption-provider-config=/etc/kubernetes/enc.yaml (aescbc/secretbox; KMS if an HSM exists)
# c) back up /etc/kubernetes/pki SEPARATELY — etcd snapshots do NOT contain the CA. Total control-plane
# loss needs the CA to rebuild and to keep the break-glass cert valid.
# 1.6 verify HA: drain+reboot ONE control-plane node → VIP fails over, quorum holds, API stays up.

⚠️ Lesson (CG-001). An etcd with no auto-compaction filled to its 2 GiB quota (4.97M revisions for 4 keys), went read-only (NOSPACE), and froze the cluster for 12 days. On NOSPACE you must etcdctl alarm disarm after compact+defrag. Set the hardening in 1.4 and alert on DB-size. Put etcd on low-latency NVMe — CG-011’s “apply request took too long” is a slow-disk symptom.

⚠️ Lesson (kubeadm cert time-bomb). kubeadm control-plane certs expire in 1 year — the classic outage is all certs expiring at once. Automate kubeadm certs renew all (or renew on every upgrade), monitor with kubeadm certs check-expiration, alert weeks ahead. RKE2/Talos rotate automatically.

Checklist §1 — 3 control-plane (odd) · API VIP in cert SANs before init · etcd on dedicated NVMe · etcd snapshots off-cluster and a restore tested · auto-compaction + quota + alerts · pki/ backed up separately · apiserver audit logging on · etcd secret encryption on · control-plane taint retained · cert rotation automated + monitored · on a supported minor, upgrade one at a time CP-first.


2. Cilium CNI — make it actually come up

Section titled “2. Cilium CNI — make it actually come up”

This is the step the team could not get working. The #1 missed setting: kubeProxyReplacement=true without k8sServiceHost/k8sServicePort pointing at the API VIP. Fix that and it works.

Terminal window
# Prereqs: §1 used --skip-phases=addon/kube-proxy (nothing to tear down). Kernel >= 5.10 everywhere.
helm repo add cilium https://helm.cilium.io/ && helm repo update
helm install cilium cilium/cilium --version 1.19.6 -n kube-system \
--set kubeProxyReplacement=true \
--set k8sServiceHost=10.x.x.10 --set k8sServicePort=6443 \ # ← THE API VIP, not a node IP
--set routingMode=native --set autoDirectNodeRoutes=true \
--set ipv4NativeRoutingCIDR=10.244.0.0/16 \
--set bpf.masquerade=true \
--set ipam.mode=cluster-pool \
--set hubble.enabled=true --set hubble.relay.enabled=true --set hubble.ui.enabled=true \
--set hubble.metrics.enableOpenMetrics=true
cilium status --wait # all green
cilium connectivity test # must pass before declaring done
kubectl -n kube-system get pods -l k8s-app=kube-dns # CoreDNS Running

⚠️ De-risk (critique). Do not enable Cilium BGP and kube-proxy-replacement simultaneously on the first pass — that combination is exactly what trips teams. Bring up native routing + kube-proxy-replacement here, validate, then add LB advertisement in §3 as a separate step, L2 announcements first, BGP only once the upstream peering is agreed with the client’s network.

⚠️ Source-IP preservation (critique — collides with your eIntegritate real-client-IP need). With kubeProxyReplacement + externalTrafficPolicy: Local, decide DSR vs SNAT explicitly so ingress sees the true client IP (loadBalancer.mode=dsr or hybrid). This is the same class as the eintegritate-real-client-ip requirement — get it right here.

Calico → Cilium migration (the team runs Calico 3.31.2): use the per-node CiliumNodeConfig live migration — a distinct pod CIDR and tunnel port from Calico, policyEnforcementMode=never until all nodes are Cilium-managed, workers first. Never chain two CNIs permanently.

Checklist §2 — kubeProxyReplacement=true with k8sServiceHost/Port = API VIP · bpf.masquerade=true + ipv4NativeRoutingCIDR set together · kube-proxy DaemonSet gone + stale iptables flushed · Cilium pod CIDR ≠ Calico CIDR during migration · DSR mode chosen for client-IP · cilium connectivity test passes · Hubble up with flow metrics · hostNetwork pods are not covered by NetworkPolicy — enable Cilium host-firewall separately if needed.


Terminal window
# 3.1 Service LoadBalancer — ONE announcer per pool. MetalLB is NOT installed alongside Cilium LB.
# Start with L2 announcements (simplest, no upstream router config):
kubectl apply -f - <<'EOF'
apiVersion: cilium.io/v2alpha1
kind: CiliumLoadBalancerIPPool
metadata: { name: pool-ingress }
spec:
blocks: [ { start: "10.x.x.200", stop: "10.x.x.220" } ]
# reserve the ingress VIP; set allowFirstLastIPs / serviceSelector as needed
---
apiVersion: cilium.io/v2alpha1
kind: CiliumL2AnnouncementPolicy
metadata: { name: l2 }
spec: { interfaces: [ "^eth0" ], externalIPs: true, loadBalancerIPs: true }
EOF
# ⚠️ (critique) Cilium BGP (CiliumBGPClusterConfig, ECMP+BFD, sub-second failover) is the HARDENING
# step — it needs the client's switches to peer (ASN/session). Do it after L2 works, with the
# network team. If Cilium BGP isn't viable, MetalLB v0.16 with frr-k8s (BGP) is the fallback —
# but never MetalLB and Cilium advertising the same pool.
# 3.2 ingress — ingress-nginx was ARCHIVED (2026-03-24). Build NEW ingress on Gateway API v1.6.
kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.0/standard-install.yaml
# data plane: Cilium Gateway API (already have Cilium) or Envoy Gateway v1.8.x
# >= 2 replicas, hard pod anti-affinity, PDB minAvailable>=1, externalTrafficPolicy: Local
# (the team's existing ingress-nginx 1.14.0 stays PINNED behind a migration item — no new use)
# 3.3 cert-manager v1.21 — internal CA via Vault PKI (§5) + ACME DNS-01 for public/wildcard
helm install cert-manager jetstack/cert-manager --version v1.21 -n cert-manager \
--create-namespace --set crds.enabled=true
# ClusterIssuers: vault-pki (internal, short-lived leafs) + letsencrypt-dns01 (public/wildcard).
# ⚠️ HTTP-01 only for public single hosts; wildcard/internal MUST use DNS-01.
# Distribute the root CA to client trust stores; verify a test Certificate issues BEFORE wiring ingress.

⚠️ Lesson (CG-008). cert-manager installed with zero ClusterIssuers meant every host served plain HTTP for months while the annotations looked fine. Add an independent Prometheus alert on cert expiry < 14 d and Ready != True — never rely on auto-renewal being silent.

Checklist §3 — one VIP announcer per pool (no MetalLB+Cilium overlap) · reserved ingress VIP, autoAssign:false on reserved pools · new ingress on Gateway API (nginx pinned/legacy only) · data plane ≥2 replicas + anti-affinity + PDB · both ClusterIssuers Ready · CA distributed · cert-expiry alert independent of auto-renew.


Each layer below is verified end-to-end before the next.

⚠️ Lesson (recurring, the biggest storage risk). The team’s memory records Longhorn/iSCSI block-device I/O errors on worker nodes for months, and disk pressure on a shared root volume flips EXT4 read-only and kills the node (CLAUDE.md). Before expanding Longhorn: diagnose and fix the pre-existing iSCSI/kernel errors, and give Longhorn a dedicated data disk (§0.4) — never the OS/containerd root. Otherwise the new platform inherits the same chronic failure.

Terminal window
longhornctl check preflight # dedicated disk registered, iscsid up, multipath ok
helm install longhorn longhorn/longhorn --version 1.12.0 -n longhorn-system --create-namespace \
--set defaultSettings.replicaSoftAntiAffinity=false \
--set defaultSettings.allowVolumeCreationWithDegradedAvailability=false \
--set defaultSettings.replicaAutoBalance=least-effort \
--set defaultSettings.defaultDataPath=/var/lib/longhorn
# Tiered StorageClasses: reclaimPolicy: Retain, allowVolumeExpansion: true,
# numberOfReplicas 3 (prod-like) / 2 (stage), dataLocality, diskSelector/nodeSelector tags.
# Mark exactly ONE default class.
# BackupTarget → S3/NFS (+ credentials Secret); verify a MANUAL backup reaches Completed.
# RecurringJob: daily backup (retain 7) + snapshot. DR volumes in a second cluster for critical data.

⚠️ Only the V1 (iSCSI) data engine — do not enable V2 without NVMe + hugepages + kernel 6.7+. RWX PVCs route through a single share-manager pod per volume (a per-volume SPOF) — state the limit if any app needs ReadWriteMany. Under app-replicated databases (Postgres/RabbitMQ already replicate), drop Longhorn replicas to 1–2 to avoid ~9× write amplification on flaky iSCSI.

Checklist §4 — dedicated data disk · pre-existing iSCSI errors diagnosed · iscsid up + multipath excludes longhorn · numberOfReplicas set · reclaimPolicy: Retain on all data classes · exactly one default class · BackupTarget verified with a real Completed backup · expansion tested · only V1 engine · over-provisioning/min-available thresholds alerted.

5. Secrets — Vault (evaluate OpenBao) + VSO

Section titled “5. Secrets — Vault (evaluate OpenBao) + VSO”

⚠️ License (critique — World Bank open-standards requirement). HashiCorp Vault is BSL (not OSI-approved) since 2023 — at odds with an anti-lock-in bid. Evaluate OpenBao (the OSI-licensed fork; drop-in) for anything tender-facing. The team already runs Vault 1.20.4; note the swap.

Terminal window
helm install vault hashicorp/vault -n adm --create-namespace \
--set server.ha.enabled=true --set server.ha.raft.enabled=true \ # Integrated Storage (Raft)
--set server.ha.replicas=3 # 3 stage / 5 prod, anti-affined
# TLS on 8200/8201; enable auto-unseal, Kubernetes auth, DB + PKI engines, >=2 audit devices, OIDC.

⚠️ On-prem auto-unseal is the hard problem (critique). “Transit auto-unseal” needs another Vault (turtles); there is no cloud KMS on-prem. The real answers are a physical HSM via PKCS#11 (procurement/cost) or a documented manual unseal with a key-holder quorum (Shamir keys split offline). Decide and document — a Vault that boots sealed on every restart is a hard SPOF, and cert-manager (Vault PKI) and VSO have hard runtime dependencies on it: a seal stops cert renewal and secret sync platform-wide.

⚠️ Audit fail-closed footgun. Vault blocks all requests if an audit write fails (e.g. disk full). Two local file devices don’t help — require ≥1 remote/non-blocking sink + disk-usage alerting.

Consume secrets via VSO 1.5.0 (VaultStaticSecret/VaultDynamicSecret/VaultPKISecret with rolloutRestartTargets). Use dynamic DB credentials and the PKI engine as the internal CA feeding cert-manager (§3.3). OIDC (Keycloak) for operator access — but keep the offline recovery-key break-glass outside OIDC.

Checklist §5 — odd node count, anti-affined · auto-unseal solved concretely (HSM or documented quorum) · Raft snapshots scheduled + restore tested · root token revoked · ≥2 audit devices with a redundant remote sink + disk alert · every mount has finite TTLs · PKI intermediate issuing short leafs · OpenBao evaluated for tender use.

⚠️ Lesson (CG-002). ArgoCD with expired git credentials silently stopped reconciling all 8 apps (ComparisonError), drift unknown, for an unknown time. Alert on sync_status != Synced, health != Healthy, and ComparisonError. And because Git becomes production-critical, the Git repo needs HA/backup/mirror — a hard dependency the platform now has.

Terminal window
kubectl create ns argocd
kubectl apply -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.4.5/manifests/ha/install.yaml
# redis-ha 3+3; server/repo-server >= 2; application-controller StatefulSet replicas == ARGOCD_CONTROLLER_REPLICAS
# OIDC → Keycloak (groups claim) in argocd-cm; argocd-rbac-cm policy.default = role:readonly + group→role
# One AppProject per tenant (sourceRepos/destinations/resource allow-lists). Default project unused for real workloads.
# Secrets resolved on the destination via VSO — NEVER plaintext in Git (gitleaks clean).
# Install Argo Rollouts 1.8.x; convert prod Deployments to Rollout with VictoriaMetrics AnalysisTemplate.

Checklist §6 — HA, pinned v3.4.5, ≥3 nodes, redis-ha healthy · controller sharding correct · OIDC login + groups work · AppProject per tenant, allow-lists · no plaintext Secret in Git · repo creds centralised · alerts on sync/health/ComparisonError · allowEmpty:false; prune/selfHeal off for stateful/storage apps · Git repo backed up/mirrored.

Terminal window
helm install kyverno kyverno/kyverno --version 1.18.2 -n kyverno --create-namespace \
--set admissionController.replicas=3 # HA; separate reports/background/cleanup controllers
# 1) Pod Security Admission: label every tenant ns pod-security.kubernetes.io/{enforce,audit,warn}=restricted (version-pinned)
# 2) Kyverno GENERATE on namespace-create: default-deny NetworkPolicy + ResourceQuota + LimitRange + RoleBinding
# 3) Kyverno VALIDATE (Audit→Enforce): require limits/requests, disallow privileged + :latest, restrict registries, require labels
# 4) Kyverno verifyImages: cosign keyless (or key-pair air-gapped) + required SLSA provenance

⚠️ Critique fixes — do not skip these:

  • PSA is bypassable. Anyone who can edit a namespace can downgrade the enforce label. Add a Kyverno policy that forbids mutating the pod-security.kubernetes.io/* labels — otherwise the “non-bypassable floor” isn’t.
  • Default-deny breaks DNS. A generated deny-all with no explicit egress to CoreDNS (UDP/TCP 53) breaks every pod. Put the CoreDNS egress allow in the generate bundle.
  • Namespace ≠ security boundary. For hostile tenants, a shared kernel/node means a container escape (privileged, hostPath, hostPID/IPC/Network, nodes/proxy) reaches other tenants. Document the threat model; use vCluster / per-tenant node pools for real isolation.
  • Webhook lockout. failurePolicy: Fail means a Kyverno outage blocks all admission — exclude kube-system and kyverno, spread replicas, set timeoutSeconds. Pick one engine (Kyverno) — don’t also run Gatekeeper.

⚠️ Keyless cosign is impossible air-gapped (needs public Fulcio/Rekor). An on-prem gov env needs self-hosted Sigstore or key-pair signing. Choose per environment.

Checklist §7 — Kyverno HA · PSA restricted on every tenant ns · policy forbidding PSA-label downgrade · validate policies Audit→Enforce · generate bundle active with CoreDNS egress · default-deny in 100% of tenant ns + cluster-wide Cilium backstop · verifyImages enforced (signing model chosen for air-gap) · failurePolicy scoped away from control plane · Policy Reporter → monitoring.

⚠️ Gap the base story omitted (critique): a private container registry (Harbor or equivalent) is where signed images, SBOMs and cosign attestations actually live, and it must be mirrored for air-gap. Add Harbor as a platform component; it is a prerequisite for the supply-chain claims in §7 and the tender.

9. Observability — VictoriaMetrics (build this EARLY)

Section titled “9. Observability — VictoriaMetrics (build this EARLY)”
Terminal window
helm install vmks vm/victoria-metrics-k8s-stack --version 1.148.0 -n monitoring --create-namespace
# operator + vmagent + vmalert + Alertmanager + Grafana + node-exporter + kube-state-metrics.

⚠️ Critique — do not ship VMSingle for the platform monitor. A single replica / single PVC for the very system whose job is “never be blind again” is self-defeating. Use VMCluster (RF≥2) for the platform monitor, or explicitly own the risk. Set retention + durable PVC; vmagent on-disk buffering.

Scrape: control plane (etcd fsync/commit p99, apiserver SLO, scheduler, controller-manager), Longhorn, cert-manager, ArgoCD, Vault-sealed, Cilium/Hubble, Keycloak. Alert packs: node/pod, etcd latency + DB size, apiserver SLO, Longhorn volume health, cert expiry 21 d/7 d, ArgoCD OutOfSync, Vault sealed, clock skew. Watchdog dead-man’s-switch → an external heartbeat.

⚠️ Critique — the dead-man’s-switch and notifications must be real on-prem. SaaS heartbeats (Dead Man’s Snitch/healthchecks.io) assume internet egress; a restricted gov network needs an on-prem heartbeat receiver and a named notification transport (SMTP relay? webhook?). “One tested channel” is exactly what was missing in the 12-day outage — name the actual transport. VictoriaLogs receives logs; OTel Collector stops exporting to debug. 1-year audit retention needs downsampling (VM Enterprise) or unbounded growth — note the license/cost.

Checklist §9 — VMCluster RF≥2 (not VMSingle) · retention + durable storage · control-plane scrape live · Watchdog firing on an on-prem heartbeat · Alertmanager HA with a tested on-prem channel · alert packs loaded · VictoriaLogs receiving · OTel no longer debug-only · cardinality guardrails · all GitOps-managed.

10. SSO — Keycloak for the cluster AND every UI

Section titled “10. SSO — Keycloak for the cluster AND every UI”
Terminal window
# Keycloak Operator + EXTERNAL clustered PostgreSQL (CloudNativePG) — never start-dev/H2.
# Keycloak CR: instances >= 2, TLS, strict hostname, proxy headers.
# One reusable 'groups' client scope (Group Membership mapper, Full path OFF) on every client.

kube-apiserver → Keycloak (k8s ≥ 1.30 Structured Authentication Configuration): --authentication-config with issuer = <realm>, username = preferred_username, groups = groups (both oidc: prefixed); RBAC binds kind: Group. kubectl via kubelogin (oidc-login, PKCE).

Per-UI integration — native OIDC where it exists, oauth2-proxy where it doesn’t:

UIMethodGotcha to get right
ArgoCDnative OIDCneeds a Keycloak group mapper or the groups claim is empty
Grafananative OAuthset role_attribute_strict=true, default Viewer — else everyone becomes Admin
Vault/OpenBaonative OIDCexternal groups → policies; keep recovery-key break-glass outside OIDC
MinIO Consolenative OIDCpolicy claim required — no claim = zero access (looks like a broken login)
RabbitMQ mgmtnative oauth2aud must equal resource_server_id, not the client id, or login fails
pgAdminnative OAuth2OIDC logs into pgAdmin only — DB creds + MASTER_PASSWORD remain; no in-app roles
Longhorn UIoauth2-proxyno native auth — a bare Ingress is a full storage console; gate on a group. All-or-nothing (no read-only)
Hubble UIoauth2-proxyno native auth; must not be exposed unprotected
K8s Dashboardbearer token / oauth2-proxyreconsider deploying it at all; HA oauth2-proxy mandatory

⚠️ Critique — the two biggest traps:

  1. Break-glass + circular dependency. Cluster auth depends on Keycloak, which runs inside the cluster on CloudNativePG. If storage/DB/Keycloak is down, nobody can kubectl — except via a sealed static cluster-admin client-cert kubeconfig kept offline. Keep exactly one, monitor its expiry, document the retrieval procedure. Each UI also needs a documented sealed local-admin fallback (ArgoCD admin, Grafana admin, Vault recovery keys, pgAdmin internal). Humans lose access within one 5–15 min token lifetime of a Keycloak outage; bound SA tokens keep GitOps/controllers alive.
  2. iss split-horizon. The apiserver’s --authentication-config issuer must byte-match the token iss; an internal-vs-external hostname mismatch silently fails all OIDC. And the apiserver must reach Keycloak’s JWKS through Cilium’s service LB — if k8sServiceHost is wrong or DNS is split-horizon, service LB and OIDC break together. oauth2-proxy is itself a SPOF — ≥2 replicas + shared session store (Redis) + cookie secret; watch for large-group header overflow (many groups → 502 on login).

Checklist §10 — Keycloak from Operator, ≥2 replicas, external clustered PG · TLS + strict hostname · one groups scope on every client · apiserver --authentication-config, RBAC kind:Group · kubectl via kubelogin · every UI behind Keycloak (native or oauth2-proxy) · sealed break-glass cert + per-UI local-admin fallbacks documented · iss byte-matches · oauth2-proxy HA · MFA on *-admin · Keycloak events → logging · Keycloak realm-as-code + its Postgres in the DR plan.


11. Data services (in-cluster, stage only)

Section titled “11. Data services (in-cluster, stage only)”

Run only disposable/stage state in-cluster; keep the production system-of-record DB external (the client’s Patroni+HAProxy+etcd+Keepalived — see [[PostgreSQL-HA-Production-Install-Guide]]), referenced via ExternalName/Endpoints so manifests stay identical across environments.

ServiceOperatorNotes
PostgreSQLCloudNativePG 1.30.x (or existing CrunchyData PGO)1 primary + 2 replicas; Barman/pgBackRest PITR to off-cluster S3; Longhorn replicas 1–2
Redis/Valkeyredis-operator 0.26.x, Sentinel (1+2+3)reconstructable cache; Valkey avoids Redis 8 AGPLv3
RabbitMQCluster Operator 2.22, RabbitMQ 4.xquorum queues (classic mirrored removed in 4.0); definitions exported to git
Object storageMinIO EOL risk (repo archived, last v7.1.1)pin-frozen or adopt SeaweedFS / Garage / Rook-Ceph RGW

Checklist §11 — each service via its operator CR, version-pinned in git · Retain-reclaim, expandable StorageClass · Longhorn replicas reduced under replicated services · operator-native off-cluster backup + tested restore · PDBs allow one-at-a-time drain (avoid the 0-disruption trap seen on cg-stage) · exporters/ServiceMonitors wired · prod DB external · MinIO path chosen.

12. Platform disaster recovery (the hard part — critique)

Section titled “12. Platform disaster recovery (the hard part — critique)”

Backing up etcd, Vault, Longhorn, Keycloak and PKI in isolation is not a DR plan. Total-cluster recovery has cross-dependencies and a required order.

  • Consolidated DR runbook with restore ordering. Rough order: restore CA/pki → etcd snapshot (--force-new-cluster, re-add members) → bring up control plane + Cilium → restore Longhorn volumes (etcd references PVs that need Longhorn back) → unseal Vault (secrets that manifests reference) → Keycloak realm + its Postgres → verify. Rehearse it (game-days), not once.
  • Consistent backup point. etcd@T1 + Longhorn@T2 + Vault@T3 restore to an internally inconsistent cluster (PVCs with no backing volume, secrets pointing at rolled Vault paths). Coordinate a point-in-time or accept and document the reconciliation steps.
  • Add Velero for k8s-native, granular namespace/PVC (CSI-snapshot) backup — etcd snapshot is all-or-nothing.
  • Quorum-loss runbook (lose 2 of 3 etcd): etcdctl snapshot restore single-member bootstrap → re-add. Written and rehearsed.
  • Immutability/offsite/encryption (S3 Object Lock vs ransomware) on every backup; site-level DR posture stated (single-DC on-prem = correlated failures — a WB bid usually needs a multi-site story).
  • RPO/RTO per tier stated (etcd snapshot interval = data-loss window; Longhorn large-volume restore = full download; Vault/Keycloak restore windows).

13. Operating rules (from CLAUDE.md — non-negotiable)

Section titled “13. Operating rules (from CLAUDE.md — non-negotiable)”
  • No bulk cluster-wide mutations. PV/reclaim changes one volume at a time; snapshot to backup/ first; verify between each. Order: backups → Retain → then GitOps.
  • Never trust an L4/TCP health check as a real signal — probe the actual protocol/SLI (the CG-012 false-UP trap that hid a 12-day outage).
  • Monitoring + alerting from day 0. Build §9 before you need it.
  • Log state-changing commands to the cluster’s ops trail; be explicit about which cluster (--context) every command targets.
  • PriorityClasses for platform components (Cilium, monitoring, Vault, ingress) so app load can’t evict them on a fixed 6–9 node cluster.

Idempotent install scripts live beside this guide under docs/Playbooks/scripts/k8s/ (to be generated per environment from the parameterised steps above): 00-prep-node.sh, 10-bootstrap-cp.sh, 20-cilium.sh, 30-lb-ingress-tls.sh, 40-longhorn.sh, 50-vault.sh, 60-argocd.sh, 70-kyverno.sh, 80-victoriametrics.sh, 90-keycloak-sso.sh, dr/etcd-snapshot.sh, dr/etcd-restore.sh. Each is a thin, commented wrapper around the commands above with the environment values (VIPs, CIDRs, versions, domains) pulled from a single env.sh. Keep versions in env.sh and re-validate them at install time.

Appendix B — the team’s real stack vs this guide

Section titled “Appendix B — the team’s real stack vs this guide”
LayerAs-built (k8s-u2 / cg-stage)This guide
Control planeEKS (u2) / single kubeadm master (cg-stage, a SPOF)3 CP + API VIP, self-managed
Versionv1.30 / v1.29.14v1.36.x (re-validate)
CNICalico 3.31.2 / 3.25Cilium (native routing, kube-proxy-free) + migration path
LBMetalLBCilium LB-IPAM (L2 → BGP); MetalLB fallback
Ingressingress-nginx 1.14/1.12Gateway API (nginx pinned/legacy)
StorageLonghorn 1.9.2 / local-pathLonghorn 1.12 on dedicated disk + Retain + backups
GitOpsArgoCD 3.2.1 / 3.0.6ArgoCD 3.4.5 HA + OIDC + alerts
SecretsVault 1.20.4+VSO / noneVault/OpenBao HA + solved auto-unseal
PolicynoneKyverno + PSA + tenant isolation
Observabilitypartial / none (CG-006)VictoriaMetrics from day 0
SSOKeycloak 26.3.3 present / noneKeycloak on the API + every UI + break-glass

Consolidated from the 2026-07 multi-agent research pass (14 dimensions, authoritative primaries: kubernetes.io, docs.cilium.io, longhorn.io, cert-manager.io, argo-cd.readthedocs.io, developer.hashicorp.com/openbao.org, kyverno.io, docs.victoriametrics.com, keycloak.org, cloudnative-pg.io, CIS/NSA-CISA/NIST/ISO, worldbank.org procurement). Per-dimension source URLs are in the research transcript under subagents/workflows/…/journal.jsonl.