Kubernetes Platform — Production Installation Guide
Acest conținut nu este încă disponibil în limba selectată.
Reusable engineering playbook for a self-managed, on-prem, HA Kubernetes platform: 3 control-plane
- 6–9 workers, Cilium CNI, Longhorn storage, LoadBalancer VIPs, cert-manager, ArgoCD GitOps, Vault/OpenBao secrets, Kyverno policy + tenant isolation, VictoriaMetrics observability, and Keycloak SSO for the cluster and every management UI. Structured as a minimal BASE cluster, then a CHECKLIST of platform layers, then operations & DR. Grounded in the team’s real clusters (
docs/K8S-U2/,docs/CG-Stage/) and hardened against the incidents recorded there.
Companion (the “why”, for bids): [[Kubernetes-Platform-Tender-Response]]. Sibling playbooks: [[PostgreSQL-HA-Production-Install-Guide]] · [[Database-Tender-NFR-Responses]]. Last revised: 2026-07-25.
⚠️ Version currency. Versions below reflect the 2026-07 research pass (Kubernetes v1.36.2, Cilium 1.19.6, Longhorn 1.12.0, ArgoCD 3.4.5, Kyverno 1.18.2, VictoriaMetrics 1.148.0, Keycloak 26.7.0, CloudNativePG 1.30.x). Re-validate every version’s support window / EOL at install time — several will have moved. The team’s clusters currently run k8s v1.30, Calico 3.31.2, Longhorn 1.9.2, ArgoCD 3.2.1, Vault 1.20.4, Keycloak 26.3.3.
How to read this guide
Section titled “How to read this guide”Three parts:
- BASE (§0–3): the minimum for a working, HA, networked cluster. Each step is verified end-to-end with observed evidence before the next — the discipline the whole workspace runs on.
- CHECKLIST (§4–10): the platform layers. Each is independently installable and has its own post-install checklist.
- OPERATE (§11–13): day-2, disaster recovery, and the safety rules.
⚠️ Lesson boxes are failure modes we have actually hit on the real clusters — they are why a step
exists. Two rules govern everything (from CLAUDE.md):
- Monitoring + alerting live from day 0. A client database sat dead for 12 days unnoticed
because there was no metrics/alerting (
CG-006). Build §9 early, not last. - No bulk cluster-wide mutations. Storage/reclaim changes go one volume at a time; order of
operations is backups →
Retain→ then GitOps; snapshot before any state-changing op.
De-risking note (from the critique). The team wanted Cilium but could not configure it. This guide gets the one setting they were missing right (§2), and deliberately sequences the risky pieces: Cilium L2 announcements first, BGP later; diagnose the pre-existing Longhorn/iSCSI errors before expanding Longhorn (§4); introduce kube-proxy replacement and LB advertisement as separate, verified steps, not simultaneously.
PART A — BASE CLUSTER
Section titled “PART A — BASE CLUSTER”0. Prerequisites and node preparation (all nodes)
Section titled “0. Prerequisites and node preparation (all nodes)”# Topology: 3 control-plane + 6-9 workers on one L2/L3 fabric, static IPs, DNS, NTP.# Reserve TWO VIPs in DNS now: the API VIP and the ingress VIP.
# 0.1 kernel + sysctl (Cilium needs kernel >= 5.10; cgroup v2)uname -r # must be >= 5.10 on EVERY nodecat >/etc/modules-load.d/k8s.conf <<'EOF'overlaybr_netfilterEOFmodprobe overlay && modprobe br_netfiltercat >/etc/sysctl.d/k8s.conf <<'EOF'net.ipv4.ip_forward = 1net.bridge.bridge-nf-call-iptables = 1net.bridge.bridge-nf-call-ip6tables = 1EOFsysctl --systemswapoff -a && sed -i '/ swap / s/^/#/' /etc/fstab # swap OFF
# 0.2 time sync — TLS, OIDC and etcd all fail on clock skew (critique). Monitor the offset later.apt-get install -y chrony && systemctl enable --now chrony && chronyc tracking
# 0.3 containerd 2.x + runc, systemd cgroup driver on cgroup v2# (containerd 2.x plugin path is io.containerd.cri.v1.runtime)apt-get install -y containerd runccontainerd config default >/etc/containerd/config.tomlsed -i 's/SystemdCgroup = false/SystemdCgroup = true/' /etc/containerd/config.tomlsystemctl restart containerd && crictl info | grep -i systemdCgroup
# 0.4 DEDICATED Longhorn data disk — never the OS/containerd root (see §4 Lesson)# mkfs.ext4 /dev/sdb ; mount at /var/lib/longhorn ; add to /etc/fstabapt-get install -y open-iscsi nfs-common && systemctl enable --now iscsid# multipath must NOT grab longhorn devices:grep -q longhorn /etc/multipath.conf || cat >>/etc/multipath.conf <<'EOF'blacklist { devnode "^sd[a-z0-9]+" } # or specifically exclude /dev/longhorn/*EOFChecklist §0
- Kernel ≥ 5.10 and cgroup v2 on every node · swap off · sysctls persisted
- containerd 2.x with
SystemdCgroup=true, verified viacrictl info - A dedicated disk mounted for Longhorn (not the OS/containerd root);
iscsidrunning - chrony synced; API VIP and ingress VIP reserved in DNS before any init
1. HA control plane + etcd
Section titled “1. HA control plane + etcd”Installer decision (record it): RKE2 (upstream v1.36, embedded etcd, static-pod control plane, built-in scheduled etcd snapshots, CIS-hardened by default) is the pragmatic default for lowest day-2 toil; Talos where an immutable, SSH-less appliance OS is acceptable; kubeadm where you need an existing golden image or maximal control — this is what both team clusters already use, so the kubeadm path is shown here.
# 1.1 API VIP FIRST — set as controlPlaneEndpoint so it is baked into every cert SAN and kubeconfig.# Retrofitting later means regenerating all certs.# kube-vip static pod (ARP/leader-election) on each control-plane node, OR HAProxy+Keepalived VRRP.# ⚠️ (critique) kube-vip ARP mode: L2 failover only, leader election needs the API reachable# (chicken-and-egg) — deploy the static pod BEFORE the CNI, and prefer BGP where the network allows.export VIP=10.x.x.10 KVVERSION=v1.2.xctr image pull ghcr.io/kube-vip/kube-vip:$KVVERSION# generate /etc/kubernetes/manifests/kube-vip.yaml (arp, controlplane, services=false), probing /readyz
# 1.2 init first control-planekubeadm init --control-plane-endpoint "$VIP:6443" \ --upload-certs --skip-phases=addon/kube-proxy \ # ← no kube-proxy: Cilium replaces it (§2) --pod-network-cidr=10.244.0.0/16# add the VIP + any DNS names to apiServer.certSANs in the kubeadm config
# 1.3 join the other 2 control-plane nodes (--control-plane --certificate-key ...),# then join the 6-9 workers (kubeadm join $VIP:6443 --token ...)
# 1.4 etcd health + HARDENINGetcdctl endpoint health --cluster && etcdctl endpoint status --cluster -w table # leader, sane DB size# add to the etcd static-pod manifest / kubeadm extraArgs:# --auto-compaction-mode=periodic --auto-compaction-retention=1h# --quota-backend-bytes=8589934592 (8 GiB; default 2 GiB is a cliff — see the CG-001 Lesson)# schedule off-cluster snapshots + rehearse a restore BEFORE production cutover
# 1.5 ⚠️ (critique) SECURITY FUNDAMENTALS the base MUST include# a) API-server audit logging → ship to VictoriaLogs (who-did-what; WB table stakes)# --audit-policy-file=/etc/kubernetes/audit-policy.yaml --audit-log-path=/var/log/kube-audit.log# b) etcd secret encryption-at-rest (secrets are plaintext in etcd by default)# --encryption-provider-config=/etc/kubernetes/enc.yaml (aescbc/secretbox; KMS if an HSM exists)# c) back up /etc/kubernetes/pki SEPARATELY — etcd snapshots do NOT contain the CA. Total control-plane# loss needs the CA to rebuild and to keep the break-glass cert valid.
# 1.6 verify HA: drain+reboot ONE control-plane node → VIP fails over, quorum holds, API stays up.⚠️ Lesson (CG-001). An etcd with no auto-compaction filled to its 2 GiB quota (4.97M revisions for 4 keys), went read-only (
NOSPACE), and froze the cluster for 12 days. OnNOSPACEyou mustetcdctl alarm disarmaftercompact+defrag. Set the hardening in 1.4 and alert on DB-size. Put etcd on low-latency NVMe —CG-011’s “apply request took too long” is a slow-disk symptom.
⚠️ Lesson (kubeadm cert time-bomb). kubeadm control-plane certs expire in 1 year — the classic outage is all certs expiring at once. Automate
kubeadm certs renew all(or renew on every upgrade), monitor withkubeadm certs check-expiration, alert weeks ahead. RKE2/Talos rotate automatically.
Checklist §1 — 3 control-plane (odd) · API VIP in cert SANs before init · etcd on dedicated NVMe ·
etcd snapshots off-cluster and a restore tested · auto-compaction + quota + alerts · pki/ backed
up separately · apiserver audit logging on · etcd secret encryption on · control-plane taint
retained · cert rotation automated + monitored · on a supported minor, upgrade one at a time CP-first.
2. Cilium CNI — make it actually come up
Section titled “2. Cilium CNI — make it actually come up”This is the step the team could not get working. The #1 missed setting: kubeProxyReplacement=true
without k8sServiceHost/k8sServicePort pointing at the API VIP. Fix that and it works.
# Prereqs: §1 used --skip-phases=addon/kube-proxy (nothing to tear down). Kernel >= 5.10 everywhere.helm repo add cilium https://helm.cilium.io/ && helm repo updatehelm install cilium cilium/cilium --version 1.19.6 -n kube-system \ --set kubeProxyReplacement=true \ --set k8sServiceHost=10.x.x.10 --set k8sServicePort=6443 \ # ← THE API VIP, not a node IP --set routingMode=native --set autoDirectNodeRoutes=true \ --set ipv4NativeRoutingCIDR=10.244.0.0/16 \ --set bpf.masquerade=true \ --set ipam.mode=cluster-pool \ --set hubble.enabled=true --set hubble.relay.enabled=true --set hubble.ui.enabled=true \ --set hubble.metrics.enableOpenMetrics=true
cilium status --wait # all greencilium connectivity test # must pass before declaring donekubectl -n kube-system get pods -l k8s-app=kube-dns # CoreDNS Running⚠️ De-risk (critique). Do not enable Cilium BGP and kube-proxy-replacement simultaneously on the first pass — that combination is exactly what trips teams. Bring up native routing + kube-proxy-replacement here, validate, then add LB advertisement in §3 as a separate step, L2 announcements first, BGP only once the upstream peering is agreed with the client’s network.
⚠️ Source-IP preservation (critique — collides with your eIntegritate real-client-IP need). With
kubeProxyReplacement+externalTrafficPolicy: Local, decide DSR vs SNAT explicitly so ingress sees the true client IP (loadBalancer.mode=dsrorhybrid). This is the same class as theeintegritate-real-client-iprequirement — get it right here.
Calico → Cilium migration (the team runs Calico 3.31.2): use the per-node CiliumNodeConfig
live migration — a distinct pod CIDR and tunnel port from Calico, policyEnforcementMode=never
until all nodes are Cilium-managed, workers first. Never chain two CNIs permanently.
Checklist §2 — kubeProxyReplacement=true with k8sServiceHost/Port = API VIP ·
bpf.masquerade=true + ipv4NativeRoutingCIDR set together · kube-proxy DaemonSet gone + stale
iptables flushed · Cilium pod CIDR ≠ Calico CIDR during migration · DSR mode chosen for client-IP ·
cilium connectivity test passes · Hubble up with flow metrics · hostNetwork pods are not covered by
NetworkPolicy — enable Cilium host-firewall separately if needed.
3. LoadBalancer VIPs + ingress + TLS
Section titled “3. LoadBalancer VIPs + ingress + TLS”# 3.1 Service LoadBalancer — ONE announcer per pool. MetalLB is NOT installed alongside Cilium LB.# Start with L2 announcements (simplest, no upstream router config):kubectl apply -f - <<'EOF'apiVersion: cilium.io/v2alpha1kind: CiliumLoadBalancerIPPoolmetadata: { name: pool-ingress }spec: blocks: [ { start: "10.x.x.200", stop: "10.x.x.220" } ] # reserve the ingress VIP; set allowFirstLastIPs / serviceSelector as needed---apiVersion: cilium.io/v2alpha1kind: CiliumL2AnnouncementPolicymetadata: { name: l2 }spec: { interfaces: [ "^eth0" ], externalIPs: true, loadBalancerIPs: true }EOF# ⚠️ (critique) Cilium BGP (CiliumBGPClusterConfig, ECMP+BFD, sub-second failover) is the HARDENING# step — it needs the client's switches to peer (ASN/session). Do it after L2 works, with the# network team. If Cilium BGP isn't viable, MetalLB v0.16 with frr-k8s (BGP) is the fallback —# but never MetalLB and Cilium advertising the same pool.
# 3.2 ingress — ingress-nginx was ARCHIVED (2026-03-24). Build NEW ingress on Gateway API v1.6.kubectl apply -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.0/standard-install.yaml# data plane: Cilium Gateway API (already have Cilium) or Envoy Gateway v1.8.x# >= 2 replicas, hard pod anti-affinity, PDB minAvailable>=1, externalTrafficPolicy: Local# (the team's existing ingress-nginx 1.14.0 stays PINNED behind a migration item — no new use)
# 3.3 cert-manager v1.21 — internal CA via Vault PKI (§5) + ACME DNS-01 for public/wildcardhelm install cert-manager jetstack/cert-manager --version v1.21 -n cert-manager \ --create-namespace --set crds.enabled=true# ClusterIssuers: vault-pki (internal, short-lived leafs) + letsencrypt-dns01 (public/wildcard).# ⚠️ HTTP-01 only for public single hosts; wildcard/internal MUST use DNS-01.# Distribute the root CA to client trust stores; verify a test Certificate issues BEFORE wiring ingress.⚠️ Lesson (CG-008). cert-manager installed with zero ClusterIssuers meant every host served plain HTTP for months while the annotations looked fine. Add an independent Prometheus alert on cert expiry < 14 d and
Ready != True— never rely on auto-renewal being silent.
Checklist §3 — one VIP announcer per pool (no MetalLB+Cilium overlap) · reserved ingress VIP,
autoAssign:false on reserved pools · new ingress on Gateway API (nginx pinned/legacy only) · data
plane ≥2 replicas + anti-affinity + PDB · both ClusterIssuers Ready · CA distributed · cert-expiry
alert independent of auto-renew.
PART B — PLATFORM CHECKLIST
Section titled “PART B — PLATFORM CHECKLIST”Each layer below is verified end-to-end before the next.
4. Storage — Longhorn
Section titled “4. Storage — Longhorn”⚠️ Lesson (recurring, the biggest storage risk). The team’s memory records Longhorn/iSCSI block-device I/O errors on worker nodes for months, and disk pressure on a shared root volume flips EXT4 read-only and kills the node (
CLAUDE.md). Before expanding Longhorn: diagnose and fix the pre-existing iSCSI/kernel errors, and give Longhorn a dedicated data disk (§0.4) — never the OS/containerd root. Otherwise the new platform inherits the same chronic failure.
longhornctl check preflight # dedicated disk registered, iscsid up, multipath okhelm install longhorn longhorn/longhorn --version 1.12.0 -n longhorn-system --create-namespace \ --set defaultSettings.replicaSoftAntiAffinity=false \ --set defaultSettings.allowVolumeCreationWithDegradedAvailability=false \ --set defaultSettings.replicaAutoBalance=least-effort \ --set defaultSettings.defaultDataPath=/var/lib/longhorn# Tiered StorageClasses: reclaimPolicy: Retain, allowVolumeExpansion: true,# numberOfReplicas 3 (prod-like) / 2 (stage), dataLocality, diskSelector/nodeSelector tags.# Mark exactly ONE default class.# BackupTarget → S3/NFS (+ credentials Secret); verify a MANUAL backup reaches Completed.# RecurringJob: daily backup (retain 7) + snapshot. DR volumes in a second cluster for critical data.⚠️ Only the V1 (iSCSI) data engine — do not enable V2 without NVMe + hugepages + kernel 6.7+. RWX PVCs route through a single share-manager pod per volume (a per-volume SPOF) — state the limit if any app needs ReadWriteMany. Under app-replicated databases (Postgres/RabbitMQ already replicate), drop Longhorn replicas to 1–2 to avoid ~9× write amplification on flaky iSCSI.
Checklist §4 — dedicated data disk · pre-existing iSCSI errors diagnosed · iscsid up +
multipath excludes longhorn · numberOfReplicas set · reclaimPolicy: Retain on all data classes
· exactly one default class · BackupTarget verified with a real Completed backup · expansion tested ·
only V1 engine · over-provisioning/min-available thresholds alerted.
5. Secrets — Vault (evaluate OpenBao) + VSO
Section titled “5. Secrets — Vault (evaluate OpenBao) + VSO”⚠️ License (critique — World Bank open-standards requirement). HashiCorp Vault is BSL (not OSI-approved) since 2023 — at odds with an anti-lock-in bid. Evaluate OpenBao (the OSI-licensed fork; drop-in) for anything tender-facing. The team already runs Vault 1.20.4; note the swap.
helm install vault hashicorp/vault -n adm --create-namespace \ --set server.ha.enabled=true --set server.ha.raft.enabled=true \ # Integrated Storage (Raft) --set server.ha.replicas=3 # 3 stage / 5 prod, anti-affined# TLS on 8200/8201; enable auto-unseal, Kubernetes auth, DB + PKI engines, >=2 audit devices, OIDC.⚠️ On-prem auto-unseal is the hard problem (critique). “Transit auto-unseal” needs another Vault (turtles); there is no cloud KMS on-prem. The real answers are a physical HSM via PKCS#11 (procurement/cost) or a documented manual unseal with a key-holder quorum (Shamir keys split offline). Decide and document — a Vault that boots sealed on every restart is a hard SPOF, and cert-manager (Vault PKI) and VSO have hard runtime dependencies on it: a seal stops cert renewal and secret sync platform-wide.
⚠️ Audit fail-closed footgun. Vault blocks all requests if an audit write fails (e.g. disk full). Two local file devices don’t help — require ≥1 remote/non-blocking sink + disk-usage alerting.
Consume secrets via VSO 1.5.0 (VaultStaticSecret/VaultDynamicSecret/VaultPKISecret with
rolloutRestartTargets). Use dynamic DB credentials and the PKI engine as the internal CA
feeding cert-manager (§3.3). OIDC (Keycloak) for operator access — but keep the offline recovery-key
break-glass outside OIDC.
Checklist §5 — odd node count, anti-affined · auto-unseal solved concretely (HSM or documented quorum) · Raft snapshots scheduled + restore tested · root token revoked · ≥2 audit devices with a redundant remote sink + disk alert · every mount has finite TTLs · PKI intermediate issuing short leafs · OpenBao evaluated for tender use.
6. GitOps — ArgoCD (HA)
Section titled “6. GitOps — ArgoCD (HA)”⚠️ Lesson (CG-002). ArgoCD with expired git credentials silently stopped reconciling all 8 apps (
ComparisonError), drift unknown, for an unknown time. Alert onsync_status != Synced,health != Healthy, andComparisonError. And because Git becomes production-critical, the Git repo needs HA/backup/mirror — a hard dependency the platform now has.
kubectl create ns argocdkubectl apply -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/v3.4.5/manifests/ha/install.yaml# redis-ha 3+3; server/repo-server >= 2; application-controller StatefulSet replicas == ARGOCD_CONTROLLER_REPLICAS# OIDC → Keycloak (groups claim) in argocd-cm; argocd-rbac-cm policy.default = role:readonly + group→role# One AppProject per tenant (sourceRepos/destinations/resource allow-lists). Default project unused for real workloads.# Secrets resolved on the destination via VSO — NEVER plaintext in Git (gitleaks clean).# Install Argo Rollouts 1.8.x; convert prod Deployments to Rollout with VictoriaMetrics AnalysisTemplate.Checklist §6 — HA, pinned v3.4.5, ≥3 nodes, redis-ha healthy · controller sharding correct · OIDC
login + groups work · AppProject per tenant, allow-lists · no plaintext Secret in Git · repo creds
centralised · alerts on sync/health/ComparisonError · allowEmpty:false; prune/selfHeal off for
stateful/storage apps · Git repo backed up/mirrored.
7. Policy + tenant isolation — Kyverno
Section titled “7. Policy + tenant isolation — Kyverno”helm install kyverno kyverno/kyverno --version 1.18.2 -n kyverno --create-namespace \ --set admissionController.replicas=3 # HA; separate reports/background/cleanup controllers# 1) Pod Security Admission: label every tenant ns pod-security.kubernetes.io/{enforce,audit,warn}=restricted (version-pinned)# 2) Kyverno GENERATE on namespace-create: default-deny NetworkPolicy + ResourceQuota + LimitRange + RoleBinding# 3) Kyverno VALIDATE (Audit→Enforce): require limits/requests, disallow privileged + :latest, restrict registries, require labels# 4) Kyverno verifyImages: cosign keyless (or key-pair air-gapped) + required SLSA provenance⚠️ Critique fixes — do not skip these:
- PSA is bypassable. Anyone who can edit a namespace can downgrade the
enforcelabel. Add a Kyverno policy that forbids mutating thepod-security.kubernetes.io/*labels — otherwise the “non-bypassable floor” isn’t.- Default-deny breaks DNS. A generated deny-all with no explicit egress to CoreDNS (UDP/TCP 53) breaks every pod. Put the CoreDNS egress allow in the generate bundle.
- Namespace ≠ security boundary. For hostile tenants, a shared kernel/node means a container escape (privileged, hostPath, hostPID/IPC/Network,
nodes/proxy) reaches other tenants. Document the threat model; use vCluster / per-tenant node pools for real isolation.- Webhook lockout.
failurePolicy: Failmeans a Kyverno outage blocks all admission — excludekube-systemandkyverno, spread replicas, settimeoutSeconds. Pick one engine (Kyverno) — don’t also run Gatekeeper.
⚠️ Keyless cosign is impossible air-gapped (needs public Fulcio/Rekor). An on-prem gov env needs self-hosted Sigstore or key-pair signing. Choose per environment.
Checklist §7 — Kyverno HA · PSA restricted on every tenant ns · policy forbidding PSA-label downgrade · validate policies Audit→Enforce · generate bundle active with CoreDNS egress · default-deny in 100% of tenant ns + cluster-wide Cilium backstop · verifyImages enforced (signing model chosen for air-gap) · failurePolicy scoped away from control plane · Policy Reporter → monitoring.
8. — (reserved: private registry)
Section titled “8. — (reserved: private registry)”⚠️ Gap the base story omitted (critique): a private container registry (Harbor or equivalent) is where signed images, SBOMs and cosign attestations actually live, and it must be mirrored for air-gap. Add Harbor as a platform component; it is a prerequisite for the supply-chain claims in §7 and the tender.
9. Observability — VictoriaMetrics (build this EARLY)
Section titled “9. Observability — VictoriaMetrics (build this EARLY)”helm install vmks vm/victoria-metrics-k8s-stack --version 1.148.0 -n monitoring --create-namespace# operator + vmagent + vmalert + Alertmanager + Grafana + node-exporter + kube-state-metrics.⚠️ Critique — do not ship VMSingle for the platform monitor. A single replica / single PVC for the very system whose job is “never be blind again” is self-defeating. Use VMCluster (RF≥2) for the platform monitor, or explicitly own the risk. Set retention + durable PVC; vmagent on-disk buffering.
Scrape: control plane (etcd fsync/commit p99, apiserver SLO, scheduler, controller-manager), Longhorn, cert-manager, ArgoCD, Vault-sealed, Cilium/Hubble, Keycloak. Alert packs: node/pod, etcd latency + DB size, apiserver SLO, Longhorn volume health, cert expiry 21 d/7 d, ArgoCD OutOfSync, Vault sealed, clock skew. Watchdog dead-man’s-switch → an external heartbeat.
⚠️ Critique — the dead-man’s-switch and notifications must be real on-prem. SaaS heartbeats (Dead Man’s Snitch/healthchecks.io) assume internet egress; a restricted gov network needs an on-prem heartbeat receiver and a named notification transport (SMTP relay? webhook?). “One tested channel” is exactly what was missing in the 12-day outage — name the actual transport. VictoriaLogs receives logs; OTel Collector stops exporting to
debug. 1-year audit retention needs downsampling (VM Enterprise) or unbounded growth — note the license/cost.
Checklist §9 — VMCluster RF≥2 (not VMSingle) · retention + durable storage · control-plane scrape live · Watchdog firing on an on-prem heartbeat · Alertmanager HA with a tested on-prem channel · alert packs loaded · VictoriaLogs receiving · OTel no longer debug-only · cardinality guardrails · all GitOps-managed.
10. SSO — Keycloak for the cluster AND every UI
Section titled “10. SSO — Keycloak for the cluster AND every UI”# Keycloak Operator + EXTERNAL clustered PostgreSQL (CloudNativePG) — never start-dev/H2.# Keycloak CR: instances >= 2, TLS, strict hostname, proxy headers.# One reusable 'groups' client scope (Group Membership mapper, Full path OFF) on every client.kube-apiserver → Keycloak (k8s ≥ 1.30 Structured Authentication Configuration):
--authentication-config with issuer = <realm>, username = preferred_username, groups = groups
(both oidc: prefixed); RBAC binds kind: Group. kubectl via kubelogin (oidc-login, PKCE).
Per-UI integration — native OIDC where it exists, oauth2-proxy where it doesn’t:
| UI | Method | Gotcha to get right |
|---|---|---|
| ArgoCD | native OIDC | needs a Keycloak group mapper or the groups claim is empty |
| Grafana | native OAuth | set role_attribute_strict=true, default Viewer — else everyone becomes Admin |
| Vault/OpenBao | native OIDC | external groups → policies; keep recovery-key break-glass outside OIDC |
| MinIO Console | native OIDC | policy claim required — no claim = zero access (looks like a broken login) |
| RabbitMQ mgmt | native oauth2 | aud must equal resource_server_id, not the client id, or login fails |
| pgAdmin | native OAuth2 | OIDC logs into pgAdmin only — DB creds + MASTER_PASSWORD remain; no in-app roles |
| Longhorn UI | oauth2-proxy | no native auth — a bare Ingress is a full storage console; gate on a group. All-or-nothing (no read-only) |
| Hubble UI | oauth2-proxy | no native auth; must not be exposed unprotected |
| K8s Dashboard | bearer token / oauth2-proxy | reconsider deploying it at all; HA oauth2-proxy mandatory |
⚠️ Critique — the two biggest traps:
- Break-glass + circular dependency. Cluster auth depends on Keycloak, which runs inside the cluster on CloudNativePG. If storage/DB/Keycloak is down, nobody can
kubectl— except via a sealed static cluster-admin client-cert kubeconfig kept offline. Keep exactly one, monitor its expiry, document the retrieval procedure. Each UI also needs a documented sealed local-admin fallback (ArgoCD admin, Grafana admin, Vault recovery keys, pgAdmin internal). Humans lose access within one 5–15 min token lifetime of a Keycloak outage; bound SA tokens keep GitOps/controllers alive.isssplit-horizon. The apiserver’s--authentication-configissuer must byte-match the tokeniss; an internal-vs-external hostname mismatch silently fails all OIDC. And the apiserver must reach Keycloak’s JWKS through Cilium’s service LB — ifk8sServiceHostis wrong or DNS is split-horizon, service LB and OIDC break together. oauth2-proxy is itself a SPOF — ≥2 replicas + shared session store (Redis) + cookie secret; watch for large-group header overflow (many groups → 502 on login).
Checklist §10 — Keycloak from Operator, ≥2 replicas, external clustered PG · TLS + strict hostname
· one groups scope on every client · apiserver --authentication-config, RBAC kind:Group · kubectl
via kubelogin · every UI behind Keycloak (native or oauth2-proxy) · sealed break-glass cert +
per-UI local-admin fallbacks documented · iss byte-matches · oauth2-proxy HA · MFA on *-admin ·
Keycloak events → logging · Keycloak realm-as-code + its Postgres in the DR plan.
PART C — OPERATE
Section titled “PART C — OPERATE”11. Data services (in-cluster, stage only)
Section titled “11. Data services (in-cluster, stage only)”Run only disposable/stage state in-cluster; keep the production system-of-record DB external
(the client’s Patroni+HAProxy+etcd+Keepalived — see [[PostgreSQL-HA-Production-Install-Guide]]),
referenced via ExternalName/Endpoints so manifests stay identical across environments.
| Service | Operator | Notes |
|---|---|---|
| PostgreSQL | CloudNativePG 1.30.x (or existing CrunchyData PGO) | 1 primary + 2 replicas; Barman/pgBackRest PITR to off-cluster S3; Longhorn replicas 1–2 |
| Redis/Valkey | redis-operator 0.26.x, Sentinel (1+2+3) | reconstructable cache; Valkey avoids Redis 8 AGPLv3 |
| RabbitMQ | Cluster Operator 2.22, RabbitMQ 4.x | quorum queues (classic mirrored removed in 4.0); definitions exported to git |
| Object storage | MinIO EOL risk (repo archived, last v7.1.1) | pin-frozen or adopt SeaweedFS / Garage / Rook-Ceph RGW |
Checklist §11 — each service via its operator CR, version-pinned in git · Retain-reclaim,
expandable StorageClass · Longhorn replicas reduced under replicated services · operator-native
off-cluster backup + tested restore · PDBs allow one-at-a-time drain (avoid the 0-disruption
trap seen on cg-stage) · exporters/ServiceMonitors wired · prod DB external · MinIO path chosen.
12. Platform disaster recovery (the hard part — critique)
Section titled “12. Platform disaster recovery (the hard part — critique)”Backing up etcd, Vault, Longhorn, Keycloak and PKI in isolation is not a DR plan. Total-cluster recovery has cross-dependencies and a required order.
- Consolidated DR runbook with restore ordering. Rough order: restore CA/
pki→ etcd snapshot (--force-new-cluster, re-add members) → bring up control plane + Cilium → restore Longhorn volumes (etcd references PVs that need Longhorn back) → unseal Vault (secrets that manifests reference) → Keycloak realm + its Postgres → verify. Rehearse it (game-days), not once. - Consistent backup point. etcd@T1 + Longhorn@T2 + Vault@T3 restore to an internally inconsistent cluster (PVCs with no backing volume, secrets pointing at rolled Vault paths). Coordinate a point-in-time or accept and document the reconciliation steps.
- Add Velero for k8s-native, granular namespace/PVC (CSI-snapshot) backup — etcd snapshot is all-or-nothing.
- Quorum-loss runbook (lose 2 of 3 etcd):
etcdctl snapshot restoresingle-member bootstrap → re-add. Written and rehearsed. - Immutability/offsite/encryption (S3 Object Lock vs ransomware) on every backup; site-level DR posture stated (single-DC on-prem = correlated failures — a WB bid usually needs a multi-site story).
- RPO/RTO per tier stated (etcd snapshot interval = data-loss window; Longhorn large-volume restore = full download; Vault/Keycloak restore windows).
13. Operating rules (from CLAUDE.md — non-negotiable)
Section titled “13. Operating rules (from CLAUDE.md — non-negotiable)”- No bulk cluster-wide mutations. PV/reclaim changes one volume at a time; snapshot to
backup/first; verify between each. Order: backups →Retain→ then GitOps. - Never trust an L4/TCP health check as a real signal — probe the actual protocol/SLI (the CG-012
false-
UPtrap that hid a 12-day outage). - Monitoring + alerting from day 0. Build §9 before you need it.
- Log state-changing commands to the cluster’s ops trail; be explicit about which cluster
(
--context) every command targets. - PriorityClasses for platform components (Cilium, monitoring, Vault, ingress) so app load can’t evict them on a fixed 6–9 node cluster.
Appendix A — scripts
Section titled “Appendix A — scripts”Idempotent install scripts live beside this guide under docs/Playbooks/scripts/k8s/ (to be generated
per environment from the parameterised steps above): 00-prep-node.sh, 10-bootstrap-cp.sh,
20-cilium.sh, 30-lb-ingress-tls.sh, 40-longhorn.sh, 50-vault.sh, 60-argocd.sh,
70-kyverno.sh, 80-victoriametrics.sh, 90-keycloak-sso.sh, dr/etcd-snapshot.sh,
dr/etcd-restore.sh. Each is a thin, commented wrapper around the commands above with the environment
values (VIPs, CIDRs, versions, domains) pulled from a single env.sh. Keep versions in env.sh and
re-validate them at install time.
Appendix B — the team’s real stack vs this guide
Section titled “Appendix B — the team’s real stack vs this guide”| Layer | As-built (k8s-u2 / cg-stage) | This guide |
|---|---|---|
| Control plane | EKS (u2) / single kubeadm master (cg-stage, a SPOF) | 3 CP + API VIP, self-managed |
| Version | v1.30 / v1.29.14 | v1.36.x (re-validate) |
| CNI | Calico 3.31.2 / 3.25 | Cilium (native routing, kube-proxy-free) + migration path |
| LB | MetalLB | Cilium LB-IPAM (L2 → BGP); MetalLB fallback |
| Ingress | ingress-nginx 1.14/1.12 | Gateway API (nginx pinned/legacy) |
| Storage | Longhorn 1.9.2 / local-path | Longhorn 1.12 on dedicated disk + Retain + backups |
| GitOps | ArgoCD 3.2.1 / 3.0.6 | ArgoCD 3.4.5 HA + OIDC + alerts |
| Secrets | Vault 1.20.4+VSO / none | Vault/OpenBao HA + solved auto-unseal |
| Policy | none | Kyverno + PSA + tenant isolation |
| Observability | partial / none (CG-006) | VictoriaMetrics from day 0 |
| SSO | Keycloak 26.3.3 present / none | Keycloak on the API + every UI + break-glass |
Appendix C — sources
Section titled “Appendix C — sources”Consolidated from the 2026-07 multi-agent research pass (14 dimensions, authoritative primaries:
kubernetes.io, docs.cilium.io, longhorn.io, cert-manager.io, argo-cd.readthedocs.io,
developer.hashicorp.com/openbao.org, kyverno.io, docs.victoriametrics.com, keycloak.org, cloudnative-pg.io,
CIS/NSA-CISA/NIST/ISO, worldbank.org procurement). Per-dimension source URLs are in the research
transcript under subagents/workflows/…/journal.jsonl.