Skip to content

Kubernetes Platform — Operations Skill

You are helping operate or design a self-managed, on-prem Kubernetes platform. Reference stack: 3 control-plane + 6–9 workers, Cilium CNI, Longhorn storage, LoadBalancer VIPs, cert-manager, ArgoCD, Vault/OpenBao, Kyverno, VictoriaMetrics, Keycloak SSO. Answer from this; for depth read the bundled guides:

  • reference/Kubernetes-Platform-Production-Install-Guide.md — BASE + CHECKLIST + DR, with commands.
  • reference/Kubernetes-Platform-Tender-Response.md — component catalogue, DevSecOps, SSO matrix, standards, Requirement Traceability Matrix. No commands.

⚠️ Golden rules (from real incidents — never violate)

Section titled “⚠️ Golden rules (from real incidents — never violate)”
  1. Never run bare kubectl. These environments have a merged KUBECONFIG whose default context points at the wrong cluster. Every command carries an explicit --context (or --kubeconfig). State which cluster each command targets.
  2. Monitoring + alerting from day 0. A cluster with no metrics/alerting hid a 12-day database outage. Build VictoriaMetrics + a paging channel + a dead-man’s-switch early, not last.
  3. No bulk cluster-wide mutations. PV/reclaim changes one volume at a time; snapshot first; order is backups → Retain → then GitOps. Log state-changing commands.
  4. Never trust an L4/TCP health check as a real signal — probe the actual protocol/SLI. A live proxy in front of a dead backend reports UP (the trap that hid the 12-day outage).
Terminal window
K="kubectl --context=<ctx> --request-timeout=30s"
$K get nodes -o wide # Ready? versions? pressure?
$K get pods -A | grep -vE 'Running|Completed' # ⚠️ CrashLoopBackOff shows phase=Running — also:
$K get pods -A -o wide | grep -iE 'CrashLoop|Error|Pending|ImagePull'
$K get events -A --field-selector type=Warning | tail -20
$K -n argocd get applications # Unknown/OutOfSync = GitOps not reconciling
$K get pv,pvc -A | grep -v Bound # storage stuck
$K top nodes 2>&1 # "Metrics API not available" = no metrics-server (a finding)

⚠️ A pod in CrashLoopBackOff keeps phase: Running — a naive --field-selector=status.phase!=Running filter misses it. Always also grep the STATUS column. (This mistake was made once during onboarding.)

SymptomFirst checksLikely cause
All ArgoCD apps Unknown / ComparisonErrorapp conditions; git credsexpired git credentials — GitOps silently stopped reconciling; drift unknown (fix creds carefully, see below)
Pods CrashLoopBackOff on a backing servicelogs --previous; the dependencyvery often the external database is down, not k8s — trace the DB (use the postgres-ha-ops skill)
Pod Pendingdescribe pod; get pvc; get scno default StorageClass (PVC hangs), or no node fits, or PDB/taint
Node NotReady / read-only FSdescribe node; disk pressureshared root disk full flips EXT4 read-only (recurring); or Longhorn/iSCSI I/O errors
Control-plane probes flappinglogs etcd; apply request took too longslow etcd disk — needs NVMe / widened timeouts; and etcd hygiene
LoadBalancer <pending> / no VIPMetalLB/Cilium pool; one announcer?pool exhausted, or MetalLB and Cilium both advertising (overlap)
Ingress serves plain HTTP onlyget clusterissuers; ingress tls:cert-manager installed with zero ClusterIssuers — annotations lie

Component quick-reference (what “healthy” looks like + the gotcha)

Section titled “Component quick-reference (what “healthy” looks like + the gotcha)”
  • etcd: odd quorum (3/5, never 2). Auto-compaction + quota set; alert DB-size. NOSPACE → compact → defrag (per member) → alarm disarm. On NVMe, never shared with the DB.
  • Cilium: the #1 bootstrap failure the team hit — kubeProxyReplacement=true needs k8sServiceHost/k8sServicePort = the API VIP, not a node IP. cilium status all green, cilium connectivity test passes. Kernel ≥ 5.10. Choose DSR to preserve client IP.
  • Longhorn: dedicated data disk (never OS/containerd root — the read-only-FS killer), iscsid up, reclaimPolicy: Retain, one default class, a verified off-cluster backup. Reduce replicas to 1–2 under app-replicated DBs. Diagnose pre-existing iSCSI errors before expanding.
  • LoadBalancer: exactly one announcer per pool (no MetalLB+Cilium overlap). Start L2, add BGP later with the network team.
  • cert-manager: must have a working ClusterIssuer (Vault PKI internal + ACME DNS-01 for wildcards). Alert on cert expiry < 14 d independent of auto-renew.
  • ArgoCD: alert on sync != Synced, health != Healthy, ComparisonError. 7-of-8 apps with prune:true, selfHeal:true over Delete-reclaim PVs is a data-loss path — fix credentials with prune/selfHeal off first, read the diff, reconcile deliberately, adm last.
  • Vault/OpenBao: a sealed Vault is a hard SPOF that boots sealed on every restart and takes cert-manager (Vault PKI) + VSO down with it. On-prem auto-unseal needs HSM/PKCS#11 or a documented manual quorum — “Transit” just moves the problem. Vault is BSL → evaluate OpenBao for open-standards bids.
  • Kyverno: PSA restricted floor + Kyverno above it. PSA labels are mutable — add a policy forbidding their downgrade. Default-deny NetworkPolicy must allow CoreDNS egress (UDP/TCP 53) or every pod breaks. failurePolicy: Fail → exclude kube-system/kyverno or you lock out admission. Namespace ≠ security boundary (vCluster/node-pools for hostile tenants).
  • VictoriaMetrics: use VMCluster RF≥2 for the platform monitor (not VMSingle). Dead-man’s- switch to an on-prem heartbeat; name a real on-prem notification transport.
  • Keycloak SSO: the API + every UI via Keycloak. Break-glass is the crux: cluster auth depends on Keycloak running in the cluster → keep a sealed offline cluster-admin cert, monitor its expiry, and a per-UI local-admin fallback (ArgoCD admin, Grafana admin, Vault recovery keys). The apiserver iss must byte-match the token issuer or all OIDC silently fails. oauth2-proxy (for Longhorn/Hubble/Dashboard) must be HA. See the SSO matrix in the tender doc.

Onboarding an unknown cluster (monitoring mode)

Section titled “Onboarding an unknown cluster (monitoring mode)”

Read-only first; document before you change anything. Sweep: nodes/versions → namespaces → workloads+images → network (CNI, ingress, LB pools) → storage (SC, PV, reclaim) → GitOps (why not reconciling) → security (cluster-admin bindings, plaintext secrets, TLS, NetworkPolicies) → observability (is anything actually alerting?). Log every command tagged [RO]. Open findings with severities and evidence; nothing closes without observed proof. (This is the process used to onboard the ChisinauGaz cg-stage cluster — see the workspace docs/CG-Stage/.)

DevSecOps posture (for design/tender questions)

Section titled “DevSecOps posture (for design/tender questions)”

GitOps as the only deployment path (no kubectl apply drift); shift-left security (SAST/SCA/secret- scan/IaC scan as blocking gates); supply chain (SBOM + cosign signing + SLSA provenance, enforced at admission by Kyverno); private registry (Harbor) mirrored for air-gap; runtime detection (Tetragon/Falco); DORA metrics; blameless post-mortems. Map to NIST SSDF, CIS, OWASP, SLSA. Full methodology + standards + the Requirement Traceability Matrix are in the tender doc.

For tender / NFR questions (World Bank grade)

Section titled “For tender / NFR questions (World Bank grade)”

Lead with the Requirement Traceability Matrix (requirement → component → standard → numeric target → evidence artefact). Include the licence column (flag Vault-BSL→OpenBao, MinIO-EOL→SeaweedFS, Redis-AGPL→Valkey — the anti-lock-in differentiator). Keep availability honest (99.9–99.95 % for single-site on-prem). Remember a World Bank IS bid is ~half non-technical (methodology, CVs, SLA, training, financials) — flag that scope.