Kubernetes Platform — Operations Skill
Acest conținut nu este încă disponibil în limba selectată.
You are helping operate or design a self-managed, on-prem Kubernetes platform. Reference stack: 3 control-plane + 6–9 workers, Cilium CNI, Longhorn storage, LoadBalancer VIPs, cert-manager, ArgoCD, Vault/OpenBao, Kyverno, VictoriaMetrics, Keycloak SSO. Answer from this; for depth read the bundled guides:
reference/Kubernetes-Platform-Production-Install-Guide.md— BASE + CHECKLIST + DR, with commands.reference/Kubernetes-Platform-Tender-Response.md— component catalogue, DevSecOps, SSO matrix, standards, Requirement Traceability Matrix. No commands.
⚠️ Golden rules (from real incidents — never violate)
Section titled “⚠️ Golden rules (from real incidents — never violate)”- Never run bare
kubectl. These environments have a mergedKUBECONFIGwhose default context points at the wrong cluster. Every command carries an explicit--context(or--kubeconfig). State which cluster each command targets. - Monitoring + alerting from day 0. A cluster with no metrics/alerting hid a 12-day database outage. Build VictoriaMetrics + a paging channel + a dead-man’s-switch early, not last.
- No bulk cluster-wide mutations. PV/reclaim changes one volume at a time; snapshot first; order
is backups →
Retain→ then GitOps. Log state-changing commands. - Never trust an L4/TCP health check as a real signal — probe the actual protocol/SLI. A live
proxy in front of a dead backend reports
UP(the trap that hid the 12-day outage).
Fast cluster health sweep (read-only)
Section titled “Fast cluster health sweep (read-only)”K="kubectl --context=<ctx> --request-timeout=30s"$K get nodes -o wide # Ready? versions? pressure?$K get pods -A | grep -vE 'Running|Completed' # ⚠️ CrashLoopBackOff shows phase=Running — also:$K get pods -A -o wide | grep -iE 'CrashLoop|Error|Pending|ImagePull'$K get events -A --field-selector type=Warning | tail -20$K -n argocd get applications # Unknown/OutOfSync = GitOps not reconciling$K get pv,pvc -A | grep -v Bound # storage stuck$K top nodes 2>&1 # "Metrics API not available" = no metrics-server (a finding)⚠️ A pod in
CrashLoopBackOffkeepsphase: Running— a naive--field-selector=status.phase!=Runningfilter misses it. Always also grep the STATUS column. (This mistake was made once during onboarding.)
Diagnosis by symptom
Section titled “Diagnosis by symptom”| Symptom | First checks | Likely cause |
|---|---|---|
All ArgoCD apps Unknown / ComparisonError | app conditions; git creds | expired git credentials — GitOps silently stopped reconciling; drift unknown (fix creds carefully, see below) |
Pods CrashLoopBackOff on a backing service | logs --previous; the dependency | very often the external database is down, not k8s — trace the DB (use the postgres-ha-ops skill) |
Pod Pending | describe pod; get pvc; get sc | no default StorageClass (PVC hangs), or no node fits, or PDB/taint |
Node NotReady / read-only FS | describe node; disk pressure | shared root disk full flips EXT4 read-only (recurring); or Longhorn/iSCSI I/O errors |
| Control-plane probes flapping | logs etcd; apply request took too long | slow etcd disk — needs NVMe / widened timeouts; and etcd hygiene |
LoadBalancer <pending> / no VIP | MetalLB/Cilium pool; one announcer? | pool exhausted, or MetalLB and Cilium both advertising (overlap) |
| Ingress serves plain HTTP only | get clusterissuers; ingress tls: | cert-manager installed with zero ClusterIssuers — annotations lie |
Component quick-reference (what “healthy” looks like + the gotcha)
Section titled “Component quick-reference (what “healthy” looks like + the gotcha)”- etcd: odd quorum (3/5, never 2). Auto-compaction + quota set; alert DB-size.
NOSPACE→ compact → defrag (per member) →alarm disarm. On NVMe, never shared with the DB. - Cilium: the #1 bootstrap failure the team hit —
kubeProxyReplacement=trueneedsk8sServiceHost/k8sServicePort= the API VIP, not a node IP.cilium statusall green,cilium connectivity testpasses. Kernel ≥ 5.10. Choose DSR to preserve client IP. - Longhorn: dedicated data disk (never OS/containerd root — the read-only-FS killer),
iscsidup,reclaimPolicy: Retain, one default class, a verified off-cluster backup. Reduce replicas to 1–2 under app-replicated DBs. Diagnose pre-existing iSCSI errors before expanding. - LoadBalancer: exactly one announcer per pool (no MetalLB+Cilium overlap). Start L2, add BGP later with the network team.
- cert-manager: must have a working ClusterIssuer (Vault PKI internal + ACME DNS-01 for wildcards). Alert on cert expiry < 14 d independent of auto-renew.
- ArgoCD: alert on
sync != Synced,health != Healthy,ComparisonError. 7-of-8 apps withprune:true, selfHeal:trueoverDelete-reclaim PVs is a data-loss path — fix credentials with prune/selfHeal off first, read the diff, reconcile deliberately,admlast. - Vault/OpenBao: a sealed Vault is a hard SPOF that boots sealed on every restart and takes cert-manager (Vault PKI) + VSO down with it. On-prem auto-unseal needs HSM/PKCS#11 or a documented manual quorum — “Transit” just moves the problem. Vault is BSL → evaluate OpenBao for open-standards bids.
- Kyverno: PSA
restrictedfloor + Kyverno above it. PSA labels are mutable — add a policy forbidding their downgrade. Default-deny NetworkPolicy must allow CoreDNS egress (UDP/TCP 53) or every pod breaks.failurePolicy: Fail→ excludekube-system/kyvernoor you lock out admission. Namespace ≠ security boundary (vCluster/node-pools for hostile tenants). - VictoriaMetrics: use VMCluster RF≥2 for the platform monitor (not VMSingle). Dead-man’s- switch to an on-prem heartbeat; name a real on-prem notification transport.
- Keycloak SSO: the API + every UI via Keycloak. Break-glass is the crux: cluster auth depends
on Keycloak running in the cluster → keep a sealed offline cluster-admin cert, monitor its
expiry, and a per-UI local-admin fallback (ArgoCD admin, Grafana admin, Vault recovery keys). The
apiserver
issmust byte-match the token issuer or all OIDC silently fails. oauth2-proxy (for Longhorn/Hubble/Dashboard) must be HA. See the SSO matrix in the tender doc.
Onboarding an unknown cluster (monitoring mode)
Section titled “Onboarding an unknown cluster (monitoring mode)”Read-only first; document before you change anything. Sweep: nodes/versions → namespaces →
workloads+images → network (CNI, ingress, LB pools) → storage (SC, PV, reclaim) → GitOps (why not
reconciling) → security (cluster-admin bindings, plaintext secrets, TLS, NetworkPolicies) →
observability (is anything actually alerting?). Log every command tagged [RO]. Open findings with
severities and evidence; nothing closes without observed proof. (This is the process used to
onboard the ChisinauGaz cg-stage cluster — see the workspace docs/CG-Stage/.)
DevSecOps posture (for design/tender questions)
Section titled “DevSecOps posture (for design/tender questions)”GitOps as the only deployment path (no kubectl apply drift); shift-left security (SAST/SCA/secret-
scan/IaC scan as blocking gates); supply chain (SBOM + cosign signing + SLSA provenance, enforced
at admission by Kyverno); private registry (Harbor) mirrored for air-gap; runtime detection
(Tetragon/Falco); DORA metrics; blameless post-mortems. Map to NIST SSDF, CIS, OWASP, SLSA. Full
methodology + standards + the Requirement Traceability Matrix are in the tender doc.
For tender / NFR questions (World Bank grade)
Section titled “For tender / NFR questions (World Bank grade)”Lead with the Requirement Traceability Matrix (requirement → component → standard → numeric target → evidence artefact). Include the licence column (flag Vault-BSL→OpenBao, MinIO-EOL→SeaweedFS, Redis-AGPL→Valkey — the anti-lock-in differentiator). Keep availability honest (99.9–99.95 % for single-site on-prem). Remember a World Bank IS bid is ~half non-technical (methodology, CVs, SLA, training, financials) — flag that scope.