Kubernetes¶
On this page
The container orchestrator: control-plane architecture, the core objects (Pods, Deployments, Services, config, storage), scaling, and how to debug a broken workload.
Basics¶
Kubernetes (K8s) automates deploying, scaling, and operating containerized apps across a cluster of machines. You declare desired state in YAML; controllers continuously reconcile actual state toward it.
Cluster architecture¶
┌──────────── Control Plane ────────────┐
│ kube-apiserver (the front door / API)│
│ etcd (cluster state store) │
│ kube-scheduler (places Pods on nodes) │
│ controller-mgr (reconciliation loops) │
└────────────────────────────────────────┘
│ (API)
┌────────────────┼────────────────┐
Node1 Node2 Node3
kubelet kubelet kubelet ← runs Pods, talks to API
kube-proxy kube-proxy kube-proxy ← service networking
containerd containerd containerd ← container runtime
Core objects¶
| Object | Purpose |
|---|---|
| Pod | Smallest unit: one or more containers sharing network/storage. Usually managed, not created directly. |
| ReplicaSet | Keeps N identical Pods running. |
| Deployment | Manages ReplicaSets; rolling updates & rollbacks. The default for stateless apps. |
| StatefulSet | Stable identity/storage for stateful apps (databases). |
| DaemonSet | One Pod per node (agents, log shippers). |
| Job / CronJob | Run-to-completion / scheduled tasks. |
| Service | Stable virtual IP + DNS load-balancing across Pods (ClusterIP / NodePort / LoadBalancer). |
| Ingress | HTTP(S) routing into the cluster (host/path rules) via an ingress controller. |
| ConfigMap / Secret | Inject configuration / sensitive data. |
| Namespace | Virtual cluster partition for isolation & quotas. |
| PersistentVolume / Claim | Decouple storage from Pods. |
Networking model¶
Every Pod gets its own IP; all Pods can reach each other (flat network via a CNI plugin
like Calico/Cilium/Flannel). Services provide stable access; kube-dns/CoreDNS
resolves svc.namespace.svc.cluster.local. NetworkPolicies restrict traffic.
Scaling¶
- HPA — Horizontal Pod Autoscaler scales Pod count on CPU/memory/custom metrics.
- VPA — adjusts Pod resource requests.
- Cluster Autoscaler — adds/removes nodes.
Cheatsheet¶
kubectl essentials¶
kubectl get pods -A # all namespaces
kubectl get deploy,svc,ingress -n app
kubectl describe pod mypod -n app # events at the bottom = gold
kubectl logs -f mypod -c container
kubectl exec -it mypod -- sh
kubectl apply -f manifest.yaml # declarative create/update
kubectl delete -f manifest.yaml
kubectl rollout status deploy/web
kubectl rollout undo deploy/web # roll back
kubectl scale deploy/web --replicas=5
kubectl get events --sort-by=.lastTimestamp -n app
kubectl top pod ; kubectl top node # needs metrics-server
kubectl port-forward svc/web 8080:80 # local access
kubectl config get-contexts # which cluster am I on?
A Deployment + Service¶
apiVersion: apps/v1
kind: Deployment
metadata: { name: web, namespace: app }
spec:
replicas: 3
selector: { matchLabels: { app: web } }
template:
metadata: { labels: { app: web } }
spec:
containers:
- name: web
image: ghcr.io/moin/web:1.0
ports: [{ containerPort: 8080 }]
resources:
requests: { cpu: "100m", memory: "128Mi" }
limits: { cpu: "500m", memory: "256Mi" }
readinessProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 5
livenessProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 15
---
apiVersion: v1
kind: Service
metadata: { name: web, namespace: app }
spec:
selector: { app: web }
ports: [{ port: 80, targetPort: 8080 }]
type: ClusterIP
Config & secrets¶
kubectl create configmap appcfg --from-literal=LOG_LEVEL=info
kubectl create secret generic dbcreds --from-literal=PASSWORD=s3cret
Thumb Rules¶
Rules of thumb
- Don't create bare Pods. Use a Deployment/StatefulSet so they self-heal.
- Always set resource requests/limits. Requests drive scheduling; missing requests cause noisy-neighbor chaos.
- Set readiness + liveness probes. Readiness gates traffic; liveness restarts hung containers. Don't make liveness too aggressive.
- Labels are the glue. Services/selectors/operators all work via labels — keep them consistent.
kubectl describeand events first when debugging, before reading logs.- Namespaces + quotas to isolate teams/environments.
- Secrets are base64, not encrypted at rest by default — enable encryption and/or external secret stores.
- Stateless on Deployments, stateful on StatefulSets with real PersistentVolumes.
- Pin image digests/tags; rolling updates need immutable images.
Use Cases¶
- Running microservices at scale with self-healing and rolling updates.
- Multi-tenant platforms — namespaces, quotas, RBAC, network policies.
- Batch & scheduled jobs (Jobs/CronJobs).
- Hybrid/multi-cloud — a consistent abstraction across providers.
- Autoscaling workloads — scale to demand, scale to zero with add-ons (KEDA).
- Platform engineering — operators/CRDs extend K8s for databases, CI, ML, etc.
Common Issues¶
Pod stuck in Pending
Scheduler can't place it: insufficient CPU/memory, no node matches nodeSelector/taints,
or an unbound PVC. kubectl describe pod shows the reason in Events.
CrashLoopBackOff
The container starts then exits/crashes repeatedly. kubectl logs --previous; common
causes: bad config/env, missing dependency, failing liveness probe, or app error.
ImagePullBackOff / ErrImagePull
Wrong image name/tag, private registry without an imagePullSecret, or rate limits.
Check the image path and registry credentials.
Service has no endpoints
The Service selector doesn't match Pod labels, or Pods aren't Ready (failing readiness
probe). kubectl get endpoints svc — empty means selector/readiness problem.
DNS not resolving inside the cluster
CoreDNS issues or NetworkPolicy blocking. Test from a debug Pod:
nslookup web.app.svc.cluster.local. Check CoreDNS pods/logs.
OOMKilled
Container exceeded its memory limit. Raise the limit or fix the leak; check
kubectl describe pod (Last State: OOMKilled).
Node NotReady
kubelet, container runtime, or network plugin problem, or disk/memory pressure.
kubectl describe node and check kubelet logs.
Best Practices¶
- GitOps: store manifests in Git, reconcile with Argo CD/Flux; no manual
kubectl applyin prod. - Set requests/limits, probes, and PodDisruptionBudgets for reliable rollouts.
- Use RBAC least-privilege, namespaces, NetworkPolicies, and Pod security standards.
- Manage secrets externally (External Secrets/Vault) and enable etcd encryption.
- Health, not hope: readiness/liveness/startup probes on every workload.
- Resource governance: ResourceQuotas + LimitRanges per namespace.
- Observability: metrics-server, Prometheus, centralized logs, and tracing.
- Package with Helm/Kustomize; keep environments DRY (see the Helm page).
- Plan upgrades: K8s releases roughly every ~4 months; test and stay within supported skew.
Official Sources¶
- Kubernetes Documentation — https://kubernetes.io/docs/home/
- Concepts — https://kubernetes.io/docs/concepts/
- kubectl reference — https://kubernetes.io/docs/reference/kubectl/
- Configuration best practices — https://kubernetes.io/docs/concepts/configuration/overview/
- Production best practices / cluster setup — https://kubernetes.io/docs/setup/best-practices/
- CNCF landscape — https://landscape.cncf.io/
- Argo CD (GitOps) — https://argo-cd.readthedocs.io/