The Exploded Cluster teaches how modern container platforms work by taking them apart - +
The Exploded Cluster teaches how modern container platforms work by taking them apart, literally. Each course is one machine drawn as a single exploded illustration, sliced into its real components and wired to your scroll, so the architecture moves while the words explain it. Start with the aperitif's four terminal commands - three that deliver, one that checks - and @@ -74,7 +74,7 @@ kubectl in the docs, type whichever your cluster gives you.
That 1/1 reads as containers-ready over containers-wanted: a pod can hold more than one, which is Course IIIb.
-Three of those were delivery actions - build, push, apply - and the +
Three of those were delivery actions - build, push, apply, and the last, oc get pods, only checked the result. Built of what, exactly? Pushed to where, and what travelled? Applied, which is not the same as launched. Running, according to whom? Each course below takes one of those words apart.
@@ -87,7 +87,7 @@Course I · Podman
Scroll, and the thing you keep calling "a container image" comes apart in your - hands. Four layers. Each one only stores what changed from the layer under it - and here they + hands. Four layers. Each one only stores what changed from the layer under it, and here they rise one at a time, bottom up.
If you arrived here from Linux, take this translation first: a running container is an - ordinary process on your kernel. No guest OS, no hypervisor. The kernel gives it - namespaces so it sees its own PID tree, mounts, network and hostname, and - cgroups so its CPU and memory can be capped. ps on the host + ordinary process on your kernel. No guest OS, no hypervisor. The kernel gives it + namespaces so it sees its own PID tree, mounts, network and hostname, and + cgroups so its CPU and memory can be capped. ps on the host lists it. kill on the host kills it. Podman leans into that: no - daemon sits in the middle - the container is a child of your own shell - and rootless mode + daemon sits in the middle - the container is a child of your own shell, and rootless mode maps your user onto root inside the container through a user namespace, so root in there is an unprivileged UID out here.
-So what does the image provide? The filesystem that process sees. That is the whole +
So what does the image provide? The filesystem that process sees. That is the whole job, and it is why an image is a stack of layers rather than a disk image.
Field note. Prove it on your own box. @@ -123,19 +123,19 @@ it is exotic. It is your kernel, described differently.
An image is not a copy of a machine. It is a stack of read-only layers, each +
An image is not a copy of a machine. It is a stack of read-only layers, each recording only what changed from the one beneath. Layers are content-addressed, so an identical layer is stored once and reused by every image that references it, so a pull fetches just the layers you do not already have. A base sits at the bottom, your dependencies on it, your code - usually the smallest layer, and typically the most volatile - above that.
-Those layers become one filesystem through a union mount - overlayfs, the same +
Those layers become one filesystem through a union mount - overlayfs, the same kernel feature you can mount by hand. The read-only image layers are the lower dirs; the - container gets a fresh upper dir of its own. Writes land in the upper, and editing an + container gets a fresh upper dir of its own. Writes land in the upper, and editing an existing file copies it up there first, leaving the image layer untouched underneath.
-That upper dir is the top slab in the scene, and it is the odd one out: the writable - layer is not part of the image. The runtime creates it with the container and +
That upper dir is the top slab in the scene, and it is the odd one out: the writable + layer is not part of the image. The runtime creates it with the container and discards it when that container dies. Nothing written there ships, and nothing written there - survives - which is the entire reason volumes exist.
+ survives, which is the entire reason volumes exist.Change a layer and every layer above it must be rebuilt.
The builder caches layer by layer, and a cached layer survives only while everything @@ -160,7 +160,7 @@
Course II · Kubernetes
You never tell Kubernetes how to run your app. You describe what you - want - declarative intent - and the cluster works continuously to make it true. Scroll, and + want - declarative intent, and the cluster works continuously to make it true. Scroll, and the formation splits: the half that decides rises, the halves that run spread below.
The habit underneath everything: the reconciliation loop - compare desired state +
The habit underneath everything: the reconciliation loop - compare desired state against actual state, fix the difference, repeat. That oc apply didn't launch anything; it filed paperwork. The machine took it from there, and it keeps taking it from there: kill a pod and it returns, not because something noticed the crash but because the loop noticed the difference.
Kubernetes doesn't run your app - it reconciles it.
The split in the scene is the split that makes everything else possible: a control - plane that decides - holds the truth, schedules, reconciles - and worker nodes +
The split in the scene is the split that makes everything else possible: a control + plane that decides - holds the truth, schedules, reconciles, and worker nodes that run pods. The workers are deliberately interchangeable: identical, replaceable, cattle from day one. Authority does not live where the work happens.
Field note. If you SSH into a node to "fix" a workload, a controller may reconcile your change away. Change the desired state instead, and let the loop carry it.
The core control plane is the API server, scheduler and - controller-manager, backed by etcd; cloud deployments may also run a +
The core control plane is the API server, scheduler and + controller-manager, backed by etcd; cloud deployments may also run a cloud-controller-manager. Naming them makes the later error messages readable. The API server (kube-apiserver) is the only door: requests authenticate, are - authorised and are admitted there, and it is the one component that talks to etcd, the + authorised and are admitted there, and it is the one component that talks to etcd, the key-value store holding Kubernetes API state. Lose etcd and you have lost the cluster, which is why backing it up is the homework to not skip. The scheduler decides which node a new pod belongs on - packing by the resources a pod requests, not by what it currently uses - @@ -202,10 +202,10 @@ reconciliation loops. Control-plane state converges through the API server: the scheduler and built-in controllers watch API objects and write decisions back through the API rather than modifying etcd directly.
-One of those loops is the chain you will debug most: a Deployment creates a - ReplicaSet, and the ReplicaSet creates pods. That is why the deploy in the +
One of those loops is the chain you will debug most: a Deployment creates a + ReplicaSet, and the ReplicaSet creates pods. That is why the deploy in the aperitif printed deployment.apps/api created and you then went - looking for a pod - and why, when a rollout is stuck with no pod at all, the answer is + looking for a pod, and why, when a rollout is stuck with no pod at all, the answer is upstream in that chain rather than on any node.
etcd is the durable store for Kubernetes API state; controllers reconcile from it continually.
Everything the delivery arc teaches from Course VI onward is the same @@ -222,14 +222,14 @@
Course IIIa · The node
Everything so far was decision. This is the machine where a pod stops being - paperwork and starts being a process - and the chain assembles link by link as you scroll.
+ paperwork and starts being a process, and the chain assembles link by link as you scroll.The scheduler decides which node a pod belongs on; the kubelet on that node - makes the assigned PodSpec real; the runtime creates and starts the containers. The - kubelet does not place pods, and does not itself create containers - it speaks CRI (the - Container Runtime Interface) to containerd or CRI-O, and that runtime pulls the image +
The scheduler decides which node a pod belongs on; the kubelet on that node + makes the assigned PodSpec real; the runtime creates and starts the containers. The + kubelet does not place pods, and does not itself create containers - it speaks CRI (the + Container Runtime Interface) to containerd or CRI-O, and that runtime pulls the image (through the mirror of Course VII) and starts it, handing the low-level work to runc or crun. - A CNI plugin (Container Network Interface) gives the pod a real IP - called by the - runtime, not the kubelet - and kube-proxy, or an eBPF datapath replacing it, makes + A CNI plugin (Container Network Interface) gives the pod a real IP - called by the + runtime, not the kubelet, and kube-proxy, or an eBPF datapath replacing it, makes Service addresses route to real pods.
The scheduler places the pod. The kubelet realises it. The runtime runs it.
Field note. Node NotReady? The kubelet is a systemd unit @@ -279,26 +279,26 @@
Everything inside the jacket shares a network namespace: one IP, localhost between +
Everything inside the jacket shares a network namespace: one IP, localhost between friends. Volumes are declared once on the pod, but each container mounts the ones it needs - - sharing storage is opt-in, not automatic. regular init containers run in order and - complete before the app containers start. Native sidecars are the deliberate exception: + sharing storage is opt-in, not automatic. regular init containers run in order and + complete before the app containers start. Native sidecars are the deliberate exception: restartable init containers that keep running alongside the app, for jobs like proxy, logs and config reload.
startup owns warm-up, readiness gates traffic, liveness restarts the +
startup owns warm-up, readiness gates traffic, liveness restarts the truly hung. Confusing them is how healthy pods get executed - a slow start killed by an impatient liveness probe looks exactly like a crash.
containerPort is documentation - unless a Service targets it by name, or you use hostPort. Either way the app still has to bind the port itself.
-Field note. Exit 137 is 128 + 9: the process was SIGKILLed. +
Field note. Exit 137 is 128 + 9: the process was SIGKILLed. It does not say by whom. The kernel's OOM killer surfaces as reason OOMKilled; a failed liveness probe shows up in the pod's events; eviction and node pressure look different again. One exit code, several possible crimes - read the termination reason and the events, never the number alone: oc describe pod <name> shows both together, and oc logs --previous shows what the dead container said on its way out. Worth connecting to what you already know: a memory limit becomes a - cgroup ceiling, and the kernel's OOM killer enforces it just as it would for any + cgroup ceiling, and the kernel's OOM killer enforces it just as it would for any other process on the box.
The pod is the smallest schedulable unit - the jacket, never the
container. Once that distinction lands, half of Kubernetes networking stops being
@@ -314,7 +314,7 @@
Course IV · The traffic Pods die, respawn and change addresses - and traffic still arrives. Scroll,
+ Pods die, respawn and change addresses, and traffic still arrives. Scroll,
and the delivery route assembles checkpoint by checkpoint; watch what happens to the pod
that stops answering.Pods die constantly.
-
The address does not.
+ ready pod, and a dark unlit pod fallen out of the line.">
A Service is the fixed point: a ClusterIP inside, a LoadBalancer at the edge, +
A Service is the fixed point: a ClusterIP inside, a LoadBalancer at the edge, Ingress or the Gateway API doing host- and path-routing above. Clients hold the name; the pods behind it come and go without anyone being told.
A Service holds no list of pods. It holds a label selector - match - app: api - and a controller continuously matches that against every - pod in the namespace - a Kubernetes namespace this time, a naming boundary for - objects, no relation to the kernel namespaces of Course I - keeping the passing set in an - EndpointSlice. Labels are how +
A Service holds no list of pods. It holds a label selector (say + app: api), and a controller continuously matches that against every + pod in the namespace. That is a Kubernetes namespace, a naming boundary for objects, unrelated + to the kernel namespaces of Course I; the passing set lands in an EndpointSlice. Labels are how everything here finds everything else, from a Service picking pods to a fleet hub picking whole clusters (Course VI). Wear the label and you are eligible; readiness decides whether you stay.
-A ClusterIP is a virtual Service address implemented by the node dataplane rather +
A ClusterIP is a virtual Service address implemented by the node dataplane rather than an application process listening on that IP. In iptables or nftables mode, - kube-proxy installs rules that steer Service traffic to endpoint IPs; if you have + kube-proxy installs rules that steer Service traffic to endpoint IPs; if you have written a DNAT rule by hand, that mode will feel familiar. In IPVS mode, kube-proxy binds Service IPs to the kube-ipvs0 dummy interface and creates IPVS virtual servers. eBPF implementations can replace kube-proxy with their own dataplane. Cluster - DNS (CoreDNS) resolves api.myns.svc.cluster.local to the + DNS (CoreDNS) resolves api.myns.svc.cluster.local to the Service address.
When a configured readiness probe fails, the pod becomes unready and normal +
When a configured readiness probe fails, the pod becomes unready and normal Kubernetes Service traffic stops selecting it for new connections. There is no error at the client - the pod drops out of the endpoint set. That is the feature: broken instances take themselves out of rotation. It is also the first place to look when traffic "disappears".
@@ -363,7 +362,7 @@ the only address worth knowing.For ordinary Service traffic: address the Service, not a pod.
Field note. "The network is broken" after a deploy is usually - readiness telling the truth about your app - not the network lying about your packets.
+ readiness telling the truth about your app, not the network lying about your packets.The dark pod in the scene is not an error state - it is the system working. For pods with readiness probes, eligibility for Service traffic is continually re-evaluated for the life of the pod.
@@ -393,20 +392,20 @@The SCC - Security Context Constraint - is admission deciding what a pod may BE, checked +
The SCC - Security Context Constraint - is admission deciding what a pod may BE, checked by the api-server when the pod is created and before any node sees it. Under OpenShift's restricted - SCCs, workloads normally run as a non-root UID drawn from the project's allocated range, + SCCs, workloads normally run as a non-root UID drawn from the project's allocated range, so an image has to work with an arbitrary permitted UID rather than assuming a fixed user. Workloads that genuinely need privilege get a dedicated ServiceAccount bound to the minimum - SCC that grants it - not the stock one.
+ SCC that grants it, not the stock one.Routes predate Ingress and still rule here. An Operator is a controller paired with a +
Routes predate Ingress and still rule here. An Operator is a controller paired with a custom resource: you describe what you want in YAML, and its controller builds it and keeps it - true - an Operator applies the same reconciliation pattern to an application or platform + true. An Operator applies the same reconciliation pattern to an application or platform capability. The split matters: the platform's own operators are driven by the Cluster Version - Operator, while OLM installs and upgrades the add-on Operators you choose from - OperatorHub. MachineConfig is the supported declarative path for node OS configuration; - ad-hoc SSH changes create drift and are not the intended operating model. And SNO - + Operator, while OLM installs and upgrades the add-on Operators you choose from + OperatorHub. MachineConfig is the supported declarative path for node OS configuration; + ad-hoc SSH changes create drift and are not the intended operating model. And SNO - single-node OpenShift - puts a whole cluster on one box at the edge. At fleet scale the labels from Course VI decide which of these boxes runs what.
On OpenShift, admission is the interview - the SCC is the dress code.
@@ -445,22 +444,22 @@C-II taught the reconciliation loop: desired versus actual, fix the difference, repeat. - Argo CD applies the same habit to delivery. An Application names a repo, a path - and a revision - watch this branch of this repository - and the controller renders what it + Argo CD applies the same habit to delivery. An Application names a repo, a path + and a revision - watch this branch of this repository, and the controller renders what it finds there, compares it against the live cluster, and syncs the difference. The deploy button is a Git commit; the change history is git log; code review is change control. Git records intent - the cluster's own audit log and Argo CD's sync history record what actually happened, including what Git does not see.
oc reads the running system; Git records the change.
Here the cluster pulls. Push-based delivery exists and is still GitOps to many; the +
Here the cluster pulls. Push-based delivery exists and is still GitOps to many; the pull variant is worth choosing deliberately, because no CI system, laptop or build pipeline holds a credential that can touch the cluster - the agent inside holds a read-only deploy key and the trust arrow points out. Compromise the build system and you can propose a change, which is visible; you cannot reach into production through Git. It can still push images, which is the other half of why a manifest should name the digest, not the tag. - Hand-edit a live object and the controller flags it OutOfSync. With self-heal enabled, - Argo CD can restore the desired state; with prune enabled, objects removed from Git can + Hand-edit a live object and the controller flags it OutOfSync. With self-heal enabled, + Argo CD can restore the desired state; with prune enabled, objects removed from Git can be deleted. Both are opt-in: without them the reconciler reports the drift and waits for a sync. Rollback is git revert, which is why commit hygiene is an operational skill.
@@ -470,8 +469,8 @@ go missing.Run many OpenShift clusters under a hub - RHACM, Red Hat's fleet manager, with edge clusters - arriving through zero-touch provisioning - and nobody applies apps to clusters by hand. Each app carries a - Placement that selects cluster labels; the hub matches placements against the + arriving through zero-touch provisioning, and nobody applies apps to clusters by hand. Each app carries a + Placement that selects cluster labels; the hub matches placements against the labels a cluster wears, and the chosen cluster is given its assignment - push-style from the hub by default, or in a pull model where each cluster runs its own reconciler. Labelling a cluster changes its placement eligibility; the lifecycle and cleanup of what lands still follow @@ -481,11 +480,11 @@
In this architecture, plaintext secret values do not live in Git: Git stores the ExternalSecret reference, and the vault stores the value. Commit a plaintext secret and you - should treat it as compromised from that moment - deleting it later does not guarantee + should treat it as compromised from that moment - deleting it later does not guarantee it is gone from history, forks, clones, CI caches or backups. So Git carries an ExternalSecret - naming a logical key; a vault (Azure Key Vault in the worked example) holds the value; + naming a logical key; a vault (Azure Key Vault in the worked example) holds the value; an operator inside the cluster exchanges one for the other at runtime. Rotation happens in the - vault, not as a commit - which is why the vault sits apart from the pipeline in the scene + vault, not as a commit, which is why the vault sits apart from the pipeline in the scene above.
Git holds the shape of the secret. The vault holds the secret.
Delivery becomes a property rather than an event: the cluster converges on
@@ -506,11 +505,11 @@
Course VII · The image supply chain myapp:latest feels like a name. It is a sticky note - a
+ myapp:latest feels like a name. It is a sticky note: a
mutable reference anyone with push rights can move from one image to another. Registries may
- audit tag updates, but the tag itself carries no immutability guarantee. The digest - the
- sha256 of the image manifest, the small JSON index listing an image's layers, and no
- relation to the YAML manifests you apply to a cluster - identifies the exact uploaded content.
+ audit tag updates, but the tag itself carries no immutability guarantee. The digest is the
+ sha256 of the image manifest (the small JSON index listing an image's layers, unrelated
+ to the YAML manifests you apply to a cluster). It identifies the exact uploaded content.
A local image ID printed by a build is not necessarily the same value, so pin the registry
digest that consumers actually pull.A tag is a promise.
-
A digest is a fact.
Between the build and the node sits the registry chain. Upstream, a managed registry - Azure - Container Registry in the worked example - holds what CI built. In front of the cluster sits a - mirror: a pull-through cache like zot. On a cache miss the mirror fetches the artefact +
Between the build and the node sits the registry chain. Upstream, a managed registry (Azure + Container Registry in the worked example) holds what CI built. In front of the cluster sits a + mirror: a pull-through cache like zot. On a cache miss the mirror fetches the artefact from upstream; subsequent requests can be served locally while that content stays cached. Rate limits, egress cost, disconnected sites, control - one place to gate and audit what enters. OpenShift formalises the re-route with - image mirror rules, and carries a sharp edge: digest-mirror rules rewrite digest pulls - only, so a by-tag pull silently skips them - unless you also add an - ImageTagMirrorSet, which is the rule type built for tag pulls. Pinning by digest is + image mirror rules, and carries a sharp edge: digest-mirror rules rewrite digest pulls + only, so a by-tag pull silently skips them - unless you also add an + ImageTagMirrorSet, which is the rule type built for tag pulls. Pinning by digest is still the habit that makes the digest rules catch everything.
-At real fleet scale the mirror itself tiers: a central mirror in the cloud fronts +
At real fleet scale the mirror itself tiers: a central mirror in the cloud fronts upstream, and each site's mirror pulls from the centre rather than from upstream directly. A new image ripples outward in layers - upstream to the centre, centre to each site as it asks, - site to its nodes over the LAN - rather than every site hitting upstream at the same moment. + site to its nodes over the LAN, rather than every site hitting upstream at the same moment. With a site configured to consume only its local mirror, its node image traffic stays on the local network.
Tiered mirrors fan the load out in layers rather than all at once.
Every rebuild is a new artefact - in practice a different digest (reproducible builds are the deliberate exception), untested by the stages before it. - So build once, then promote the same digest through environments by copying, registry to + So build once, then promote the same digest through environments by copying, registry to registry - dev proves the exact bytes prod will run. For a multi-architecture image, skopeo copy --all copies the complete image list; add --preserve-digests when promotion requires the destination to keep @@ -600,15 +599,15 @@
Helm templates are Go text/template: {{ .Values.device.address }} is a pipeline walking a values object. _helpers.tpl commonly holds - named templates - reusable partials that manifests can include - - rather than ordinary functions invoked directly. You are not writing YAML; you are writing a + named templates - reusable partials that manifests can include + , rather than ordinary functions invoked directly. You are not writing YAML; you are writing a program whose output is YAML, so render locally, read the output, and lint what came out rather than what went in.
Review the render, not just the template.
Field note. Argo CD renders charts with helm template rather than running helm install, so the lifecycle differs from Helm's own. lookup - comes back empty - there is no live cluster at render time. Hooks are not dead, though: Argo + comes back empty, because there is no live cluster at render time. Hooks are not dead, though: Argo maps Helm hooks onto its sync phases (pre-install and pre-upgrade become PreSync, post-install and post-upgrade become PostSync), while a few - rollback and test hooks - have no equivalent at all. Render the way your deployer renders, and check where your hooks actually land.
@@ -623,7 +622,7 @@A template that renders whatever it is given just moves the failure downstream. A production-grade chart carries a values.schema.json: a context missing a required - value fails at render time, in the pipeline, with a message naming the field - not months + value fails at render time, in the pipeline, with a message naming the field, not months later as enforcement pointed at nothing.
Field note. The failure you want is the render that refuses. It costs a red pipeline. The alternative reports healthy the whole time.
@@ -644,7 +643,7 @@Appendix · The dependency ledger
The arc reads like a closed machine: repo to reconciler to registry to node. It - is not closed. Three load-bearing pieces live outside the cluster - and the honest move is to + is not closed. Three load-bearing pieces live outside the cluster, and the honest move is to write down what leans on them, and what happens when they are down.