The Exploded Cluster · The Delivery Arc

The machine was the easy part.
Now watch how software reaches it.

The foundations first - what an image is, what a cluster is - then the toolchain that delivers to them: how change travels, how images are named, and how one definition serves a fleet. Scroll, and each machine comes apart.

What this teaches

The Exploded Cluster teaches how modern container platforms work by taking them apart, literally. Each course is one machine drawn as a single exploded illustration, sliced into its real components and wired to your scroll, so the architecture moves while the words explain it. OpenShift is the teaching lens. oc is the command language. Kubernetes primitives explain the machinery underneath. Occasional lab notes show a separate Kubernetes environment where a concept was observed; they are not a claim that the lab runs OpenShift. Start with the aperitif's four terminal commands - three that deliver, one that checks - and finish knowing how a change travels from a Git commit to a running, secret-fed, digest-pinned workload on a fleet. The library at the end links only to official documentation, so every claim here can be checked against its source.

Course 00 · Aperitif

Three delivery actions. One check.
What actually just happened?

The whole ceremony of shipping software fits in four lines of terminal. They work on your first day and stay mysterious for years. Scroll - the shell comes off first.

Six dark armour fragments with neon seams framing a large empty centre - the
                  casing of a machine caught the instant before it comes apart.

A note on the commands: this site uses oc, OpenShift's CLI. Examples use oc, the OpenShift CLI, against Kubernetes API resources - for those, kubectl behaves identically. Where a step needs OpenShift APIs (SCCs, Routes, ClusterOperators, MachineConfig), it will not exist on a plain Kubernetes cluster, and the text says so.

That 1/1 reads as containers-ready over containers-wanted: a pod can hold more than one, which is Course IIIb.

Three of those were delivery actions - build, push, apply, and the last, oc get pods, only checked the result. Built of what, exactly? Pushed to where, and what travelled? Applied, which is not the same as launched. Running, according to whom? Each course below takes one of those words apart.

Course I · Podman

An image is not a box.
It is a stack of frozen diffs.

Scroll, and the thing you keep calling "a container image" comes apart in your hands. Four layers. Each one only stores what changed from the layer under it, and here they rise one at a time, bottom up.

An exploded isometric view of a container image: four stacked slabs floating
                  apart - a heavy metal base, a circuit-etched dependency layer, a magenta-traced
                  code layer, and a thin frosted-glass writable layer on top.

    A container is a process, not a machine

    If you arrived here from Linux, take this translation first: a running container is an ordinary process on your kernel. No guest OS, no hypervisor. The kernel gives it namespaces so it sees its own PID tree, mounts, network and hostname, and cgroups so its CPU and memory can be capped. ps on the host lists it. kill on the host kills it. Podman leans into that: no daemon sits in the middle - the container is a child of your own shell, and rootless mode maps your user onto root inside the container through a user namespace, so root in there is an unprivileged UID out here.

    So what does the image provide? The filesystem that process sees. That is the whole job, and it is why an image is a stack of layers rather than a disk image.

    Field note. Prove it on your own box. podman run -d --name t alpine sleep 300, then on the host: ps -ef | grep sleep finds the process, lsns -p <pid> lists the namespaces it was handed, cat /proc/<pid>/cgroup shows where its limits live, and mount | grep overlay shows the layers stitched together. None of it is exotic. It is your kernel, described differently.

    What an image is made of

    An image is not a copy of a machine. It is a stack of read-only layers, each recording only what changed from the one beneath. Layers are content-addressed, so an identical layer is stored once and reused by every image that references it, so a pull fetches just the layers you do not already have. A base sits at the bottom, your dependencies on it, your code - usually the smallest layer, and typically the most volatile - above that.

    Those layers become one filesystem through a union mount - overlayfs, the same kernel feature you can mount by hand. The read-only image layers are the lower dirs; the container gets a fresh upper dir of its own. Writes land in the upper, and editing an existing file copies it up there first, leaving the image layer untouched underneath.

    That upper dir is the top slab in the scene, and it is the odd one out: the writable layer is not part of the image. The runtime creates it with the container and discards it when that container dies. Nothing written there ships, and nothing written there survives, which is the entire reason volumes exist.

    Change a layer and every layer above it must be rebuilt.

    The order is a caching decision

    The builder caches layer by layer, and a cached layer survives only while everything beneath it is unchanged. Put COPY . . above your dependency install and you have told the builder to throw the dependency cache away every time one line of code changes. Dependencies first, code last - a Containerfile is a cache policy that happens to build software.

    Field note. A rebuild that takes twenty minutes and one that takes twenty seconds are usually the same Containerfile with the lines swapped.

    This stack explains a lot of what sits above it: shared layers make pulls fast, immutable content lets a digest name exact bytes (Course VII), and the throwaway top layer is why persistent state needs volumes.

    Pre-reads: none - start here  Further: Podman get-started · Podman docs + builds

    Course II · Kubernetes

    A cluster is a promise,
    not a place.

    You never tell Kubernetes how to run your app. You describe what you want - declarative intent, and the cluster works continuously to make it true. Scroll, and the formation splits: the half that decides rises, the halves that run spread below.

    A small fleet mid-explosion: one wide command slab with a strong cyan seam
                  hovering above a row of three identical worker blocks, all floating apart in
                  the void.
    Generic OpenShift/Kubernetes cluster model - not a diagram of the homelab.

      Desired against actual, on a loop

      The habit underneath everything: the reconciliation loop - compare desired state against actual state, fix the difference, repeat. That oc apply didn't launch anything; it filed paperwork. The machine took it from there, and it keeps taking it from there: kill a pod and it returns, not because something noticed the crash but because the loop noticed the difference.

      Kubernetes doesn't run your app - it reconciles it.

      One half decides, one half runs

      The split in the scene is the split that makes everything else possible: a control plane that decides - holds the truth, schedules, reconciles, and worker nodes that run pods. The workers are deliberately interchangeable: identical, replaceable, cattle from day one. Authority does not live where the work happens.

      Field note. If you SSH into a node to "fix" a workload, a controller may reconcile your change away. Change the desired state instead, and let the loop carry it.

      What the control plane is made of

      The core control plane is the API server, scheduler and controller-manager, backed by etcd; cloud deployments may also run a cloud-controller-manager. Naming them makes the later error messages readable. The API server (kube-apiserver) is the only door: requests authenticate, are authorised and are admitted there, and it is the one component that talks to etcd, the key-value store holding Kubernetes API state. Lose etcd and you have lost the cluster, which is why backing it up is the homework to not skip. The scheduler decides which node a new pod belongs on - packing by the resources a pod requests, not by what it currently uses - and records that decision; it does not start containers. The controller-manager runs the reconciliation loops. Control-plane state converges through the API server: the scheduler and built-in controllers watch API objects and write decisions back through the API rather than modifying etcd directly.

      One of those loops is the chain you will debug most: a Deployment creates a ReplicaSet, and the ReplicaSet creates pods. That is why the deploy in the aperitif printed deployment.apps/api created and you then went looking for a pod, and why, when a rollout is stuck with no pod at all, the answer is upstream in that chain rather than on any node.

      etcd is the durable store for Kubernetes API state; controllers reconcile from it continually.

      Everything the delivery arc teaches from Course VI onward is the same loop at a larger scale: Git as the desired state, whole fleets as reconciled objects.

      Pre-reads: C-I  Further: Kubernetes overview · cluster architecture

      Course IIIa · The node

      Where intent
      becomes a process.

      Everything so far was decision. This is the machine where a pod stops being paperwork and starts being a process, and the chain assembles link by link as you scroll.

      A diagonal chain of node machinery: a visor-lit kubelet module, a layered
                  container-runtime engine, a ported CNI ring, a magenta routing prism, and a
                  small glowing pod capsule descending toward the engine.

        Where a pod becomes processes

        The scheduler decides which node a pod belongs on; the kubelet on that node makes the assigned PodSpec real; the runtime creates and starts the containers. The kubelet does not place pods, and does not itself create containers - it speaks CRI (the Container Runtime Interface) to containerd or CRI-O, and that runtime pulls the image (through the mirror of Course VII) and starts it, handing the low-level work to runc or crun. A CNI plugin (Container Network Interface) gives the pod a real IP - called by the runtime, not the kubelet, and kube-proxy, or an eBPF datapath replacing it, makes Service addresses route to real pods.

        The scheduler places the pod. The kubelet realises it. The runtime runs it.

        Field note. Node NotReady? oc describe node shows its conditions and the last heartbeat, and the node's events say what the platform already knows. When you need the node itself, oc debug node/<name> is the supported doorway - not SSH. A node that cannot phone home is presumed lost.

        The control plane does not start container processes directly. It records and reconciles intent; node-side components execute it - the kubelet turning intent into instructions, the runtime turning instructions into processes.

        Pre-reads: C-II  Further: cluster architecture · node components

        Course IIIb · The pod

        One IP,
        shared fate.

        A pod is not a container - it is a jacket around one or more. Scroll, and the capsule opens like a clamshell: shells apart, contents on display.

        A pod capsule blown open: two frosted shell halves floating apart, a cyan app
                  container column standing on an amber-lit init gate, a smaller magenta sidecar
                  beside it, and a stack of translucent volume discs.

          The jacket, not the container

          Everything inside the jacket shares a network namespace: one IP, localhost between friends. Volumes are declared once on the pod, but each container mounts the ones it needs - sharing storage is opt-in, not automatic. regular init containers run in order and complete before the app containers start. Native sidecars are the deliberate exception: restartable init containers that keep running alongside the app, for jobs like proxy, logs and config reload.

          Three probes, three different jobs

          startup owns warm-up, readiness gates traffic, liveness restarts the truly hung. Confusing them is how healthy pods get executed - a slow start killed by an impatient liveness probe looks exactly like a crash.

          containerPort is documentation - unless a Service targets it by name, or you use hostPort. Either way the app still has to bind the port itself.

          Field note. Exit 137 is 128 + 9: the process was SIGKILLed. It does not say by whom. The kernel's OOM killer surfaces as reason OOMKilled; a failed liveness probe shows up in the pod's events; eviction and node pressure look different again. One exit code, several possible crimes - read the termination reason and the events, never the number alone: oc describe pod <name> shows both together, and oc logs --previous shows what the dead container said on its way out. Worth connecting to what you already know: a memory limit becomes a cgroup ceiling, and the kernel's OOM killer enforces it just as it would for any other process on the box.

          The pod is the smallest schedulable unit - the jacket, never the container. Once that distinction lands, half of Kubernetes networking stops being mysterious.

          Pre-reads: C-III's node scene above  Further: pods · the three probes

          Course IV · The traffic

          Pods die constantly.
          The address does not.

          Pods die, respawn and change addresses, and traffic still arrives. Scroll, and the delivery route assembles checkpoint by checkpoint; watch what happens to the pod that stops answering.

          A delivery route in the void: a glowing client orb, a fanned load-balancer
                  wedge, an open ingress doorframe, a Service prism with a bright core, a lit
                  ready pod, and a dark unlit pod fallen out of the line.

            The stable name in front of the churn

            A Service is the fixed point: a ClusterIP inside, a LoadBalancer at the edge, Ingress or the Gateway API doing host- and path-routing above. Clients hold the name; the pods behind it come and go without anyone being told.

            How a Service finds its pods

            A Service holds no list of pods. It holds a label selector (say app: api), and a controller continuously matches that against every pod in the namespace. That is a Kubernetes namespace, a naming boundary for objects, unrelated to the kernel namespaces of Course I; the passing set lands in an EndpointSlice. Labels are how everything here finds everything else, from a Service picking pods to a fleet hub picking whole clusters (Course VI). Wear the label and you are eligible; readiness decides whether you stay.

            A ClusterIP is a virtual Service address implemented by the node dataplane rather than an application process listening on that IP. In iptables or nftables mode, kube-proxy installs rules that steer Service traffic to endpoint IPs; if you have written a DNAT rule by hand, that mode will feel familiar. In IPVS mode, kube-proxy binds Service IPs to the kube-ipvs0 dummy interface and creates IPVS virtual servers. eBPF implementations can replace kube-proxy with their own dataplane. Cluster DNS (CoreDNS) resolves api.myns.svc.cluster.local to the Service address.

            Readiness decides membership

            When a configured readiness probe fails, the pod becomes unready and normal Kubernetes Service traffic stops selecting it for new connections. There is no error at the client - the pod drops out of the endpoint set. That is the feature: broken instances take themselves out of rotation. It is also the first place to look when traffic "disappears".

            The rule has deliberate exceptions: you address pods directly when debugging a specific instance, and headless Services exist precisely so StatefulSet members can be reached individually by stable DNS name. For ordinary application traffic, though, the Service is the only address worth knowing.

            For ordinary Service traffic: address the Service, not a pod.

            Field note. "The network is broken" after a deploy is usually readiness telling the truth about your app, not the network lying about your packets.

            The dark pod in the scene is not an error state - it is the system working. For pods with readiness probes, eligibility for Service traffic is continually re-evaluated for the life of the pod.

            Pre-reads: C-III's pod scene (readiness lives there)  Further: Services and networking

            Course V · OpenShift

            Kubernetes with opinions -
            and a security guard.

            OpenShift is a distribution of Kubernetes: same engine, opinionated chassis. Scroll, and the opinions bloom outward from the core they orbit.

            A glowing geodesic core ringed by six satellites: a magenta admission shield,
                  an open route arch, interlocking operator rings, a handheld console, an
                  amber-lit stack of machine-config plates and a compact single-node box.

              The doorman interviews every pod

              The SCC - Security Context Constraint - is admission deciding what a pod may BE, checked by the api-server when the pod is created and before any node sees it. Under OpenShift's restricted SCCs, workloads normally run as a non-root UID drawn from the project's allocated range, so an image has to work with an arbitrary permitted UID rather than assuming a fixed user. Workloads that genuinely need privilege get a dedicated ServiceAccount bound to the minimum SCC that grants it, not the stock one.

              The platform runs itself

              Routes predate Ingress and still rule here. An Operator is a controller paired with a custom resource: you describe what you want in YAML, and its controller builds it and keeps it true. An Operator applies the same reconciliation pattern to an application or platform capability. The split matters: the platform's own operators are driven by the Cluster Version Operator, while OLM installs and upgrades the add-on Operators you choose from OperatorHub. MachineConfig is the supported declarative path for node OS configuration; ad-hoc SSH changes create drift and are not the intended operating model. And SNO - single-node OpenShift - puts a whole cluster on one box at the edge. At fleet scale the labels from Course VI decide which of these boxes runs what.

              On OpenShift, admission is the interview - the SCC is the dress code.

              Field note. Deployment stuck at 0/1 with no pod at all? The refusal happened above scheduling - read the ReplicaSet events. The error lives a level up.

              Everything in the ring wraps the same core you already know; what OpenShift adds is admission, routing and lifecycle opinions on top of it.

              Pre-reads: C-II  Further: OpenShift documentation · Red Hat OpenShift

              Course VI · GitOps

              Nobody deploys anything.
              The cluster syncs itself.

              The mental model everyone arrives with: someone with credentials pushes manifests at the cluster. In the architecture this course teaches, nothing is pushed: a repository holds the desired state, an agent inside the cluster watches it, and the cluster pulls its own future from Git.

              An exploded chain: an etched repository crystal, a twin-ring reconciler engine,
                  a stack of rendered manifests and a cluster slab - with a drift shard falling
                  away and an armoured secret vault floating deliberately apart.

                The loop you already know, one level up

                C-II taught the reconciliation loop: desired versus actual, fix the difference, repeat. Argo CD applies the same habit to delivery. An Application names a repo, a path and a revision - watch this branch of this repository, and the controller renders what it finds there, compares it against the live cluster, and syncs the difference. The deploy button is a Git commit; the change history is git log; code review is change control. Git records intent - the cluster's own audit log and Argo CD's sync history record what actually happened, including what Git does not see.

                oc reads the running system; Git records the change.

                Pull, not push - the security inversion

                Here the cluster pulls. Push-based delivery exists and is still GitOps to many; the pull variant is worth choosing deliberately, because no CI system, laptop or build pipeline holds a credential that can touch the cluster - the agent inside holds a read-only deploy key and the trust arrow points out. Compromise the build system and you can propose a change, which is visible; you cannot reach into production through Git. It can still push images, which is the other half of why a manifest should name the digest, not the tag. Hand-edit a live object and the controller flags it OutOfSync. With self-heal enabled, Argo CD can restore the desired state; with prune enabled, objects removed from Git can be deleted. Both are opt-in: without them the reconciler reports the drift and waits for a sync. Rollback is git revert, which is why commit hygiene is an operational skill.

                Field note. With self-heal enabled, a live hand-patch may be reverted on the next reconciliation. Without self-heal, Argo CD reports the drift until someone syncs it. Either way, land the change in Git in the same hour, or the fix and the reason for it go missing.

                At fleet scale, the label is the deploy button

                Run many OpenShift clusters under a hub - RHACM, Red Hat's fleet manager, with edge clusters arriving through zero-touch provisioning, and nobody applies apps to clusters by hand. Each app carries a Placement that selects cluster labels; the hub matches placements against the labels a cluster wears, and the chosen cluster is given its assignment - push-style from the hub by default, or in a pull model where each cluster runs its own reconciler. Labelling a cluster changes its placement eligibility; the lifecycle and cleanup of what lands still follow the configured propagation and deletion policy, so removing a label does not, on its own, guarantee an app is deleted.

                Label the cluster; the app follows.

                Keep plaintext secret values out of Git

                In this architecture, plaintext secret values do not live in Git: Git stores the ExternalSecret reference, and the vault stores the value. Commit a plaintext secret and you should treat it as compromised from that moment - deleting it later does not guarantee it is gone from history, forks, clones, CI caches or backups. So Git carries an ExternalSecret naming a logical key; a vault (Azure Key Vault in the worked example) holds the value; an operator inside the cluster exchanges one for the other at runtime. Rotation happens in the vault, not as a commit, which is why the vault sits apart from the pipeline in the scene above.

                Git holds the shape of the secret. The vault holds the secret.

                Delivery becomes a property rather than an event: the cluster converges on what the repository says. With protected branches and attributable identities, changes should be traceable to reviewed commits.

                Pre-reads: C-II · Kubernetes concepts · git + pull requestsFurther: Argo CD · OpenShift GitOps · External Secrets · Azure Key Vault

                Course VII · The image supply chain

                A tag is a promise.
                A digest is a fact.

                myapp:latest feels like a name. It is a sticky note: a mutable reference anyone with push rights can move from one image to another. Registries may audit tag updates, but the tag itself carries no immutability guarantee. The digest is the sha256 of the image manifest (the small JSON index listing an image's layers, unrelated to the YAML manifests you apply to a cluster). It identifies the exact uploaded content. A local image ID printed by a build is not necessarily the same value, so pin the registry digest that consumers actually pull.

                The image journey: a layered image stack, an upstream registry tower, a squat
                  pull-through mirror and a node core - beneath a ghost tag plate and an engraved
                  digest seal floating side by side.

                  Say the name properly

                  Three ways to name an image, in rising order of honesty: :latest (a moving target), :1.4.2 (a label someone maintains, until they re-push it), and name:1.4.2@sha256:... - a fact. The tag stays for human eyes; the digest does the pulling. Pin by digest and "what is running?" has a single answer.

                  Field note. :latest is how two nodes run different code from one manifest - the second node pulled an hour later, after a re-push. Nobody changed the YAML.

                  Why a fleet pulls once

                  Between the build and the node sits the registry chain. Upstream, a managed registry (Azure Container Registry in the worked example) holds what CI built. In front of the cluster sits a mirror: a pull-through cache like zot. On a cache miss the mirror fetches the artefact from upstream; subsequent requests can be served locally while that content stays cached. Rate limits, egress cost, disconnected sites, control - one place to gate and audit what enters. OpenShift formalises the re-route with image mirror rules, and carries a sharp edge: digest-mirror rules rewrite digest pulls only, so a by-tag pull silently skips them - unless you also add an ImageTagMirrorSet, which is the rule type built for tag pulls. Pinning by digest is still the habit that makes the digest rules catch everything.

                  At real fleet scale the mirror itself tiers: a central mirror in the cloud fronts upstream, and each site's mirror pulls from the centre rather than from upstream directly. A new image ripples outward in layers - upstream to the centre, centre to each site as it asks, site to its nodes over the LAN, rather than every site hitting upstream at the same moment. With a site configured to consume only its local mirror, its node image traffic stays on the local network.

                  Tiered mirrors fan the load out in layers rather than all at once.

                  Build once, promote by copy

                  Every rebuild is a new artefact - in practice a different digest (reproducible builds are the deliberate exception), untested by the stages before it. So build once, then promote the same digest through environments by copying, registry to registry - dev proves the exact bytes prod will run. For a multi-architecture image, skopeo copy --all copies the complete image list; add --preserve-digests when promotion requires the destination to keep the source digests, so the copy fails if that cannot be done rather than silently landing a different digest. Human tags ride along; the digest is the through-line.

                  If the digest changed, it is not a promotion - it is a new candidate.

                  Names that can move are convenient right up until they move. Address content by its digest and the supply chain rests on a hash anyone can check, not on trust that a tag stayed put.

                  Pre-reads: C-I · Kubernetes imagesFurther: zot · Azure Container Registry · OpenShift image mirroring · skopeo

                  Course VIII · Helm

                  Stop copying YAML between clusters.
                  Ship the function instead.

                  The default way to run one app on five clusters is five copies of the YAML, and the default result is five slightly different apps. The inversion: stop copying outputs and ship the function. A chart is a template with holes; each cluster supplies one small values file that fills them.

                  The Helm press: an engraved chart plate with empty sockets, four values crystals
                  feeding in, three rendered sheets fanned out in different hues, and a schema gate
                  wedge with a rejected grey sheet stopped behind it.

                    You are writing a program, not YAML

                    Helm templates are Go text/template: {{ .Values.device.address }} is a pipeline walking a values object. _helpers.tpl commonly holds named templates - reusable partials that manifests can include , rather than ordinary functions invoked directly. You are not writing YAML; you are writing a program whose output is YAML, so render locally, read the output, and lint what came out rather than what went in.

                    Review the render, not just the template.

                    Field note. Argo CD renders charts with helm template rather than running helm install, so the lifecycle differs from Helm's own. lookup comes back empty, because there is no live cluster at render time. Hooks are not dead, though: Argo maps Helm hooks onto its sync phases (pre-install and pre-upgrade become PreSync, post-install and post-upgrade become PostSync), while a few - rollback and test hooks - have no equivalent at all. Render the way your deployer renders, and check where your hooks actually land.

                    Contexts: one small file per cluster

                    The chart owns everything structural - resources, probes, security, policy. Each cluster owns one values file: names, addresses, sizes, flags. The context is deliberately values-only; the moment it carries its own manifests there are two owners for one object, and they will disagree. One value can feed several rendered artefacts: the app config, network attachment and two policies can all render from the same field, so those copies cannot diverge - the lived version is on the blog.

                    The chart owns the shape. The context owns the numbers.

                    Make the template refuse

                    A template that renders whatever it is given just moves the failure downstream. A production-grade chart carries a values.schema.json: a context missing a required value fails at render time, in the pipeline, with a message naming the field, not months later as enforcement pointed at nothing.

                    Field note. The failure you want is the render that refuses. It costs a red pipeline. The alternative reports healthy the whole time.

                    Consistency stops being something you police and becomes something the tooling cannot express.

                    Pre-reads: C-II · Kubernetes objectsFurther: Helm docs · chart template guide · Go text/template · Helm on OpenShift

                    Appendix · The dependency ledger

                    Every toolchain stands on
                    services it does not run.

                    The arc reads like a closed machine: repo to reconciler to registry to node. It is not closed. Three load-bearing pieces live outside the cluster, and the honest move is to write down what leans on them, and what happens when they are down.

                    Three familiar machines at rest: the etched repository crystal, the armoured secret
              vault, and the mirror way-station - the supporting cast of the delivery arc.
                    You have met these three before.

                    GitHub

                    Where the desired state lives - the system of record the whole loop watches, through a read-only deploy key.

                    Leans on it: sync, rollback, change review, the "who merged this" answer.

                    When it is down: existing workloads keep running. Repository-backed refreshes and new changes become unavailable; what the controller can still evaluate depends on state already available to it locally. Do not design around an assumed cache lifetime. What stops is change; the running system is unaffected.

                    Azure Key Vault

                    Where the secret values live - git carries the reference, the vault carries the value, an operator keeps them synced.

                    Leans on it: secret sync, rotation, the first deploy of anything that needs a credential.

                    When it is down: already-materialised Kubernetes Secrets typically remain usable while the vault is unavailable. What fails is refresh, rotation and the creation of new secret material - survivable, unless you are inside a rotation window.

                    zot

                    Where the fleet pulls from - a pull-through mirror between the cluster and the internet, and the control point for what enters. At fleet scale it tiers: one central mirror in the cloud fans out to per-site mirrors, layering the load.

                    Leans on it: every image pull on every node - boot, reschedule, scale-up, recovery.

                    When it is down: the sharpest edge. Upstream down + mirror up = nobody notices, provided the image is already cached - a cold entry still needs upstream. Mirror down on a mirror-only pull path: a pod whose image is already in the node's store can still start under IfNotPresent, but imagePullPolicy: Always must resolve the reference against the registry on each launch, so a mirror outage stops it even when the layers are present.

                    None of these outages stop what is already running - they stop change, rotation and recovery. Cache what you pull, keep secret values in a vault rather than Git, and the cluster can hold its last applied state while a dependency is away.

                    Further: GitHub docs · Azure Key Vault · zot · Kubernetes · Helm · Red Hat OpenShift

                    The whole arc · end to end

                    One flow, no gaps.

                    Every course above is one stretch of the same journey. Here is the full run, drawn in the house blueprint style: the change lane, the shape lane, the artefact lane and the secret lane, all converging on one running workload.

                    End-to-end delivery flow. Change lane: a commit lands in the GitHub repository, Argo CD
              renders and diffs it, and syncs the cluster - the cluster pulls rather than being pushed to.
              Shape lane: the Helm chart plus a per-cluster values context passes the schema gate
              and renders the manifests Argo CD applies. Artefact lane: CI builds once, pushes to
              Azure Container Registry, the zot mirror caches it, and the node pulls by digest.
              Secret lane: Azure Key Vault holds the values, the External Secrets operator syncs
              them in - Git holds only the reference. All four lanes converge on the running
              workload.

                    The library

                    Go to the sources.

                    Every technology this site teaches, one sentence each, official documentation only.