Dossier / 005 · operations

The Mirror

Outcome — a pull-through OCI registry taken apart: why this lab runs its own mirror, what the deployed version actually guarantees, the retention trap that quietly evicts digest-pinned content, and how each lesson translates into production platform engineering. Claims are labelled: observed here, documented upstream, or not yet verified.

Registry
zot v2.1.17 (digest-pinned image; latest upstream release at last check: v2.1.20)
Platform
Talos Linux v1.13.4, containerd 2.2.4, single-node Kubernetes
Upstreams
5 (Docker Hub, GHCR, Quay, registry.k8s.io, NGC), all on-demand pull-through
Storage
50Gi local PVC; dedupe on, GC on (1h delay / 24h interval); no retention config
Auth state
anonymous read enabled; per-node pull credentials staged, not yet enforced
Fallback
origin fallback ON (platform default; deliberately not disabled)
Validated
2026-08-25 (live config read + upstream doc/source verification)
Open items
digest-entry retention (needs v2.1.19+), auth canary negative test, upstream namespace prefixes, disconnected-serve drill

The build that "hung"

An armoured way-station machine exploded into parts: hexagonal hull, roof plate, a shelf of glowing bricks, intake pods and an output nozzle - the mirror as a physical machine.
The mirror as a machine: upstream intakes, one cache shelf, one serving nozzle.

Nobody sets out to run a registry. You inherit the need the day a CI build stalls for no reason you can see.

This site's own build pulls one base image from Docker Hub, anonymously. One evening the build sat there - no visible error, no progress. The accurate version of that story: Docker Hub does not go silent when it limits you. It answers with an explicit HTTP 429 and the error code toomanyrequests; under the current published limits, anonymous pulls get 100 manifest requests per six hours per IPv4 address, and the accounting is per manifest GET. The silence was manufactured on my side, by retries swallowing that answer.

The swallower, recovered from the build script's git history, was a double retry stack: an outer shell wrapper allowing four attempts, wrapped around buildah build --retry 3. Buildah's flag counts retries, not attempts - one initial try plus up to three retries, and it applies to registry push/pull operations. Four outer attempts, each containing up to four inner attempts: up to sixteen requests for one consistently failing pull, each one billed against the same rate limit it was trying to outlast. Retry amplification turns a throttle into an outage and hides the evidence.

Receipt · the retry stack, as committed to the repo
retry() {
  local n=0 max=4
  until "$@"; do
    n=$((n+1)); [ "$n" -ge "$max" ] && { echo ">> failed after ${max} attempts" >&2; return 1; }
    echo ">> attempt ${n} failed, retrying in $((n*8))s..." >&2; sleep $((n*8))
  done
}
retry buildah build --retry 3 --retry-delay 5s ...
# --retry 3 = one attempt + three retries; the wrapper multiplies that by four.
# A 429 answered every one of them, and the backoffs read as a hang.

The immediate fix was to pin the base image by digest - which stops tag re-resolution, but does not remove the network round-trip: an ephemeral runner with a cold local store still fetches the manifest behind that digest every build. The durable answer is structural:

SAY ITWhy does a fleet of machines ask the public internet for the same bytes, hundreds of times, forever?

A production platform answers with a mirror: one registry that fetches once and serves the estate. The lab built one. This page takes it apart - including the parts that turned out not to work the way I first believed.

FIELD NOTE
The mirror did not fix the red X in CI. That red was a separate upstream bug in the CI system's log-finalise step - cosmetic, tolerated, documented. Two problems, one symptom. Diagnose them separately or you will fix the wrong one and declare victory.

What the mirror actually does

The mirror is zot: a single OCI registry in its own namespace, one 50Gi cache volume, running the same digest-pinned deployment discipline as everything it serves. It fronts five public registries in on-demand pull-through mode: a miss fetches from upstream and caches; a hit serves from the shelf. zot also supports polled mirroring (periodic full sync of matching content) and pre-seeding; Docker Hub is on-demand only - upstream documentation is explicit that polled mirroring should not be pointed at Hub.

The storage block, from the live config:

"storage": {
  "rootDirectory": "/var/lib/registry",
  "commit": true,
  "dedupe": true,
  "gc": true,
  "gcDelay": "1h",
  "gcInterval": "24h"
}

commit fsyncs writes before acknowledging - crash safety. dedupe hard-links identical blobs, which pays for itself when five upstreams ship the same base layers under different names. And gc with its two timers looks innocuous here; section 06 is about why it is the most consequential block on this page.

Diagram A · system context System context: consumers, the mirror, its storage, five upstreams, and the trust boundaries between them CI and Kubernetes nodes pull from the zot mirror over the LAN. The mirror stores content on a 50Gi volume and syncs on demand from five upstream registries across the internet boundary. Anonymous read is currently allowed inbound; outbound upstream credentials live in one mounted secret. The origin-fallback path from nodes directly to upstreams is shown dashed. CI (buildah)bastion host k8s nodescontainerd + local store zot v2.1.17anon read (today)sync: onDemand x5 50Gi PVCdedupe + GC 1h/24h docker.io ghcr.io quay.io registry.k8s.io nvcr.io origin fallback (dashed = only when the mirror cannot serve) internet boundary - outbound creds in ONE mounted secret LAN - anonymous read today; per-node creds staged

solid = normal pull path · dashed cyan = on-demand sync on miss · dashed magenta = origin fallback

The decision this explains: one cache serves two consumer classes, and the two credential domains (nodes-to-mirror, mirror-to-upstreams) never mix.

Consumers reach the mirror through the platform's machine-level registry config:

machine:
  registries:
    mirrors:
      docker.io:
        endpoints:
          - https://zot.bztmon.org
      ghcr.io:
        endpoints:
          - https://zot.bztmon.org

The mirror is the only listed endpoint, and that is still not a hard dependency, because Talos documents an implicit final fallback: endpoints are tried in order, "and by default the last implicit endpoint is the original upstream registry", unless skipFallback: true says otherwise. The lab's first version listed the upstream as an explicit second endpoint; the refinement that removed it came from reading the behaviour properly - the default already guaranteed it. Know which of your safety nets you built, and which are defaults you merely have not broken.

One asymmetry decides the whole rollout order in section 07: registry auth in the machine config is documented by Talos as requiring a reboot before the CRI picks it up. Mirror endpoint changes applied live in this lab (observed on Talos v1.13.4 / containerd 2.2.4; the docs are silent on this half, so treat it as an observation, not a guarantee). Credentials the registry consumes from a mounted secret rotate with a pod restart. Put each credential on the side that can move.

THE HOMELAB CLAUSE
One pod, one PVC, one node is an accepted failure domain, not a pattern. Section 09 covers what changes when the requirements change - and what genuinely does not.

Request routing and five upstreams

Five registries feed one endpoint, which raises a question the first version of this page skated past: when a node asks the mirror for pause:3.10, how does the mirror know whether that means Docker Hub, GHCR, Quay, registry.k8s.io or NGC?

Three facts, all verified:

  1. containerd tells the mirror where the request came from - a mirror request carries the original registry as a query parameter: /v2/pause/manifests/3.10?ns=registry.k8s.io. That is documented containerd behaviour.
  2. zot v2.1.17 ignores it. There is no handling of the ns parameter in the deployed version's request path, and using it for upstream selection is an open upstream feature request (zot issue 4187) with, at last check, no maintainer response. The URL path alone selects the local repository.
  3. This deployment has no per-upstream prefixes. The live sync config gives all five upstreams prefix: "**" and no destination - one flat namespace. On a miss, zot tries the configured upstreams in order until one has the path. Upstream documentation's own multi-registry example instead gives each upstream a distinct destination so the requested path selects the origin.
Upstream map - observed configuration
OriginLocal namespaceOutbound authSync modeDigest preservedOn missStatus
docker.ioflat (no prefix - collision-ambiguous by construction)Hub login (rate-limit lift)onDemandyesupstreams tried in config order; first that resolves the path winsobserved
ghcr.iononeonDemandyesobserved
quay.iononeonDemandyesobserved
registry.k8s.iononeonDemandyesobserved
nvcr.io$oauthtoken + keyonDemandyesobserved

Why has the flat namespace not bitten? Because the five origins use largely disjoint path conventions - Hub's official images live under library/ (the implicit prefix behind bare names like alpine), NGC content sits under nvidia/, registry.k8s.io has its own layout. "Largely disjoint" is a probability, not a guarantee: an organisation name that exists on both GHCR and Quay would collide silently, and the winner would be config order. That is an accepted risk in a lab; it is not a design.

PROPOSED - NOT YET DEPLOYED
The safer shape, verified against the documentation of both halves but not yet applied here: give each upstream a distinct destination prefix in zot (/docker, /ghcr, ...), and point each Talos mirror at the prefixed path with overridePath: true (which stops the automatic /v2 append so the prefix survives). Then the requested path itself names the origin, and collisions become impossible rather than improbable. The migration cost: every cached repository changes its local path, so the cache re-warms.
Diagram B · request flow Request flow from a node through containerd to the mirror, with hit, miss, revalidation and fallback branches A pull begins at the node's local content store. On local miss, containerd asks the mirror. A cached digest request is served from mirror storage without upstream contact. A tag request is revalidated against the upstream even when cached, on the deployed version. A storage miss triggers on-demand sync from the matching upstream. If the mirror cannot serve, containerd falls back to the origin registry. node content storehit = no network at all containerd -> mirror digest request, cachedserved locally (verified) tag request, cachedstill revalidates upstream storage misson-demand sync + cache origin registryrate limits live here local miss containerd fallback: mirror endpoints exhausted -> origin direct

green = fully local · amber = the surprising branch on v2.1.17 · magenta = fallback

The amber branch is section 04's punchline: a warm cache does not mean the network is out of the story.

Bootstrap, fallback and failure modes

Six machine parts arranged in a ring around an empty centre; one of them is a quarter-scale replica of the large way-station machine, representing the mirror's own image passing through the mirror.
The dependency ring: one of the orbiting parts is the mirror itself - its own image is served through the thing it is.

Every infrastructure service eventually meets the question: what do you depend on, and what happens when you ARE the dependency? The mirror has four such loops, and each got a different answer.

Loop 1 - the mirror's own image comes through the mirror

zot runs as a container whose image lives on a registry zot mirrors. When the hosting node boots, it asks the mirror - which is not running, because the node is booting. The loop breaks on the documented fallback: mirror endpoints exhaust, containerd falls through to the origin. The design decision is restraint - not setting skipFallback: true.

FIELD NOTE · OBSERVED
Proven by accident: a config change shipped with preserveDigest enabled but without its mandatory partner http.compat - a pairing the registry refuses to start without (the documentation is explicit; so was the crashloop). The mirror went down; the fleet quietly fell through to upstream and nothing user-visible broke. An unplanned failover exercise, passed. The standing rule it bought: validate the config with the registry's own verify in a throwaway pod before merging, every time.

Loop 2 - the recovery tooling deliberately ignores the mirror

The fleet's rescue tooling could pull its execution image through the mirror like everything else. It does not: its image is cached on the operations host, outside the cluster. A recovery tool that depends on the thing it recovers is not a recovery tool.

Loop 3 - authentication cannot flip everywhere at once

Covered properly in section 07; the loop shape: enforcing auth requires every node to carry credentials that only apply after a reboot, so anonymous read must survive until the last node is proven, and the proof itself must be designed not to lie.

Loop 4 - the mirror's host is also the mirror's customer

The node hosting the mirror boots its own workloads through it - including the tunnel that serves this very page. "Restart the mirror's node" therefore carries a blast radius far beyond the mirror, and the runbook for that reboot lists every public-facing thing riding on it and the order they return. Draw the dependency arrows for your own estate; the ones that surprise you are the ones that will page you.

Diagram D · two bootstrap worlds Connected fail-open bootstrap versus disconnected preseeded bootstrap, as separate paths Left: the connected lab path - node boots, mirror miss, implicit fallback to origin, mirror comes up afterwards. Right: the disconnected path - no origin exists; the registry image and release set must be preseeded onto the host or imported from disk before anything else can start. CONNECTED LAB - FAIL-OPEN (observed) node boots, asks mirror: down implicit fallback -> origin serves boot images mirror starts (its image came via fallback) estate converges back onto the mirror DISCONNECTED SITE - PRESEEDED (design, not deployed here) no origin exists - fallback is not a plan registry image preseeded / imported from disk release set, trust roots, creds staged locally site serves itself; syncs on a schedule

left = this lab, observed · right = the airgap translation, a design requirement not a deployed claim

Loop 1's trick assumes an upstream exists to fall through to. The airgap deletes that assumption, and forces Loop 2's discipline onto the mirror itself.

Fallback is a policy decision, not a universal win. This lab keeps it on: better bootstrap survivability, availability through mirror outages, and the cost is real - pulls can silently bypass the mirror (losing cache benefit and any future policy point) and land on origin rate limits. A regulated or disconnected estate makes the opposite call: skipFallback: true, egress restrictions, preseeded release stock, and a tested recovery path, because silent bypass is the failure mode there, not the safety net.

And "fails fast" is not one behaviour. The failure modes differ in symptom, path and consequence:

FailureUser-visible symptomRequest pathFallback?DetectionSecurity consequence
Mirror pod downnone (pulls slower)node -> originyes, immediatemirror probes red; origin egress risespolicy/audit bypass while down
Mirror up, sync wedgedpulls hang up to sync timeoutnode -> mirror (blocked)only after timeoutmanifest probe stalls while /v2/ still answersavailability, not integrity
Upstream 429miss/tag pulls fail or crawlmirror -> origin refusedfallback hits the same limitzot logs; retry storms amplifynone; self-inflicted DoS via retries
Upstream down, digest cachednonemirror serves locallynot needed-none (verified behaviour)
Upstream down, tag cachedfails/stalls on v2.1.17mirror insists on revalidatingorigin also downthe test matrix in section 08availability surprise - the trap of assuming "cached = offline-safe"
PVC fullnew pulls fail, cached OK-ishmirror 5xx on writesyes for missing contentcapacity metrics (proposed)availability; GC pressure
Digest entry GC-evictedsilent re-fetch on next pullmirror -> origin re-syncn/aupstream egress for "cached" contentrate-limit exposure returns (section 06)
Node auth wrong (post-flip)ImagePullBackOffmirror 401; fallback unreliable on 401unreliablethe canary gate (section 07)the outage the gate exists to prevent

Digests, manifests and the supply chain

Precision matters here, because half the traps on this page come from conflating these objects:

For a mirror, digests carry one sharp operational rule, learned here as a crashloop: zot converts Docker-schema manifests to OCI by default, and conversion changes the digest - silently breaking every digest pin and signature downstream. The pairing that prevents it (preserveDigest: true per upstream, with http.compat: ["docker2s2"]) is mandatory in both directions: the registry refuses to start with one and not the other. This deployment runs both, on all five upstreams - observed in the live config.

Preserving digests does not mean the supply chain travels whole. OCI 1.1 referrers - signatures, SBOMs, attestations attached via the subject field and served by the referrers API - are separate objects with their own discovery path. A pull-through cache that mirrors manifests and blobs does not automatically carry them. Nothing here verifies signatures today; if it did, disconnected verification would also need the trust material carried locally (a public key, or a trusted-root bundle for a private signing stack). That is stated as a boundary, not an aspiration.

THE HOMELAB CLAUSE
Production estates promote releases by digest between registries - a tag is a suggestion, a digest is a fact - and decide explicitly whether referrers travel with them. The lab runs the digest half of that discipline; the referrer half is future work and labelled as such.

Retention, GC and the digest-only trap

A stuck intake pod ringed in magenta with three glowing cargo bricks queued behind it - cached content waiting on a jammed process.
Cached bricks queued behind a jammed intake: availability problems here are quiet, not loud.

This section corrects the largest error in this page's own first edition. I wrote: "no retention config means keep everything". That is true for tags, and dangerously false for everything else - and "everything else" includes precisely the content this mirror exists to hold.

The verified semantics, from upstream documentation and the deployed version's source:

SAY ITA digest-pinned fleet, pulling through a mirror whose GC treats digest-only entries as garbage: the cache evicts exactly what the estate is built on.

Put the pieces together with section 04's revalidation fact and the deployed behaviour is this: tag pulls are cached but always phone upstream anyway; digest pulls are served locally but their cache entries are eligible for eviction within hours. On v2.1.17, this mirror is a rate-limit shield and a latency win - it is not yet a disconnection shelf, and the first edition of this page was wrong to imply otherwise.

Diagram C · what keeps an object alive The OCI object graph and which references protect content from garbage collection A tag points to an index; the index references platform manifests; manifests reference layers and a config blob. Referrers attach to a manifest via a subject field. A separate digest-only manifest sits with no tag pointing at it; it is eligible for garbage collection after the delay on the deployed version. Retention rules that keep tags do not protect the untagged manifest. tag: v1.2 index (multi-arch)own digest manifest amd64 manifest arm64 layers + config layers + config referrer (sig/SBOM)subject -> manifest a tag is a keep-alive digest-only cached manifestNO tag references it GC after gcDelay (1h here)deleteUntagged default: true keepTags rules protect TAGS in their repo. Nothing on v2.1.17/18 protects the amber box; pull-aware keepUntagged ships in v2.1.19 (2026-08-04). Node-local caches can mask the eviction for days.

green = protected by a tag · amber/red = the eviction path · dashed magenta = referrers (separate lifecycle)

The lesson: reachability, not existence, is what GC respects - and a pull-through cache full of digest pulls is a graveyard of unreachable-by-tag objects.

Retention policy semantics, stated precisely because two nearby systems use opposite rules:

Mitigations, honestly ranked: upgrade to v2.1.19+ and configure keepUntagged with pull-activity rules (the designed fix); until then, widen gcDelay/retention.delay so eviction pressure drops (the upstream maintainer's own interim suggestion), or set deleteUntagged: false and accept that the cache only grows - a capacity trade, not a free lunch. Disabling GC entirely trades eviction for guaranteed storage exhaustion with no reclaim path; it is listed here to be argued against.

SAFE VALIDATION - PROPOSED, NOT YET RUN
The eviction claim above is documented upstream and consistent with this config; it has not been reproduced in this lab yet. The safe reproduction, for a disposable zot instance only: record version + sanitised config; pull an image by digest through it; confirm on the registry side that the stored manifest is untagged; shorten GC timers (in the disposable instance only); observe the manifest before and after the sweep; then block upstream and repeat the pull from a clean runtime store, recording whether the mirror serves or re-fetches. Two warnings from upstream documentation: the retention verification tool executes orphan-blob GC for real even in dry-run, and on local storage it requires the registry stopped. Never point retention experiments at live storage.

The authentication migration

Four parts left to right: a waiting keyed cartridge, a key wedge, a tall gate frame, and a cartridge beyond the gate with its keyway lit - per-node credentials proven at a checkpoint.
The proving gate: a node counts as migrated when an authenticated pull succeeds where an anonymous one cannot.

Target state: the mirror refuses anonymous pulls. Current state: anonymous read on, with per-node credentials staged in machine configs (inert until each node's reboot - the documented behaviour). The migration is a sequencing problem with one trap at its centre:

While anonymous read is on, an ordinary pull proves nothing about auth. A node whose credentials never applied sends no authorisation header, gets served as an anonymous reader, and false-passes the gate. The canary must live in a repository that denies anonymous read, so a successful pull can only mean an authenticated pull.

That design is now upstream-validated rather than assumed: zot authorisation resolves per-repository policies by longest match - the most specific path pattern wins, and ** is explicitly the default policy for anything unmatched. A canary/** entry that grants named identities read and carries no anonymousPolicy therefore denies anonymous on that path while the glob keeps the rest of the shelf open. Maintainer guidance confirms anonymous and authenticated access are evaluated independently. The staged policy:

"accessControl": {
  "repositories": {
    "**": {
      "anonymousPolicy": ["read"],
      "policies": [
        { "users": ["zot-push"], "actions": ["read", "create", "update"] },
        { "users": ["zot-pull"], "actions": ["read"] }
      ]
    },
    "canary/**": {
      "policies": [
        { "users": ["zot-pull", "zot-push"], "actions": ["read"] }
      ]
    }
  }
}
CLIENT CAVEAT · DOCUMENTED
Mixed anonymous/authenticated policies trigger a Docker-client-specific workaround present in this exact version: /v2/ returns 401 to Docker user agents so the Docker CLI sends credentials, meaning anonymous docker users must log in even for anonymous repos. Podman and containerd are unaffected. This estate pulls with containerd, podman and buildah, so the caveat is noted rather than felt.

The gate, per node, and what each step proves:

  1. Stage credentials in the node's machine config (inert; documented as requiring reboot).
  2. Reboot the node at a planned window.
  3. Prove identity: pull the protected canary through the node's own runtime. Success = authenticated (anonymous cannot); failure = 401, back to step 1. The expected split, to be confirmed against observed responses when the test runs: 401 for missing/wrong credentials, 403 for a valid identity lacking the action - upstream docs state the 403 case explicitly only for OIDC identities, so the basic-auth 403 boundary is listed in the open-verification register rather than asserted.
  4. Prove interception separately: the canary pull is a direct reference to the mirror, so it cannot prove that docker.io/... references are being routed through the mirror at all. That second proof needs an original-upstream reference pulled on the node plus the mirror's request log showing it arrive - and a cold runtime store, because node-local content will satisfy the pull without any network and fake a pass.
  5. Only when every node passes both proofs does anonymousPolicy come off the glob. Rollback is the previous config commit, and it is named before the flip, not after.
Diagram E · the migration ladder The authentication migration ladder with its per-node proof gate and rollback point Five stages: anonymous baseline, credentials staged inert, node reboot, the two-part proof - authenticated canary pull plus logged mirror interception - and the final anonymous-off flip with a named rollback commit. A failing node loops from the proof back to staging. anon read ONbaseline (today) creds stagedinert in config node rebootsauth goes live two-part proofcanary + loggedinterception anon OFFrollback = prior commit 401: node loops back - it does not count

the flip is per-fleet; the proof is per-node - one unproven node under a fleet-wide flip is an outage wearing a green tick

Why two proofs: identity and interception fail independently, and each has a false-pass mode the other cannot detect.

Behind the gate sit five separate trust domains, deliberately not shared: node pull credentials (machine config, reboot-bound), CI push credentials (human-held login), the mirror's own upstream credentials (one mounted secret, pod-restart-bound), TLS trust (estate wildcard, standard roots), and human admin access (SSO in front of the UI). The push credential carries this page's oldest scar: its hash and plaintext once drifted apart, rotation became impossible, and pushes stayed dead for twelve days until the rebuilt flow landed - three applications shipped the same day it came back. Hash and plaintext now live side by side, rotated together by a human, and the automation identity that syncs secrets is read-only by design so it can never half-rotate them again.

Observability that means something

"The registry is up" is the least useful sentence in this page. The worst incident in this mirror's life: an in-flight sync wedged, every pull from that upstream hung, and the health endpoint returned 200 throughout. The log line was almost poetic:

Replay · observed 2026-07-07 · the stalled sync
"image already demanded, waiting on channel"
$ kubectl -n zot rollout restart deploy zot   <- recovery action (not a root cause)
# probe an actual manifest, not the process's opinion of itself:
$ curl -sfS --max-time 10 -o /dev/null -w '%{http_code}\n' \
    https://zot.bztmon.org/v2/library/busybox/manifests/latest \
    -H 'Accept: application/vnd.oci.image.index.v1+json'
200                                           <- THIS is "the mirror works"

Honesty about that incident, upgraded by reading the source: the log line is coalescing by design - concurrent requests for one image join the first sync rather than duplicating it, and since the deployed version syncs run on a detached background context with a three-hour default timeout, surviving client disconnects. Waiters blocking on a genuinely stalled sync until that timeout is consistent with what we saw; the restart was a recovery action, and the root cause was never isolated. Two knobs this config does not currently set - sync maxRetries (disabled by default upstream) and a tighter syncTimeout - are the first things a recurrence should change.

What this deployment actually exposes, verified against the running version's source: /livez, /readyz and /startupz exist as real health endpoints (absent from the docs at this version - a documentation gap, not an invention), alongside the spec's /v2/. The Prometheus metrics extension exists upstream (real series names include zot_http_requests_total, zot_repo_storage_bytes, zot_repo_downloads_total, zot_storage_lock_latency_seconds) - and is not enabled in this deployment. Wiring it, with scrape auth, is on the open list; no dashboard here pretends otherwise.

The layered checks that would actually mean something, each answering one question:

CheckQuestion it answersStatus here
/livez / /readyzis the process up / initialisedavailable
authenticated manifest GET of a local-only canarycan the right identity read real content from local storageproposed
anonymous GET of the canary expecting 401is the protection actually protectingproposed
original-reference pull on a cold node + mirror log lineis the runtime actually routed through the mirrorproposed
digest pull with upstream blocked (disposable env)does "cached" mean "served locally"open - section 06
PVC usage + growth, GC activity, sync latencywhen does capacity or eviction become the storyneeds metrics ext
certificate expiry, credential agewhat breaks on a scheduleproposed

The test matrix for the caching claims - each cell is an experiment, not an assumption (status: the two digest rows are documented upstream and consistent with observed behaviour here; the disconnected rows have not been run in this lab):

ScenarioExpected on v2.1.17What proves it
cached tag, upstream reachableserved; upstream contacted anyway (revalidation)zot log shows upstream request; latency includes round-trip
cached tag, upstream unreachablefails or stalls until timeout, then local fallback pathpull timing vs sync timeout; error text
cached digest, upstream reachableserved locally, no upstream contactabsence of upstream request in zot logs during the pull
cached digest, upstream unreachableserved locally - IF the entry survived GCthe section-06 disposable-instance drill
cold node, warm mirrormirror serves; node store fillsmirror log + no origin egress
warm node, evicted mirror entrypull succeeds from node store - masking the evictionthis is the false-pass: only mirror-side inspection reveals it
multi-arch index + child manifestsindex and per-platform manifests are separate cache entriesper-digest existence checks on the mirror

The production translation

A monolithic vault with a wall of glowing bricks visible through its door, attended by four smaller self-sufficient way-stations - a central release source and site-local mirrors.
A central curated source and site-local registries: each site holds its own shelf, so a cut WAN idles nothing.

The first edition of this section asserted what "production" does. That was the wrong register: production chooses from patterns against requirements. The honest version:

Whatever the topology: recovery time and recovery point get numbers before an incident provides them; backups are restore-tested copies off the failure domain (array snapshots and RAID protect against disks, not against the array, the site, or an operator error - they are inputs to a backup strategy, not the strategy); capacity is planned against the retention policy from section 06, because "how big does the cache get" is a policy output, not a guess; and the circular dependencies from section 04 are drawn per site, because every one of them exists at every scale - the only thing that changes is how expensive they are to ignore.

Verified, open, and where the claims come from

A wide flat node slab in an otherwise empty void; a single luminous brick hovers midway in its descent toward the slab - one small transfer of bytes as proof of life.
Proof of life is one authenticated byte transfer through the real path - not a status page.

When a node powers on, this fleet's gate makes it pull one known, first-party image through its own runtime before it counts as a member. That gate is the whole page in miniature: prove the path, not the process.

Architecture decisions
DecisionChosenWhyTrade-offRollback / alternativeStatus
Origin fallbackON (default kept)bootstrap survivability, Loop 1silent mirror bypass possibleskipFallback:true + preseed (airgap shape)observed
Sync modeonDemand, all upstreamscache follows real usage; Hub-safetag pulls revalidate upstream; digest entries untaggedpolled sync for a curated release setobserved
Digest preservationpreserveDigest + docker2s2, x5digest pins + signatures survive mirroringmandatory config pairing (crashloop scar)none - non-negotiable for digest-pinned estatesobserved
Namespace layoutflat (no destinations)simplicity at build timeordered-trial upstream selection; theoretical collisionsper-origin destinations + overridePathredesign proposed
Retentionnone configuredpredates understanding the untagged ruledigest-only entries evict within hoursv2.1.19+ keepUntagged; interim wider gcDelaycorrection owed
Authanonymous read until per-node proofcontainerd 401-fallback unreliability makes big-bang flips dangerouswindow with no pull authstaged flip w/ canary gate + named rollbackin flight
Metricsnot enabledminimal first deploymentcapacity/eviction invisiblemetrics extension + scrape authproposed

Open verification register

  1. Digest-entry eviction reproduction in a disposable instance (section 06 procedure).
  2. Canary repo negative test: anonymous 401 on canary/** while anonymous read elsewhere still succeeds; observed 401-vs-403 boundaries recorded.
  3. Mirror interception proof: original-reference pull on a cold node correlated with the mirror's request log.
  4. Disconnected-serve drill: cached digest pull with upstream blocked, disposable environment first.
  5. Upgrade evaluation: v2.1.19/v2.1.20 for keepUntagged, with the retention config written and reviewed before the upgrade, not after.
  6. Namespace redesign migration plan (per-origin destinations + overridePath), including cache re-warm cost.

Verified against