Chapter 02 · Dossier 005 · operations
A pull-through OCI registry, taken apart.
What the deployed version guarantees, the retention behaviour that evicts digest-pinned content, and how each lesson translates to production. Every claim carries an evidence label — and the corrections to my own first draft are on the page, not quietly edited out.
Evidence labels used below: lab reproduced here with a dated receipt · config present in the inspected config · source established from the deployed version's source · docs stated in version-appropriate documentation · reported described in an unresolved upstream issue · proposed not deployed · open not yet run
I added zot after a CI pull hit Docker Hub's anonymous rate limit and the retry wrapper hid the 429. At the time of the incident, this site's build-stage pull was anonymous.
Docker Hub does not go silent when it limits you. It answers with HTTP 429 and the error
code toomanyrequests. Under the
current published limits,
unauthenticated users get 100 pulls per six hours per IPv4 address (or IPv6 /64); a
single-platform image counts as one pull, a multi-architecture image counts once per
architecture pulled, version checks do not count, and a HEAD request can read
the rate-limit headers without consuming a pull. A separate abuse limiter covers all
request types with its own 429 form. The silence in this incident was
manufactured on my side, by retries swallowing the answer.
The swallower, recovered from the build script's git history, was a double retry stack:
an outer shell wrapper allowing four attempts, wrapped around
buildah build --retry 3. Buildah's flag counts retries - one initial
try plus up to three more, applying to registry push/pull operations. Four outer attempts,
each containing up to four inner attempts: up to sixteen requests for one
repeatedly failing registry operation (a build with several pulls can produce more), each
billed against the limit it was trying to outlast. Retry amplification turns a throttle
into an outage and hides the evidence while doing it.
retry() {
local n=0 max=4
until "$@"; do
n=$((n+1)); [ "$n" -ge "$max" ] && { echo ">> failed after ${max} attempts" >&2; return 1; }
echo ">> attempt ${n} failed, retrying in $((n*8))s..." >&2; sleep $((n*8))
done
}
retry buildah build --retry 3 --retry-delay 5s ...
# --retry 3 = one attempt + three retries; the wrapper multiplies that by four.
# A 429 answered every one of them, and the backoffs read as a hang.
Pinning the base image by digest stopped the tag re-resolution but not the round-trip: an ephemeral runner with a cold store still fetches the manifest behind that digest every build. The repeated origin pulls justified a shared cache. This page takes it apart, including the parts that turned out not to work the way I first believed.
The mirror is zot: one registry, one 50Gi cache volume, digest-pinned like everything it serves. Five public registries front it in on-demand pull-through mode - a miss fetches and caches, a hit serves from the shelf. zot also offers polled mirroring and pre-seeding; Docker Hub is on-demand only, per its documentation.
The storage block, from the live config:
"storage": {
"rootDirectory": "/var/lib/registry",
"commit": true,
"dedupe": true,
"gc": true,
"gcDelay": "1h",
"gcInterval": "24h"
}
commit asks zot to commit writes to disk immediately instead of relying on
buffered flushing; it narrows the buffered-write window, while end-to-end power-loss
durability still depends on the filesystem, volume and disk. dedupe uses hard
links on local filesystem storage (remote backends implement it differently), saving
capacity when five upstreams ship the same base layers under different names - at the cost
of a startup reconciliation pass that matters operationally when toggled on existing data.
And gc with its two timers looks innocuous here; section 06 is about why it is
the most consequential block on this page.
solid = normal pull path · dashed cyan = on-demand sync on miss · dashed magenta = origin fallback
One cache, two consumer classes; the two credential domains never mix.
Consumers point at it via the platform's machine-level registry config:
machine:
registries:
mirrors:
docker.io:
endpoints:
- https://zot.bztmon.org
ghcr.io:
endpoints:
- https://zot.bztmon.org
The mirror is the only listed endpoint, yet not a hard dependency: Talos tries
endpoints in order "and by default the last implicit endpoint is the original upstream
registry", unless skipFallback: true. An early version listed the upstream
explicitly as endpoint two; it was removed as redundant - origin fallback here is a Talos
default, not an endpoint this project maintains.
One asymmetry decides section 07's rollout order. Talos's v1alpha1 machine-config reference states for registry auth: "changes to the registry auth will not be picked up by the CRI containerd plugin without a reboot" - matching what this lab saw on v1.13.4. Mirror endpoint changes applied live here (same versions; the docs are silent on that half, so it stays a lab observation, not a cross-version guarantee). Secrets the registry consumes from a mount rotate with a pod restart. Put each credential on the side that can move.
One pod, one PVC, one node is an accepted failure domain here, recorded as such; section 09 covers what changes when the requirements do.
Five registries feed one endpoint: when a node asks the mirror for
pause:3.10, how does it know which origin that means? Three facts:
/v2/pause/manifests/3.10?ns=registry.k8s.io (documented).ns handling exists in the
deployed request path; using it for upstream selection is open feature request zot 4187.
The URL path alone selects the local repository.prefix: "**" and no destination. On a
miss, zot tries upstreams in config order until one has the path; the docs' own
multi-registry example instead gives each a distinct destination.| Origin | Local namespace | Outbound auth | Sync mode | Digest preserved | On miss | Status |
|---|---|---|---|---|---|---|
| docker.io | flat (no prefix - collision-ambiguous by construction) | Hub login (rate-limit lift) | onDemand | yes | upstreams tried in config order; first that resolves the path wins | observed |
| ghcr.io | none | onDemand | yes | observed | ||
| quay.io | none | onDemand | yes | observed | ||
| registry.k8s.io | none | onDemand | yes | observed | ||
| nvcr.io | $oauthtoken + key | onDemand | yes | observed |
The flat namespace has not bitten because the origins use largely disjoint paths -
Hub's official images live under library/ (the implicit prefix behind bare
names like alpine), NGC under nvidia/. Largely disjoint is a
probability: an organisation existing on both GHCR and Quay would collide, and config order
would pick the winner. Accepted risk, recorded; not a design.
destination prefix per upstream
in zot (/docker, /ghcr, ...), with each Talos mirror pointed at
the prefixed path via overridePath: true (suppressing the automatic
/v2 append). The requested path then names the origin and collisions become
impossible. Migration cost: every cached repository changes local path, so the cache
re-warms.green = fully local · amber = the surprising branch on v2.1.17 · magenta = fallback
The amber branch: a warm cache does not take the network out of the story.
Four circular dependencies live in this design; each got a different answer.
zot runs as a container whose image lives on a registry zot mirrors. When the hosting
node boots, it asks the mirror - which is not running, because the node is booting. The
loop breaks on the documented fallback: mirror endpoints exhaust, containerd falls through
to the origin. Keeping skipFallback unset is the load-bearing choice.
preserveDigest enabled but without its
mandatory partner http.compat - a pairing the registry refuses to start
without. The mirror crashlooped; the fleet fell through to upstream and nothing
user-visible broke. Since then, every config change runs the registry's own
verify in a throwaway pod before merging.The fleet's rescue tooling does not pull through the mirror: its image is cached on the operations host, outside the cluster, so it remains available while zot or its cluster is down.
Enforcing auth needs every node to carry credentials that only apply after a reboot, so anonymous read must survive until the last node is proven - and the proof must be designed not to lie. Section 07 covers it.
The node hosting the mirror boots its own workloads through it, including the tunnel serving this page - so restarting that node carries a blast radius beyond the mirror, and its reboot runbook lists every public-facing dependant and the order they return.
left = this lab, observed · right = the airgap translation, a design requirement not a deployed claim
Loop 1 assumes an upstream to fall through to; the airgap deletes that assumption and forces Loop 2's discipline onto the mirror itself.
Fallback is a policy decision. This lab keeps it on for bootstrap
survivability and availability through mirror outages; the cost is that pulls can bypass
the mirror unnoticed and land on origin rate limits. A regulated or disconnected estate
makes the opposite call - skipFallback: true, egress restriction, preseeded
release stock, a tested recovery path - because there the bypass is the failure mode.
Failure classes differ in symptom, path and consequence:
| Failure | User-visible symptom | Request path | Fallback? | Detection | Security consequence |
|---|---|---|---|---|---|
| Mirror pod down | public pulls may continue transparently; private images, origin limits, DNS or TLS issues can slow or fail them | node -> origin | after the mirror endpoint fails; elapsed time differs by failure class (refused vs DNS vs TLS vs blackhole) | mirror probes red; origin egress rises | policy/audit bypass while down |
| Mirror up, sync wedged | pulls hang up to sync timeout | node -> mirror (blocked) | only after timeout | manifest probe stalls while /v2/ still answers | availability, not integrity |
| Upstream 429 | miss/tag pulls fail or crawl | mirror -> origin refused | fallback also reaches the origin and may hit a pull or abuse limit - the quota bucket depends on the node's identity and source IP, not necessarily zot's | zot logs; retry storms amplify | self-inflicted denial of service via retries |
| Upstream down, digest cached | none expected | mirror serves locally | not needed | - | none (source-verified short-circuit; disconnected drill still open) |
| Upstream down, tag cached | reported on v2.1.17: failure or waiting until timeout, with content possibly served after it | mirror revalidates the tag upstream first | origin also down | the test matrix in section 08 | availability surprise for anyone assuming cached means offline-safe (upstream-reported; not yet reproduced here) |
| PVC full | new pulls fail, cached OK-ish | mirror 5xx on writes | yes for missing content | capacity metrics (proposed) | availability; GC pressure |
| Digest entry GC-evicted | silent re-fetch on next pull | mirror -> origin re-sync | n/a | upstream egress for "cached" content | rate-limit exposure returns (section 06) |
| Node auth wrong (post-flip) | ImagePullBackOff | mirror 401; fallback unreliable on 401 | unreliable | the canary gate (section 07) | the outage the gate exists to prevent |
Half the mistakes on this page came from conflating these objects:
name:tag@digest is client-side grammar, not wire protocol: the OCI
distribution spec takes a tag or a digest in the URL. Clients resolve the digest
and ignore the tag (Kubernetes documents this), so digest-pinned pulls are immune to tag
moves. This page once mislearned the combined form as a cache-miss bug; retesting on
v2.1.17 showed the historical zot rejection no longer reproduces.For a mirror, digests carry one sharp operational rule: zot converts Docker-schema
manifests to OCI by default, and conversion changes the digest - breaking every digest pin
and signature downstream. The pairing that prevents it (preserveDigest: true
per upstream, with http.compat: ["docker2s2"]) is mandatory in both
directions: the registry refuses to start with one and not the other (the Loop 1 field
note is this rule being learned). This deployment runs both on all five upstreams.
Digest preservation and artifact completeness are different properties. The docs tie
preserveDigest/compat to keeping mirrored manifest bytes and
media types - and therefore signature validity - aligned with upstream. Whether
signature, SBOM and attestation objects arrive is separate: OCI 1.1 referrers
ride the subject field and their own API, legacy Cosign signatures ride
tag-schema conventions, and both depend on origin support and this version's sync
behaviour. This deployment has not verified a complete referrer graph for any origin and
performs no cryptographic signature verification; disconnected verification would also
need locally held trust material. Both gaps sit in the open register. Production estates
promote releases by digest and decide per policy whether referrers travel too; this lab
runs only the digest half.
This section corrects the largest error in this page's first edition, which said "no retention config means keep everything". True for tags; false for everything else - and "everything else" includes the content this mirror exists to hold. The semantics, from upstream documentation and the deployed version's source:
gcDelay (1h here),
then deletion happens at the next sweep (gcInterval, 24h here) - the survival
window depends on where creation falls relative to that sweep, not a fixed one-hour fuse.Combined with section 04's revalidation behaviour: tag pulls are cached but contact upstream anyway; digest pulls serve locally but their entries are GC-eligible between sweeps - a digest-pinned fleet caches exactly the entry class GC may delete. On v2.1.17 this mirror is a rate-limit shield and a latency win, not yet a disconnection shelf; the first edition implied otherwise and was wrong.
green = protected by a tag · amber/red = the eviction path · dashed magenta = referrers (separate lifecycle)
GC evaluates reference reachability: an untagged digest-only manifest can be deleted even while its blobs remain present, and node-local image stores can mask the eviction for days.
Retention policy semantics - two nearby systems use opposite matching rules:
keepTags rule inverts the default:
non-matching tags in that repository become deletable. The inversion is scoped to
the repository, not global - the first edition of this page overstated it.deleteUntagged, default true),
and on the deployed version no retention rule can protect them. Pull-aware
keepUntagged exists upstream from v2.1.19.Mitigations, ranked: upgrade to v2.1.19+ for pull-aware
keepUntagged (the designed fix, schema to be validated against the actual
binary before rollout); until then, widen gcDelay/retention.delay
(the upstream maintainer's interim suggestion), or set deleteUntagged: false,
which protects every untagged manifest at the cost of unbounded cache growth. Disabling GC
entirely swaps eviction for storage exhaustion with no reclaim path.
Target state: the mirror refuses anonymous pulls. Current state: anonymous read on, per-node credentials staged inert in machine configs until each reboot. The sequencing problem has one central hazard: while anonymous read is on, an ordinary pull says nothing about auth - a node whose credentials never applied is served as an anonymous reader and false-passes the gate. The canary must live in a repository that denies anonymous read, so success can only mean an authenticated pull.
The design is upstream-supported: authorisation resolves per-repository policies by
longest match (** is the default for anything unmatched), and
maintainer guidance confirms anonymous and authenticated access evaluate independently. So
a canary/** entry granting named identities read, with no
anonymousPolicy, denies anonymous on that path while the glob keeps the rest
of the shelf open. The staged policy:
"accessControl": {
"repositories": {
"**": {
"anonymousPolicy": ["read"],
"policies": [
{ "users": ["zot-push"], "actions": ["read", "create", "update"] },
{ "users": ["zot-pull"], "actions": ["read"] }
]
},
"canary/**": {
"policies": [
{ "users": ["zot-pull", "zot-push"], "actions": ["read"] }
]
}
}
}
/v2/ returns 401 to Docker user agents so the Docker CLI
sends credentials, meaning anonymous docker users must log in even for anonymous
repos. Podman and containerd are unaffected. This estate pulls with containerd, podman and
buildah, so the caveat is noted rather than felt.The gate, per node, and what each step proves:
docker.io/... references route
through the mirror. That needs an original-upstream reference pulled on a node with a cold
runtime store, correlated with the mirror's request log - node-local content would satisfy
the pull without any network and fake a pass.the flip is per-fleet; the proof is per-node - one unproven node under a fleet-wide flip is an outage wearing a green tick
Identity and interception fail independently; each has a false-pass mode the other cannot detect.
Five trust domains never share material: node pull credentials (machine config, reboot-bound), the push credential (used by the operator-run build host for registry logins, not embedded in CI configuration), the mirror's upstream credentials (one mounted secret, pod-restart-bound), TLS trust, and human admin access (SSO in front of the UI). The push credential carries this page's oldest incident: in July 2026 its bcrypt hash and plaintext drifted apart in the secrets manager, rotation became impossible, and pushes stayed dead for twelve days - three applications shipped the day the rebuilt flow landed. Today the hash and plaintext are separate entries in the same secrets manager, rotated as a pair by a human; the registry consumes only the derived htpasswd file, each plaintext reaches only its consumer, and the sync identity is read-only so automation cannot half-rotate the pair. Residual cost, stated: one secrets-manager project holds both halves, so its compromise yields verifier and credential together - versioned, consumer-scoped secret objects remain a possible refinement.
The worst incident in this mirror's life: an in-flight sync wedged, every pull from that upstream hung, and the health endpoint returned 200 throughout:
"image already demanded, waiting on channel" $ kubectl -n zot rollout restart deploy zot <- recovery action (not a root cause) # probe an actual manifest, not the process's opinion of itself: $ curl -sfS --max-time 10 -o /dev/null -w '%{http_code}\n' \ https://zot.bztmon.org/v2/library/busybox/manifests/latest \ -H 'Accept: application/vnd.oci.image.index.v1+json' 200 <- THIS is "the mirror works"
The deployed source refines that story: the log line is coalescing by design -
concurrent requests join the first sync, which runs on a detached background context with
a three-hour default timeout and survives client disconnects. Waiters blocking on a
stalled sync until that timeout matches what we saw; the restart was recovery, and the
root cause was never isolated. Two knobs this config leaves unset - sync
maxRetries (disabled by default upstream) and a tighter
syncTimeout - are the first change on recurrence.
From the running version's source: /livez, /readyz and
/startupz exist as health endpoints (absent from this version's docs),
alongside the spec's /v2/. The Prometheus metrics extension exists upstream -
series names include zot_http_requests_total,
zot_repo_storage_bytes, zot_repo_downloads_total and
zot_storage_lock_latency_seconds - and is not enabled here;
wiring it, with scrape auth, is on the open list.
The layered checks, one question each:
| Check | Question it answers | Status here |
|---|---|---|
/livez / /readyz | is the process up / initialised | available |
| authenticated manifest GET of a local-only canary | can the right identity read real content from local storage | proposed |
| anonymous GET of the canary expecting 401 | is the protection actually protecting | proposed |
| original-reference pull on a cold node + mirror log line | is the runtime actually routed through the mirror | proposed |
| digest pull with upstream blocked (disposable env) | does "cached" mean "served locally" | open - section 06 |
| PVC usage + growth, GC activity, sync latency | when does capacity or eviction become the story | needs metrics ext |
| certificate expiry, credential age | what breaks on a schedule | proposed |
The caching test matrix (digest rows: source-verified and consistent with behaviour here; disconnected rows: not yet run in this lab):
| Scenario | Expected on v2.1.17 | What proves it |
|---|---|---|
| cached tag, upstream reachable | served; upstream contacted anyway (revalidation) | zot log shows upstream request; latency includes round-trip |
| cached tag, upstream unreachable | fails or stalls until timeout, then local fallback path | pull timing vs sync timeout; error text |
| cached digest, upstream reachable | served locally, no upstream contact | absence of upstream request in zot logs during the pull |
| cached digest, upstream unreachable | served locally - IF the entry survived GC | the section-06 disposable-instance drill |
| cold node, warm mirror | mirror serves; node store fills | mirror log + no origin egress |
| warm node, evicted mirror entry | pull succeeds from node store - masking the eviction | this is the false-pass: only mirror-side inspection reveals it |
| multi-arch index + child manifests | index and per-platform manifests are separate cache entries | per-digest existence checks on the mirror |
The first edition of this section asserted what "production" does; production chooses from patterns against requirements:
Whatever the topology: recovery time and point objectives get numbers before an incident supplies them; backups are restore-tested copies off the failure domain (array snapshots and RAID protect against disk loss, not against the array, the site or operator error); cache capacity is planned from the section 06 retention policy; and the section 04 circular dependencies get drawn per site, because each exists at every scale - only the cost of ignoring them changes.
When a node powers on, the fleet's admission check performs an authenticated manifest pull through the node's own runtime before the node counts as a member: a process-health response validates neither routing nor authorisation. Most of this page condenses into that one command.
| Decision | Chosen | Why | Trade-off | Rollback / alternative | Status |
|---|---|---|---|---|---|
| Origin fallback | ON (default kept) | bootstrap survivability, Loop 1 | silent mirror bypass possible | skipFallback:true + preseed (airgap shape) | observed |
| Sync mode | onDemand, all upstreams | cache follows real usage; Hub-safe | tag pulls revalidate upstream; digest entries untagged | polled sync for a curated release set | observed |
| Digest preservation | preserveDigest + docker2s2, x5 | digest pins + signatures survive mirroring | mandatory config pairing (crashloop scar) | none - non-negotiable for digest-pinned estates | observed |
| Namespace layout | flat (no destinations) | simplicity at build time | ordered-trial upstream selection; theoretical collisions | per-origin destinations + overridePath | redesign proposed |
| Retention | none configured | predates understanding the untagged rule | digest-only entries evict within hours | v2.1.19+ keepUntagged; interim wider gcDelay | correction owed |
| Auth | anonymous read until per-node proof | containerd 401-fallback unreliability makes big-bang flips dangerous | window with no pull auth | staged flip w/ canary gate + named rollback | in flight |
| Metrics | not enabled | minimal first deployment | capacity/eviction invisible | metrics extension + scrape auth | proposed |
canary/** while anonymous
read elsewhere still succeeds; observed 401-vs-403 boundaries recorded.keepUntagged, with the retention
config written and reviewed before the upgrade, not after.