Skip to content

Snapshot retention

What retention may select, what it may never touch, and what the two passes do. Task-shaped coverage is in Reclaiming snapshot storage.

Retention is off by default (skills-gateway.retention.enabled: false). It deletes only because an operator asked for it.

Criteria

Evaluation resolves the policy in force for each marketplace, then applies these:

Criterion Kind Selects Governed by
Held too long Selector held snapshots ingested more than held-max-age ago. held-max-age
Superseded Selector held, rejected or revoked snapshots of a marketplace that a later approved snapshot has overtaken, once older than superseded-min-age. superseded, superseded-min-age
Minimum idle Veto Nothing. It removes any candidate whose SHA was fetched from its own marketplace through the facade within min-idle. min-idle
Approved Absolute guard Nothing. An approved snapshot is never eligible, by policy or by hand. Not configurable
Administratively withdrawn Partial guard Nothing. The content of a snapshot an administrator revoked is reclaimed as usual; its record is never permanently removed. Not configurable

A snapshot is deleted when a selector picks it and no veto or guard removes it.

Administratively withdrawn snapshots

A snapshot an administrator withdrew is treated like any other for the purpose of reclaiming disk: it can be soft deleted, and its git content removed. What is never removed is the row.

The refusal that stops a withdrawn commit being approved again is derived from that row. Delete it and the withdrawal expires on a timer nobody connected to it: the same upstream commit ingests afterwards as an ordinary held snapshot, carrying no trace that anything was withdrawn, and a reviewer approves it in good faith. Nothing refuses, and nothing is recorded — which makes it a quieter failure than the one the approval gate closes.

Keeping the row is cheap. The bytes are the part worth reclaiming and the part nobody wants to keep; the row is small, and it is the only thing that remembers.

This applies to administrative withdrawals only. A snapshot revoked by re-vetting purges as before: that is a machine verdict a later run can reach again from the content itself, so nothing is lost when its row goes.

Held too long

state = 'held' and created_at older than held-max-age (default 90d). This is the unreviewed-quarantine backlog: snapshots ingested by polling that nobody will ever open.

Setting held-max-age to zero or a negative duration disables the criterion rather than making everything eligible — a fail-safe reading, since the mis-typed value is the one that would otherwise delete the whole backlog.

Superseded

state IN ('held','rejected'), older than superseded-min-age (default 30d), and some snapshot of the same marketplace with a higher id is approved.

Supersession deliberately does not extend to older approved snapshots, even though they are no longer published — main was force-updated past them. An older approved snapshot is exactly what a future team catalog would pin, and the gateway cannot yet prove that nothing references it.

The criterion is a boolean as well as an age: superseded: false turns it off without disturbing superseded-min-age.

Minimum idle

A candidate whose SHA appears in the ledger as a facade fetch of its own marketplace (source <> 'admin') within min-idle (default 30d) is dropped from the pass. The marketplace is part of the match because a SHA is not unique across marketplaces — a fork, a mirror, or the same upstream registered twice all carry it — and the facade serves per marketplace, so traffic to one marketplace says nothing about whether another's identically-pinned snapshot is still in use.

This veto cannot select anything today — and that is the point

Only approved snapshots are ever served, and approved snapshots are categorically ineligible, so "not fetched in N days" can never pick a candidate on its own in the current model. It is wired in as a veto so that the first feature which unpublishes a snapshot or pins historical approvals into a catalog inherits the protection instead of having to remember it.

Approved snapshots

An approved snapshot is what the facade serves. It is never eligible:

  • every eligibility query names the deletable states explicitly — state IN ('held', 'rejected', 'revoked') — so a policy pass cannot reach an approved one;
  • the soft-delete UPDATE itself excludes approved snapshots, so served content stays served whatever a caller asks for;
  • DELETE /api/v1/snapshots/{id} refuses an approved snapshot with 409 before anything is written.

The guard is in SQL, not only in Java, which is why the check holds for the policy pass and the manual endpoint alike. No snapshot the facade serves is ever selected, and the only snapshot ref compaction removes lives in quarantine. The one thing retention does touch on the published side is the abandoned staging reference sweep, which cannot reach a served ref at all.

Revoked snapshots

A snapshot that re-vetting revoked is not approved, so the guard above no longer covers it. That is deliberate, not a side effect: the deletable states are named explicitly in the SQL precisely so that a new state has to be added on purpose to become deletable.

Retention treats revoked exactly as it treats rejected:

  • the superseded criterion may select it — it is not being served, and a later approved snapshot has taken its place;
  • held-max-age never does, because that criterion names held itself;
  • an administrator may delete it by hand, which the approved guard refused before the revocation;
  • the min-idle veto still applies, which is what keeps a recently-revoked snapshot around while the consumers that fetched it before the revocation are still recent.

Deleting one destroys nothing anyone could fetch. What it was revoked for, and who had already fetched it, stays in the append-only ledger regardless.

The two passes

Deletion is reversible first and permanent later, and the two are separate passes so that a wrong criterion costs a mark rather than content.

stateDiagram-v2
    [*] --> Live : ingested
    Live --> SoftDeleted : evaluate() selects it,\nor DELETE /api/v1/snapshots/{id}
    SoftDeleted --> Live : POST /api/v1/snapshots/{id}/restore\n(clears deleted_at)
    SoftDeleted --> Purged : compact() after purge_after
    Purged --> [*]

    note right of Live
        state stays held | approved | rejected | revoked
        deletion is orthogonal, never a state of its own
    end note
    note right of SoftDeleted
        deleted_at, deleted_reason, purge_after set
        restorable for the whole restore window
        approved snapshots never enter this state
    end note
    note right of Purged
        quarantine ref refs/snapshots/<sha> deleted,
        repository gc'd, row deleted;
        the ledger keeps the record
    end note

Evaluate (skills-gateway.retention.poll-interval, default hourly) marks: it sets deleted_at, deleted_reason and purge_after = now + restore-window. The vetting state is untouched — a deleted snapshot was still held or rejected, and provenance and the ledger must keep saying so.

Compact (compaction-interval, default six-hourly) removes: for each snapshot whose purge_after has elapsed it deletes refs/snapshots/<sha> with JGit, clears refs/quarantine/incoming when it still points at that commit, deletes the row, and garbage-collects the quarantine repository once per marketplace per pass with the expiry set to now. It then runs the staging reference sweep on the published side. The two halves are independent: a sweep that fails does not fail the pass, and the {selected, acted} counts are about snapshots only.

Compaction is irreversible

After compaction the row is gone and the objects the deleted tip made unreachable are gone with it. There is no undo and no archive: what survives is the ledger entry recording that the SHA existed and was purged. Restore is only possible before purge_after.

Edge cases worth knowing:

  • Objects still reachable from another snapshot's ref are kept. Snapshots are usually commits on one branch, so a purge often reclaims little until the older tips go too.
  • If the ref deletion fails, the row is left in place and the next pass retries. Deleting the record while the ref survived would strand objects with nothing left to say what they were.
  • If garbage collection fails the pass still succeeds; the refs are already gone so the space stays reclaimable by the next collection.

Abandoned publication staging references

Publication copies a snapshot's objects into the published repository under an unadvertised refs/staging/<sha> and only then moves the served refs, so a transfer that completes and a transition that is then refused leaves nothing on the wire. A gateway killed between the two leaves the staging ref behind. It serves nothing — the facade advertises refs/heads/main and refs/snapshots/* and nothing else — but it holds its objects against collection, so the repository carries a whole snapshot that will never be served.

The compaction pass removes such a ref when both of these hold:

Condition Why
No live snapshot row of that marketplace names the commit the ref points at. Approval writes the approved row before it publishes, so a publication in flight always has one.
The ref has been under observation for longer than staging-ref-max-age. The ref listing and the database query are two reads taken at different instants. Without the bound, a publication starting between them presents a ref the query did not see.

The second condition is what makes this safe. Deleting a live publication's staging ref and collecting its objects would leave the marketplace published at a commit whose content is gone — a disk leak turned into data loss. The gateway records the first time each pass sees a staging ref, in a table of its own, because a ref carries no creation time either storage backend can be asked for; the age that yields is an under-estimate, which only ever delays a sweep.

Served refs are unreachable from the sweep by construction: it lists refs/staging/* and nothing else. Objects still reachable from refs/heads/main or refs/snapshots/* survive the garbage collection that follows, which is what keeps served content served.

Set staging-ref-max-age above your slowest publication

The two ways of being wrong are not symmetric. Too generous costs the disk of an abandoned snapshot for a while longer — the cost that exists today anyway. Too tight risks collecting a publication's objects mid-flight. The default of 24h is far past any plausible object transfer. Zero or negative switches the sweep off entirely rather than making everything eligible.

Endpoints

Endpoint Purpose
GET /api/v1/retention/candidates Dry run — what a pass would select right now, each with the criterion that selected it. Writes nothing. ?marketplace= restricts it.
POST /api/v1/retention/evaluate Run one evaluation pass now. ?marketplace= restricts it. 200 with {selected, acted}.
POST /api/v1/retention/compact Run one compaction pass now. 200 with {selected, acted}.
DELETE /api/v1/snapshots/{id} Soft-delete by hand, reason manual. 409 if approved or already deleted, 404 if unknown.
POST /api/v1/snapshots/{id}/restore Clear the marks. 409 if the snapshot is not deleted, 404 if unknown.

The on-demand passes work whether or not the scheduler is enabled, which is what makes a policy inspectable before it is switched on.

The candidates preview requires auditor (or admin); the passes, the delete, and the restore require admin. See Delegated administration.

Policies

Configured under skills-gateway.retention — global defaults plus a marketplaces.<name> override map whose unset fields fall back to the defaults. See Configuration for the full block.

Knob Default Effect
held-max-age 90d Age at which a held snapshot is selected; zero or negative disables the criterion.
superseded true Whether the supersession criterion applies.
superseded-min-age 30d Minimum age before a superseded snapshot is selected.
min-idle 30d Fetch-recency veto window.
restore-window 14d How long a soft-deleted snapshot stays restorable.

The restore window is resolved at deletion time from the marketplace's policy, so shortening it later does not pull in snapshots already marked with a longer window.

What is recorded

Every retention action lands in the append-only ledger with the acting identity:

Ledger event Written by
retention-evaluated:selected=<n>,deleted=<n> Each evaluation pass.
snapshot-soft-deleted:<reason> Each soft delete — reason held-too-long, superseded, or manual.
snapshot-restored Each restore.
snapshot-purged Each compaction removal, carrying the SHA.
webhook-deliveries-swept:removed=<n> Each compaction pass that removed delivered or failed webhook deliveries last updated more than 30 days ago (at most batch-size per pass; pending deliveries are never removed).
staging-refs-swept:count=<n> Each compaction pass that removed abandoned staging refs from a marketplace's published repository.

Soft delete and restore also emit the marketplace.snapshot.soft_deleted and marketplace.snapshot.restored lifecycle webhook events. A policy-driven deletion carries the actor retention-policy, so a receiver can tell a scheduled deletion from an operator's.