Reclaiming snapshot storage¶
A marketplace polled daily accumulates a snapshot per upstream commit forever, each pinning a commit in the quarantine repository. Retention reclaims that space on stated criteria, with a restore window so a mistake is recoverable.
Nothing here happens on its own until you turn it on.
This guide enables a destructive capability
Retention deletes. Work through it in the order below — preview first, enable second — and read Snapshot retention for exactly what the criteria select and what they can never touch.
1. See what a policy would delete¶
The dry run writes nothing and works with retention disabled, which is the whole point: inspect the effect before switching the scheduler on.
[{"snapshotId":17,"marketplace":"acme","sha":"3f9c2ab...","state":"held",
"reason":"held-too-long","createdAt":"2026-04-02T08:11:03Z"},
{"snapshotId":21,"marketplace":"acme","sha":"9d1e4c7...","state":"rejected",
"reason":"superseded","createdAt":"2026-05-19T14:02:47Z"}]
reason is the criterion that selected it. Add ?marketplace=acme to scope the
preview to one upstream.
If the list is longer than you expected, adjust the policy rather than the plan —
held-max-age and superseded-min-age are the two knobs that move it most.
2. Choose the policy¶
Global defaults with per-marketplace overrides, in your deployment configuration:
skills-gateway:
retention:
enabled: false # still off; step 3 turns it on
defaults:
held-max-age: 90d
superseded: true
superseded-min-age: 30d
min-idle: 30d
restore-window: 14d
marketplaces:
acme:
held-max-age: 30d # a busy upstream; everything else inherited
Unset override fields fall back to defaults, so a per-marketplace block only
states what differs. The full reference is
Configuration.
Two rules worth internalising before you tune:
- An approved snapshot is never eligible, by policy or by hand. Retention cannot delete what the facade is serving.
min-idleis a veto, not a selector. It only removes candidates that were fetched recently; it never selects anything.
3. Run a pass by hand¶
Still with the scheduler off, apply the policy once and watch what happens:
The selected snapshots are now soft-deleted: marked with a reason and a
purge_after deadline, their vetting state untouched, and restorable for the
whole restore window. Nothing has been removed from git.
Check the portal — the marketplace's Snapshots
shows each deleted snapshot with a deleted badge and its restore deadline.
4. Restore anything you did not mean to delete¶
409 means the snapshot is not deleted; 404 means compaction has already removed it, and at that point there is nothing to restore.
You can also delete a single snapshot by hand — the Delete button, or
DELETE /api/v1/snapshots/17, recorded with the reason manual. Approved
snapshots offer no delete control and the endpoint refuses them with 409.
5. Enable the schedulers¶
Once the previews and a manual pass look right:
skills-gateway:
retention:
enabled: true
poll-interval: 1h # evaluation: marks
compaction-interval: 6h # compaction: removes
batch-size: 200
Evaluation soft-deletes on its interval; compaction permanently removes snapshots whose restore window has elapsed, deleting the pinned quarantine ref and garbage-collecting the repository so the objects are actually reclaimed.
Compaction is the irreversible half
Restore works only before purge_after. After compaction the row and the
unreachable objects are gone; what remains is the ledger entry saying the
SHA existed and was purged. Give yourself a restore window you would
actually notice a mistake within — the default is 14 days.
To force a compaction pass — after a deliberate cleanup, say:
selected is what was due; acted is what was actually removed. A snapshot
whose quarantine ref could not be deleted is left in place for the next pass
rather than removed from the table.
The same pass also sweeps abandoned publication staging refs out of the
published repositories — the leftovers of a gateway killed part-way through
publishing a snapshot. They serve nothing, but they hold objects against
collection. The counts above do not include them; the ledger entry
staging-refs-swept:count=<n> does. Nothing is removed until it has been under
observation for staging-ref-max-age (default 24h), so a publication running
right now is never disturbed. See
Snapshot retention.
Webhook delivery history¶
The same pass removes webhook deliveries that are delivered or failed and were last
updated more than 30 days ago, up to batch-size per pass. A pending delivery is never
removed, whatever its age. The age is fixed, not a setting; it is not ledger-max-age. Each
pass that removes any writes webhook-deliveries-swept:removed=<n> to the ledger. With
retention off, delivery history is kept.
6. Watch it¶
Every retention action is in the append-only ledger with the acting identity:
retention-evaluated:…, snapshot-soft-deleted:<reason>, snapshot-restored,
snapshot-purged, staging-refs-swept:count=<n> and webhook-deliveries-swept:removed=<n>. Policy-driven deletions are
attributed to retention-policy rather than to a person.
Soft delete and restore also fire the marketplace.snapshot.soft_deleted and
marketplace.snapshot.restored webhook events, so an inventory system can follow deletions
the same way it follows approvals — see
Receiving lifecycle webhooks.
7. Bound the audit ledger¶
Retention compacts snapshots. The audit ledger grows independently, and it grows
faster: info-refs is appended on every git fetch whether or not anything
transfers, so a few thousand developers polling every half hour is a couple of
hundred thousand rows a day, and nothing removed them.
The same six-hourly compaction pass now trims it, behind the same
enabled switch and one setting:
A ledger no sink exports is never trimmed — whatever this is set to
An entry is eligible only when it is both older than ledger-max-age
and already taken by every enabled
audit export sink. The export position is
the gateway's only evidence that some other system holds a copy, so trimming
behind it is the one deletion that destroys no record.
With no enabled sink there is no such position, nothing is eligible, and the
trim does nothing at all. Registering an export destination is the price of
bounding the table — the gateway will not be both the sole holder of the
audit evidence and the thing that deletes it. Watch
skills_gateway.ledger.export_lag_seconds to see the ledger you are keeping
forever; see Observability.
Unset, zero or negative switches the trim off rather than making every entry instantly eligible: the mis-typed value must not be the one that deletes.
Administrative entries are never trimmed. Only the two facade read events
(info-refs, upload-pack) are eligible; approvals, rejections, revocations,
registrations, token and sink lifecycle — the compliance-bearing half, and the
low-volume one — survive every pass whatever their age. The admitted set is a
closed allowlist, so a ledger event added to the gateway later is not trimmable
until somebody deliberately admits it.
Each pass is bounded by a fixed work budget and resumes on the next one, so the first run against a ledger years deep does not hold the compaction lease until it finishes. When a pass removes anything it logs how many and which sink bounded it.
Turning it off again¶
Set enabled: false. The schedulers stop; already-marked snapshots keep their
marks and stay restorable, and nothing new is selected or purged. Restoring them
is the explicit second step. The ledger trim stops with them — it is a duty of the
same pass — and clearing ledger-max-age stops it on its own.