Skip to content

Reclaiming snapshot storage

A marketplace polled daily accumulates a snapshot per upstream commit forever, each pinning a commit in the quarantine repository. Retention reclaims that space on stated criteria, with a restore window so a mistake is recoverable.

Nothing here happens on its own until you turn it on.

This guide enables a destructive capability

Retention deletes. Work through it in the order below — preview first, enable second — and read Snapshot retention for exactly what the criteria select and what they can never touch.

1. See what a policy would delete

The dry run writes nothing and works with retention disabled, which is the whole point: inspect the effect before switching the scheduler on.

$ curl 'localhost:8080/api/v1/retention/candidates'
[{"snapshotId":17,"marketplace":"acme","sha":"3f9c2ab...","state":"held",
  "reason":"held-too-long","createdAt":"2026-04-02T08:11:03Z"},
 {"snapshotId":21,"marketplace":"acme","sha":"9d1e4c7...","state":"rejected",
  "reason":"superseded","createdAt":"2026-05-19T14:02:47Z"}]

reason is the criterion that selected it. Add ?marketplace=acme to scope the preview to one upstream.

If the list is longer than you expected, adjust the policy rather than the plan — held-max-age and superseded-min-age are the two knobs that move it most.

2. Choose the policy

Global defaults with per-marketplace overrides, in your deployment configuration:

skills-gateway:
  retention:
    enabled: false          # still off; step 3 turns it on
    defaults:
      held-max-age: 90d
      superseded: true
      superseded-min-age: 30d
      min-idle: 30d
      restore-window: 14d
    marketplaces:
      acme:
        held-max-age: 30d   # a busy upstream; everything else inherited

Unset override fields fall back to defaults, so a per-marketplace block only states what differs. The full reference is Configuration.

Two rules worth internalising before you tune:

  • An approved snapshot is never eligible, by policy or by hand. Retention cannot delete what the facade is serving.
  • min-idle is a veto, not a selector. It only removes candidates that were fetched recently; it never selects anything.

3. Run a pass by hand

Still with the scheduler off, apply the policy once and watch what happens:

$ curl -X POST 'localhost:8080/api/v1/retention/evaluate?marketplace=acme'
{"selected":2,"acted":2}

The selected snapshots are now soft-deleted: marked with a reason and a purge_after deadline, their vetting state untouched, and restorable for the whole restore window. Nothing has been removed from git.

Check the portal — the marketplace's Snapshots shows each deleted snapshot with a deleted badge and its restore deadline.

4. Restore anything you did not mean to delete

On the marketplace's Snapshots, the deleted snapshot's Restore button. It toasts Snapshot {id} restored.

$ curl -X POST localhost:8080/api/v1/snapshots/17/restore

409 means the snapshot is not deleted; 404 means compaction has already removed it, and at that point there is nothing to restore.

You can also delete a single snapshot by hand — the Delete button, or DELETE /api/v1/snapshots/17, recorded with the reason manual. Approved snapshots offer no delete control and the endpoint refuses them with 409.

5. Enable the schedulers

Once the previews and a manual pass look right:

skills-gateway:
  retention:
    enabled: true
    poll-interval: 1h         # evaluation: marks
    compaction-interval: 6h   # compaction: removes
    batch-size: 200

Evaluation soft-deletes on its interval; compaction permanently removes snapshots whose restore window has elapsed, deleting the pinned quarantine ref and garbage-collecting the repository so the objects are actually reclaimed.

Compaction is the irreversible half

Restore works only before purge_after. After compaction the row and the unreachable objects are gone; what remains is the ledger entry saying the SHA existed and was purged. Give yourself a restore window you would actually notice a mistake within — the default is 14 days.

To force a compaction pass — after a deliberate cleanup, say:

$ curl -X POST localhost:8080/api/v1/retention/compact
{"selected":2,"acted":2}

selected is what was due; acted is what was actually removed. A snapshot whose quarantine ref could not be deleted is left in place for the next pass rather than removed from the table.

The same pass also sweeps abandoned publication staging refs out of the published repositories — the leftovers of a gateway killed part-way through publishing a snapshot. They serve nothing, but they hold objects against collection. The counts above do not include them; the ledger entry staging-refs-swept:count=<n> does. Nothing is removed until it has been under observation for staging-ref-max-age (default 24h), so a publication running right now is never disturbed. See Snapshot retention.

Webhook delivery history

The same pass removes webhook deliveries that are delivered or failed and were last updated more than 30 days ago, up to batch-size per pass. A pending delivery is never removed, whatever its age. The age is fixed, not a setting; it is not ledger-max-age. Each pass that removes any writes webhook-deliveries-swept:removed=<n> to the ledger. With retention off, delivery history is kept.

6. Watch it

Every retention action is in the append-only ledger with the acting identity: retention-evaluated:…, snapshot-soft-deleted:<reason>, snapshot-restored, snapshot-purged, staging-refs-swept:count=<n> and webhook-deliveries-swept:removed=<n>. Policy-driven deletions are attributed to retention-policy rather than to a person.

Soft delete and restore also fire the marketplace.snapshot.soft_deleted and marketplace.snapshot.restored webhook events, so an inventory system can follow deletions the same way it follows approvals — see Receiving lifecycle webhooks.

7. Bound the audit ledger

Retention compacts snapshots. The audit ledger grows independently, and it grows faster: info-refs is appended on every git fetch whether or not anything transfers, so a few thousand developers polling every half hour is a couple of hundred thousand rows a day, and nothing removed them.

The same six-hourly compaction pass now trims it, behind the same enabled switch and one setting:

skills-gateway:
  retention:
    enabled: true
    ledger-max-age: 90d

A ledger no sink exports is never trimmed — whatever this is set to

An entry is eligible only when it is both older than ledger-max-age and already taken by every enabled audit export sink. The export position is the gateway's only evidence that some other system holds a copy, so trimming behind it is the one deletion that destroys no record.

With no enabled sink there is no such position, nothing is eligible, and the trim does nothing at all. Registering an export destination is the price of bounding the table — the gateway will not be both the sole holder of the audit evidence and the thing that deletes it. Watch skills_gateway.ledger.export_lag_seconds to see the ledger you are keeping forever; see Observability.

Unset, zero or negative switches the trim off rather than making every entry instantly eligible: the mis-typed value must not be the one that deletes.

Administrative entries are never trimmed. Only the two facade read events (info-refs, upload-pack) are eligible; approvals, rejections, revocations, registrations, token and sink lifecycle — the compliance-bearing half, and the low-volume one — survive every pass whatever their age. The admitted set is a closed allowlist, so a ledger event added to the gateway later is not trimmable until somebody deliberately admits it.

Each pass is bounded by a fixed work budget and resumes on the next one, so the first run against a ledger years deep does not hold the compaction lease until it finishes. When a pass removes anything it logs how many and which sink bounded it.

Turning it off again

Set enabled: false. The schedulers stop; already-marked snapshots keep their marks and stay restorable, and nothing new is selected or purged. Restoring them is the explicit second step. The ledger trim stops with them — it is a duty of the same pass — and clearing ledger-max-age stops it on its own.