Vetting — the vetter chain¶
Every snapshot the gateway ingests is quarantined and pinned to an upstream commit SHA. Between that pin and the reviewer's decision the gateway runs a vetting chain: an ordered list of vetters, each of which looks at the snapshot's content and answers with a verdict. The verdicts and their findings are recorded against the snapshot and shown to the reviewer before any approve or reject decision.
The gateway does not vet content itself. It orchestrates: it runs vetters, normalises what they answer, records it, and aggregates it into a single outcome that gates approval.
Where the chain sits¶
flowchart TD
U["Upstream repository"] -->|"POST /api/v1/marketplaces/{name}/ingest"| F["Fetch into quarantine<br/>pin refs/snapshots/<sha>"]
F --> M{"Manifest policy<br/>local sources only"}
M -->|violation| R["snapshot: rejected<br/>chain does not run"]
M -->|ok| H["snapshot: held"]
H --> C["Vetting chain"]
subgraph C["Vetting chain — ordered; every vetter by default"]
direction TB
C1["secret-scan (order 100)"] --> C2["prompt-injection (order 200)"] --> C5["executable-surface (order 250)"] --> C3["license-scan (order 300)"] --> C4["skill-conformance (order 400)"]
end
C --> A["Aggregate verdicts<br/>clear iff every verdict is pass or warn"]
A --> E["Effective outcome<br/>run + waivers active right now"]
E -->|clear| RV["Reviewer sees verdicts<br/>Approve enabled"]
E -->|clear with waivers| RW["Reviewer sees verdicts<br/>and what was accepted"]
E -->|blocked| RB["Reviewer sees the findings<br/>no waiver covers"]
RV -->|"POST /api/v1/snapshots/{id}/approve"| P["Published repository<br/>refs/heads/main"]
RW -->|"POST /api/v1/snapshots/{id}/approve"| P
RB -->|"POST /api/v1/snapshots/{id}/waivers per finding"| E
The chain never changes the snapshot's own state. A vetted snapshot is still
held; what the chain decides is whether approving it is an ordinary act or one
that has to be justified in writing.
That separation is what lets the same chain run again, later, over a snapshot that is already approved and served — a new run against unchanged state. What such a run means is a separate judgement, described in Re-vetting approved content: only a vetter that objects to the content can retract it, and a vetter that merely broke never can.
The vetter contract¶
A vetter has a stable name, a position in the chain, and one method that takes the snapshot and returns a verdict. What it is given is deliberately narrow: the snapshot's id, its marketplace, its commit SHA, and a walk over the files in that commit. It is not given a repository handle, so a vetter cannot move a ref, write to quarantine, or read another marketplace's content.
The walk takes the vetter's own selection of paths, and content is read only
for the files that selection asks for. One chain run walks the commit's tree
once and reads each selected file once however many vetters want it, so
license-scan reading two file kinds no longer costs a full pass over the tree.
What a vetter sees is unchanged by this: a file over the size limit is still
handed over unread rather than omitted. How much content a run keeps for that
reuse is bounded by
skills-gateway.vetting.content-cache-bytes;
beyond the bound content is read again rather than kept, so the setting costs
speed and never coverage.
A verdict carries a state, an optional external report URL, and a list of
findings. A finding has a stable rule id (aws-access-key-id), a severity,
a location (normally path:line), and a reviewer-facing message. The rule id is
the identity a scoped waiver
is written against, so it is part of the contract rather than a display string.
Verdict states¶
| State | Meaning | Blocks approval |
|---|---|---|
PASS |
The vetter found nothing. | No |
WARN |
Something worth showing the reviewer that does not block. | No |
FAIL |
Something that blocks. | Yes |
ERROR |
The vetter produced no verdict: it threw, or it exceeded its time limit. | Yes |
PENDING |
The vetter was triggered and has not answered yet. | Yes |
DISABLED |
An administrator switched the vetter off for this marketplace, so the chain did not run it. | No — and it does not clear either |
NOT_REACHED |
The chain stopped before this vetter, so it was not run. | No — and it does not clear either, but see below |
A vetter does not choose its state directly: it emits findings, and the
verdict follows the worst severity present — HIGH or CRITICAL fails,
LOW or MEDIUM warns, and INFO alone still passes. That way a new rule only
has to get its severity right.
PENDING exists so that an asynchronous vetter — one that is triggered over
a webhook and answers later — fits without changing the gate. No built-in
vetter returns it; a configured external connector
that answers pending does, and it blocks until it is resolved.
Switching a vetter off¶
An administrator can disable a vetter globally or for one marketplace. A
disabled vetter is not run at ingestion or re-vetting; the chain records a
DISABLED verdict in its place, so the disablement is part of the run's
evidence rather than a silently shorter chain. Because DISABLED never clears,
disabling every vetter leaves a run blocked — the switch is not a blanket
approval. The settings, the audit events and the endpoints are in
the API reference.
Stopping the chain after a failure¶
By default the chain runs every enabled vetter, whatever the ones before it concluded. An administrator can change that for one marketplace, or globally:
| Chain mode | What it does |
|---|---|
run-all |
Every enabled vetter runs. The default. |
stop-after-fail |
The chain stops after the first vetter whose verdict still objects to the content, and the vetters after it are recorded NOT_REACHED. |
What it buys is real: an expensive vetter — one billed per call, or one that detonates the snapshot in a sandbox — is not spent on content a cheap deterministic scan has already condemned, and that cost is otherwise paid again on every re-vetting pass over the whole approved estate.
What it costs a reviewer is equally real, and is the reason it is off by
default. A run that stopped early is not a shorter description of the same
answer; it is a smaller answer. The reviewer no longer sees everything that is
wrong with a snapshot in one pass, so a fix-and-re-ingest loop can take several
rounds, each revealing one more objection. Under run-all a run's verdict set is
complete by construction; under stop-after-fail it is not, and the portal says
so above the chain.
Three properties keep a shorter run from becoming a cleaner one:
- Nothing is omitted. Every vetter of the chain still has a verdict. A vetter that was not reached says so, and says which vetter stopped the chain.
NOT_REACHEDnever clears. LikeDISABLEDit is an absence rather than a conclusion, and the aggregation still requires at least one clearing verdict.- A run that stopped early is blocked, whatever is waived. This is the one that matters. Accepting the finding that stopped the chain does not clear the run — the vetters behind it never looked at the content, and a gate opened on their silence would be indistinguishable from a gate opened on a clean chain. A waiver makes the next run get further; the run that stopped stays blocked until the chain has been run again. Re-vet the snapshot, and the gate is then decided on complete evidence.
Only a FAIL stops the chain — not an ERROR, which is a fact about the
gateway rather than about the content, and not a PENDING, which has concluded
nothing. And only a FAIL that a reviewer has not already accepted: a verdict
whose findings are all waived does not condemn the snapshot, so the chain runs
on past it.
Changing the mode is an administrator-only, audited act, like the on/off switch. The endpoints are in the API reference.
The order the vetters run in¶
Order only becomes a control once the chain can stop: under run-all it decides
nothing but the sequence of the run and the left-to-right position in the
portal's chain flow. Under stop-after-fail it decides which vetters get to look
at all, so the cheap and deterministic ones belong first.
An administrator can set the order globally or for one marketplace. The arrangement need not name every vetter. The resolved order is:
- the vetters the arrangement names, in the order it names them;
- then every vetter it does not name, by its configured position, with ties broken by the vetter's name.
That rule is total and deterministic, so a recorded run is reproducible, and an arrangement can never drop a vetter by omission — including a vetter a later release of the gateway adds, or an external connector an operator configures afterwards. An arrangement naming a vetter that does not exist, or naming one twice, is refused rather than stored.
How the settings resolve¶
Both settings resolve the same way the vetter on/off switch does: the setting scoped to the marketplace when there is one, otherwise the global setting, otherwise the default. The administrative surface reports which of the three decided, so a default is never mistaken for a decision somebody made. The global level and the marketplaces that depart from it are governed together on the portal's Vetting chain page.
Both are recorded with every run, in the run's chain identity, alongside each
vetter's rule-set version — secret-scan@3,prompt-injection@1;mode=run-all. That
is what keeps re-vetting able to answer its own
question: when a snapshot that cleared last month is blocked today, was it the
content that changed, or the chain?
Fail-closed aggregation¶
A chain run is clear if and only if it produced at least one clearing
verdict and every verdict is PASS, WARN, DISABLED or NOT_REACHED.
Everything else is blocked:
- any vetter that failed;
- any vetter that crashed or timed out — a crash is a blocked snapshot, never a skipped vetter;
- any vetter that has not answered;
- and a snapshot with no chain run at all.
That last case is the one that matters most. A snapshot ingested before the chain existed, or one whose run died halfway, is blocked — absence of evidence is not evidence of safety.
- and a run that stopped early, however its verdicts aggregate and whatever waivers are active: the vetters it did not reach have said nothing either way, so the run is missing evidence rather than merely missing an objection.
By default all vetters run, in order, and the chain does not stop at the first failure, because a reviewer should see everything that is wrong with a snapshot at once. An administrator can trade that away per marketplace — see Stopping the chain after a failure.
The approval gate¶
POST /api/v1/snapshots/{id}/approve takes no request body. It refuses a snapshot
whose effective outcome is blocked, with 409 and a problem document naming
both the blocking vetters and — in uncoveredFindings — every blocking
finding that no active waiver covers. That array is the reviewer's worklist: it
is exactly the set of waivers that must exist for the approval to succeed.
A reviewer has no override. The only way a reviewer gets past objecting vetters is to accept each blocking finding individually with a waiver.
An administrator — and only an administrator — can approve a blocked snapshot outright, by supplying a mandatory reason. The override lifts only the vetting block: the policy, name-collision, minimum-release-age and four-eyes gates still decide, and it lands on the ledger as its own event, so it is never indistinguishable from an approval the chain cleared. See Administrative override of a blocked outcome.
Not a vetter: the minimum release age¶
The same gate carries one precondition the chain has nothing to do with. When
minimum-release-age is
configured, a snapshot the gateway ingested less than that long ago cannot be
approved however clear its verdicts are — a cooling-off window that gives the
world time to notice a compromised release before this gateway adopts it.
It is deliberately not a vetter. A verdict is evidence about content at the moment it was gathered; "too young" is a fact about now. Recorded as a failing verdict it would keep blocking after the age had passed, until some later re-vetting run happened to replace it. Checked at the approval request instead, it clears itself and leaves nothing behind — the same reasoning that makes waiver expiry a comparison rather than a state.
Not a vetter either: plugin names already in use¶
The gate also refuses a snapshot that would introduce a plugin name looking like
one another marketplace already serves — the typosquat, c0de-review beside
code-review. Every vetter answers a question about the pinned content; this
asks whether the rest of the estate already has something called this, which
is a fact about now, like the release age. Asked inside the chain it would make a
verdict depend on an input no run records, and a changed verdict over unchanged
content would stop meaning that the chain learned something — the property
re-vetting rests on. So a chain run never reads
the estate, and this is checked at the approval request instead
(ADR 0015).
What it compares, deliberately narrowly:
- Plugin names, never skill names. Clients reach a skill under its
plugin's name, so
initorreviewin several plugins impersonates nothing. - Only names new to the marketplace. A name one of the marketplace's earlier approved snapshots already carried is not checked again, so a collision is raised once, when the name first arrives, and accepted once.
- Against approved snapshots of other marketplaces. First come wins: the incumbent is never re-evaluated, and nothing here can withdraw what is already served.
- Lookalikes, not near-misses. Names are compared after Unicode
normalisation, the Unicode confusable skeleton (UTS #39 —
0foro, a Cyrillicоfor a Latin one,rnform), case folding and removal of-,_,.and spaces.Claude-Skills,claude_skills,claudeskillsandcIaude-skillsare one name. There is no edit distance:code-reviewandcode-reviewsdo not collide.
A refusal is accepted with a waiver on rule plugin-name-collision — a fork is
the legitimate case — and the vetting override does not lift it. Restoring a
revoked snapshot runs the check again, so one whose name another marketplace
took while it was revoked needs that waiver first. Two approvals racing each
other with colliding names cannot both succeed: the second waits for the first
and is refused as a collision with it.
Waivers: accepted risks with a scope and an expiry¶
A waiver is an accepted-risk exception for one finding rule, on one
marketplace, within one scope, until one date. All four are mandatory,
and so are a justification and the identity accepting the risk. There is no way
to express an unlimited waiver — expires_at is NOT NULL in the schema, and an
expiry in the past, or more than 90 days ahead, is refused at creation
(refused, not shortened; a lapsed waiver is renewed, not extended).
| Scope | The scope value is | It covers a finding when |
|---|---|---|
SNAPSHOT |
the snapshot's commit SHA | the finding is on that exact commit |
PATH |
a repository-relative path | the finding's path is that path, or lies under it |
Scope is matched against the path part of a finding's location, never the
line number: inserting a line above a finding moves the number, and a waiver
that evaporates on an unrelated edit trains reviewers to re-waive without
reading. Path matching is a prefix on a segment boundary — plugins/a covers
plugins/a/x.md but never plugins/ab.md — and there is no glob syntax.
Finding groups, and the waiver on one group¶
Every finding whose location names a file of the snapshot is identified by the git blob that file has in the pinned tree. The gateway reads that identity from the tree it pinned, after the vetter answers. It is never taken from the vetter. The vetting report then shows the findings of one verdict that share rule, severity, message, blob and line as one finding group, with every location. A file vendored into nineteen plugins is one group with nineteen locations, not nineteen rows. The recorded run keeps one finding per location. Content that differs by a single byte is a different blob, so it is a different group. A finding the gateway could not tie to a blob, such as an error verdict, is never grouped.
A SNAPSHOT waiver may be narrowed to one group by naming its content (the
blob id) and line. It then covers a finding only where the rule, the blob and
the line all match, in that commit. So a new file carrying the same rule, another
line of the same file, and the same bytes in the next snapshot are all outside
it. A group waiver on PATH scope is refused, because it would outlive the
content it was judged on. This is how a vendored finding is accepted in one
step without a bulk waive: a bulk waive would accept content nobody opened.
The approval refusal and the portal name each uncovered group once, with its
rule and its path:line locations.
SNAPSHOT scope, narrowed to a group, is what the portal offers first. It dies
with the SHA, so the next ingestion blocks again and the acceptance has to be
made deliberately a second time. Without a group, a SNAPSHOT waiver accepts
every finding of its rule in that commit. A PATH waiver survives re-ingestion,
which is its purpose and also its cost — it covers content that does not exist
under that path yet. That is why an expiry is mandatory rather than advisory.
The effective outcome¶
The recorded chain run is never rewritten. It stays raw evidence of what the vetters said. What gates an approval is the effective outcome, computed on every read from that run plus the waivers active at that instant:
- a
PASSorWARNverdict stays clearing — a waiver can only ever remove an objection, never create one; - a blocking verdict with no findings stays blocking.
PENDINGcan never be waived away, because there is nothing to name; - a blocking verdict with findings is re-derived from the findings that are
left, by the same severity rule the vetter's own state came from. Waive
every
HIGH/CRITICALfinding and the verdict clears.
| Effective states | A waiver suppressed something | Outcome |
|---|---|---|
| all clearing, run non-empty | no | CLEAR |
| all clearing, run non-empty | yes | CLEAR_WITH_WAIVERS |
| anything else, or no run at all | — | BLOCKED |
CLEAR_WITH_WAIVERS is a different word from CLEAR on purpose. A reviewer or
an auditor glancing at a badge must never read an accepted risk as a clean
chain.
Expiry needs no scheduler¶
A waiver is active only while expires_at is in the future and it has not been
revoked, and that is decided at the moment the effective outcome is computed —
on the approve request, on the vetting API read, on the portal poll. So an
expired waiver stops suppressing on the very next evaluation, and the snapshot's
effective outcome reverts to BLOCKED with nothing having had to run in the
background. Revoking a waiver has the same effect immediately.
An hourly sweep writes a waiver-expired entry the first time it notices a
lapsed waiver. It has no authority over the gate — the gate is already correct
without it — so it only decides whether the lapse is announced in the ledger
rather than merely observable in it.
Expiry re-closes the gate immediately; retraction waits for a re-vet
A snapshot approved while a waiver was active stays published the moment that waiver lapses. What returns instantly is the gate: the snapshot reads as blocked again, and any future approval needs a fresh acceptance.
Taking the content back is
continuous re-vetting's job. The next re-vetting
run over that snapshot finds the finding uncovered again and reports a
violation — recorded and announced in the default warn mode, and revoking
the snapshot under enforce. So under enforcement a waiver's expiry is a
real deadline, not a reminder.
vetter-error is waivable
A vetter that crashed or timed out records a vetter-error finding,
and the uniform rule above makes it waivable like any other. That is a real
operational need — an external scanner down for a day — but it means
accepting "the scanner never looked at this". It is the single most
consequential thing a reviewer can write here, and the ledger names the rule
so it can be found.
The built-in vetters¶
All five vetters ship in the gateway and run in every chain. The first four ask whether the content is dangerous; the fifth asks whether it is well formed.
secret-scan¶
Regex and entropy rules over every UTF-8 text file in the snapshot: AWS access key ids and secret keys, PEM private-key blocks, GitHub, Slack and Google tokens, JSON Web Tokens, and assignment-shaped values whose Shannon entropy is high enough to be a real credential rather than an identifier.
Findings never echo the matched value — a finding that quoted the secret would put it in the ledger and the portal.
prompt-injection¶
Pattern heuristics over the snapshot's Markdown instruction content
(SKILL.md, commands, agents):
| Rule | What it looks for |
|---|---|
instruction-override |
"ignore all previous instructions" and its close relatives |
system-prompt-disclosure |
Asking the agent to reveal its prompt or instructions |
credential-path-reference |
~/.aws/credentials, ~/.ssh, .npmrc, /etc/passwd, … |
concealment-instruction |
Telling the agent to keep something from a person, within one clause: "do not tell / inform / mention / report / reveal … the user", "hide … from the reviewer", "without telling the user", "don't let the user know" |
pipe-to-shell |
curl … \| sh inside instructions |
exfiltration-instruction |
Sending credentials or environment values to a host |
hidden-html-instruction |
Agent-directed text inside an HTML comment |
invisible-characters |
Zero-width, bidirectional, and Unicode-tag characters used to hide text from a human reading the diff |
The rules are tuned against real skill repositories as well as against payloads.
A negation never reaches into the next sentence, so "do not silently overwrite
it. Show the user the file" is not a finding. Verbs of displaying do not
count as concealment: "do not show the user the raw JSON" is a formatting
instruction and by far the commoner use of the phrase. Command names and flags
such as ignore-rule or --all-values are not prose. The price is a known
blind spot: "do not show the user the command you ran" is not caught unless it
also says hide, without telling or don't let … know. A paraphrase walks
past pattern rules anyway, which is why this vetter is triage.
executable-surface¶
Reads each plugin's hooks, MCP servers and LSP servers and flags the code a
plugin runs without anyone invoking it. It also flags code a plugin installs
or fetches after the snapshot was pinned: dependency manifests, and installs
in its skills, commands and agents. A plugin's hooks come from its hooks/hooks.json, from the hook
files or inline hooks that its plugin.json and its marketplace entry declare,
and from the frontmatter of its skills and agents. The vetter reads the same
inventory the portal shows.
Plugins come from two places: those the manifest lists, and any directory that
holds a .claude-plugin/plugin.json. A plugin left out of the manifest is still
served, so its hooks are still read.
| Rule | What it means | Severity |
|---|---|---|
auto-run-hook |
A hook runs automatically. The finding names its trigger (the event and the tool matcher) and what it runs, at the path:line where it is declared |
medium: warns |
runtime-fetch-exec |
A hook's command, or a file of the snapshot the hook launches, downloads code and executes it. That covers three shapes: a download piped into an interpreter; an interpreter fed a download through $(…) or <(…); and a file that both downloads to a file and makes that same file executable |
high: blocks |
runtime-package-run |
A hook runs a package runner (npx, uvx, pnpm dlx, pipx run, …) or installs packages (pip install, npm install, …). Either way, it fetches its code from a registry at run time |
high: blocks |
hook-config-unreadable |
A hook file, a plugin.json or the manifest could not be parsed, so the hooks it declares could not be read |
medium: warns |
hook-target-unscanned |
A hook runs a file that is binary or over the scan size limit, so no rule could read it | medium: warns |
mcp-fetch-exec |
A local MCP server's command, or a file of the plugin it launches, downloads code and executes it, in the shapes runtime-fetch-exec recognises |
high: blocks |
mcp-package-run |
A local MCP server runs a package runner or installs packages, as runtime-package-run recognises them |
medium: warns |
lsp-fetch-exec |
An LSP server's command, or a file of the plugin it launches, downloads code and executes it, as for an MCP server | high: blocks |
lsp-package-run |
An LSP server runs a package runner or installs packages, as for an MCP server | medium: warns |
runtime-dependency |
A dependency manifest in a plugin declares dependencies; or a skill's script, or a fenced code block of a skill's Markdown or of a command or agent, installs packages or runs a package runner. The code arrives from a registry after the snapshot was pinned. One finding per rule per file, at the first line, naming every line | medium: warns |
skill-fetch-exec |
A skill's script, or a fenced code block of a skill's Markdown or of a command or agent, downloads code and executes it, in the shapes runtime-fetch-exec recognises. One finding per file, at the first line, naming every line |
high: blocks |
Why a hook only warns. A hook is a legitimate plugin feature. Its code is in the snapshot, and the rest of the chain reads it. Blocking every hook would teach reviewers to waive hooks unread. Medium puts each hook on the verdict row with its trigger, and does not hold the snapshot.
Why runtime fetching blocks. Code fetched at run time was never in the
snapshot the gateway pinned. It bypasses quarantine, vetting and approval
entirely. A reviewer who has decided that the source is trusted waives the
finding group like any other. It blocks wherever it is found: in a hook, an
MCP or LSP server, or a skill's files. On a Markdown line, skill-fetch-exec
can coincide with prompt-injection's pipe-to-shell. The overlap is
deliberate: the two vetters are switched and waived apart, and
skill-fetch-exec also sees shapes pipe-to-shell does not.
Why an MCP package runner only warns. Nearly every published MCP server is
started with npx -y or uvx, so the package's current code is fetched each
time the server starts. Blocking all of them would teach reviewers to switch
the vetter off. The warning puts each one on the verdict row. Download-and-execute
blocks, as it does in a hook. The MCP rules have their own ids, so waiving one
never waives a hook that does the same.
Why a runtime dependency only warns. Claude Code documents
${CLAUDE_PLUGIN_DATA} as the place for a plugin's installed dependencies, so
installing at run time is a sanctioned pattern, as the MCP runner is. The
finding says whether a lockfile sits beside a manifest, but a lockfile does
not lower it: a lockfile fixes versions only when the install command honours
it (npm ci, not npm i), and the code is fetched after pinning either way.
A file is one finding per rule, so a skill of scaffolding commands is one
decision for the reviewer, not one per line.
Dependency manifests. package.json (any of dependencies,
devDependencies, optionalDependencies, peerDependencies),
requirements*.txt (any requirement, -r, -c or -e line),
pyproject.toml ([project] dependencies, a poetry dependency table other
than python, [dependency-groups]) and Cargo.toml (any dependency table)
anywhere in a plugin. The finding sits at the declaring line. A manifest that
declares nothing, or that does not parse, is silent. Manifests inside
node_modules/, .venv/, venv/, site-packages/ and target/ are vendored
content and are not read, and neither are the skill files there.
Skill, command and agent files. A skill's directory is the one holding its
SKILL.md. Its scripts (a known extension, or a #! first line) are scanned
whole. Its Markdown files, and the command and agent files the inventory
lists, are scanned in fenced code blocks only (``` or ~~~, at any
indentation, so a fence nested in a list is read). Prose and inline code
spans are mentions, not commands, and are not read. A file that a hook or a
server launches is not read again here, so a line carries one finding.
LSP servers. The servers come from the plugin's .lsp.json and from the
lspServers its plugin.json and its marketplace entry declare, as a path,
an inline map or an array of both. A name declared later replaces an earlier
one. Every LSP server runs as a local process, so every one is scanned exactly
as a local MCP server is, below, under the lsp- ids. A language server on
PATH (gopls) is the user's install, and is silent.
MCP servers. The servers come from the plugin's .mcp.json and from the
mcpServers its plugin.json and its marketplace entry declare. For a local
(stdio) server, the command and its args are scanned as one command line,
and the finding is located at the line of command. Before matching:
- a
${VAR:-default}is replaced by its default, which is what runs when the variable is unset; - a runner given as an absolute path (
/usr/local/bin/npx) is read by its file name, and anenvlauncher and its assignments are dropped; - the script a shell wrapper is handed (
sh -c "…",cmd /c …,powershell -Command …) is also scanned as a command of its own.
A runner whose package is a path inside the plugin (npx ./server,
npx ${CLAUDE_PLUGIN_ROOT}/server) is not a fetch. Files of the plugin that
the server's command names are followed as a hook's are, and findings in them
carry the MCP ids and severities. A binary a server runs is not reported: a
server shipped as a compiled binary is ordinary, and its bytes are pinned in
the snapshot. A remote server (http, sse, ws) runs nothing on the user's
machine and is not scanned. Neither is an MCP bundle (.mcpb, .dxt), which
is an archive.
Precision. Only the hook's own command and the files it launches are
examined, so documentation that mentions curl is never a finding. A launched
file is one the command names that exists in the snapshot under the plugin: a
${CLAUDE_PLUGIN_ROOT}/… reference, or a path-shaped token. Launched files are
followed to a depth of three, through the files they name in turn. Comment lines
are skipped. So are a download that is never made executable, and a
command -v curl probe. Quoting inserted into a command name (c''url,
"cu"rl) is removed before matching. Several hooks that reach the same script
yield one finding per line of it, not one per hook.
What it does not see
It matches shapes. A fetch reached through variable indirection
($FETCH "$url" | sh) or an encoded payload walks past it, in a hook and
in an MCP server alike. Monitor commands are listed in the inventory but
not examined for runtime fetches. For MCP servers, a package-manager
script (bun run start, npm start) is not looked up in package.json,
and a container runner (docker run image) is not classified as a fetch,
though both can fetch code when the server starts. Hook files for other
harnesses (.codex/, .cursor/) are not Claude Code plugin hooks and are
not read.
Dependencies are reported, not resolved. No lockfile is parsed, no
version range is evaluated, and no package is looked up for known
vulnerabilities or malware, which needs an advisory database and belongs
in an external connector. An install written as an
inline code span in prose (run `npm i` first) is read as a mention and
is not reported, and neither is a command inside a string argument
(os.system("npm install")). Files under bin/ are not scanned.
license-scan¶
Deterministic license detection over the pinned content, evaluated against the
configured allow/ban lists: SPDX ids
resolved from license/copying files anywhere in the tree,
SPDX-License-Identifier tags inside them, and the marketplace manifest's
license metadata fields. Exact fingerprint matching only — no scoring:
| Rule | What it means | Blocks |
|---|---|---|
license-detected |
A license was identified (always recorded, informational) | never |
license-banned |
The license is on the configured ban list | always |
license-not-allowed |
An allow list is configured and does not contain it | always |
license-unknown |
The source identifies no known license — a first-class state, never a guess | only when an allow list is configured; warns otherwise |
license-missing |
The snapshot carries no license information at all | only when an allow list is configured; warns otherwise |
With neither list configured — the default — nothing blocks: detection is
recorded, and unknown or missing licenses warn so a reviewer sees them without
any estate being blocked by an upgrade. The vetter's recorded version
carries a digest of the policy in force, so a changed list is visible in every
run's chain identity. The task-shaped walkthrough is
License compliance for skills; the same
detection is readable per snapshot at
GET /api/v1/snapshots/{id}/licenses.
What a passing verdict does not mean
These are patterns, not understanding. An attacker who paraphrases
("disregard the guidance you were given earlier"), splits an instruction
across files, or encodes it walks past every rule above. The same is true of
secret-scan: it matches shapes, so an unshaped or wrapped credential is
invisible to it.
A PASS means "no known marker matched". It is triage that tells a reviewer
where to look first — it is not a statement that the snapshot is safe, and
reading the content is still the reviewer's job. Semantic review of skill
instructions needs an LLM review vetter, which the gateway does not ship
but an operator can add as an external connector.
skill-conformance¶
Validates every SKILL.md under a plugin's skills/ directory against a
vendored, dated copy of the Agent Skills specification.
This is the one vetter that answers a question about correctness rather than
danger: whether the skill carries the frontmatter an agent needs to load and
select it.
| Rule | What it means | Blocks |
|---|---|---|
skill-frontmatter-missing |
The file does not open with a --- frontmatter block |
only under enforcement |
skill-frontmatter-malformed |
The block is never closed, is not valid YAML, or is not a mapping | only under enforcement |
skill-field-missing |
A required field — name, description — is absent |
only under enforcement |
skill-field-invalid |
A field is present and breaks its constraint: wrong type, over its length limit, a name that is not lowercase or does not match its directory |
only under enforcement |
skill-not-scanned |
The SKILL.md was over the size limit or not valid UTF-8, so conformance could not be checked |
only under enforcement |
skill-field-unknown |
A frontmatter field the pinned specification does not define | never |
Advisory by default. With
skills-gateway.vetting.conformance.enforce
at its default false, every defect above is a MEDIUM finding: the reviewer
sees it, and nothing is blocked. A verdict covers a whole snapshot, so a
blocking default would let one malformed skill hold up every other skill beside
it — and a formatting defect is not what the gateway's blocking states are for.
Set the property to true and the same defects become HIGH, blocking and
waivable like any other finding.
skill-field-unknown stays informational under both postures. The pinned
specification is a snapshot of a document that still moves, and the
specification defines a metadata mapping precisely so clients can carry
properties it has no opinion about.
The specification is pinned, not fetched. The version in force is
agentskills-2026-08-04, transcribed into
src/main/resources/vetting/agentskills-2026-08-04.json and shipped inside the
gateway. Nothing is retrieved over the network while vetting: a chain run has to
be reproducible from the release alone, and
continuous re-vetting has to be able to say whether a
changed answer about approved content came from the content or from the rules.
The vetter's recorded version names the pin, a digest of its constraint table
and the posture in force — skill-conformance@agentskills-2026-08-04+schema-2ae36a+advisory
— so a specification bump is visible in every run's chain identity.
Upstream publishes no version number
The Agent Skills specification has no tags, no version field and no
changelog. The version the gateway records is therefore its own dated pin:
the date of the upstream commit the constraint table was transcribed from.
Provenance for the current pin, and the two places it deliberately follows
the upstream reference validator rather than the prose, are recorded in
src/main/resources/vetting/README.md.
External connectors¶
A connector is the transport that lets a vetter running outside the gateway
take part: an HTTP endpoint the gateway POSTs the snapshot's scannable content
to and reads a normalized {state, reportUrl, findings[]} back from, configured
under skills-gateway.vetting.external (see
Configuration → External connectors).
Each configured connector contributes exactly one vetter to the chain — an LLM
reviewer, a sandbox detonator, a corporate scanner — and that is the only thing
the word means here; the five built-ins are vetters with no connector.
A vetter that arrives over a connector runs in the chain at its position and is
recorded, aggregated and waivable exactly like a built-in one — its findings and
its external report link surface wherever a built-in's do. The administrative
on/off switch reaches
it on the same terms too: an administrator can disable such a vetter globally or
for one marketplace, and the run then records a DISABLED verdict in its place
rather than a silently shorter chain.
Because the endpoint is a dependency the gateway does not control, the connector
is fail-closed in the strong sense: an unreachable, slow, oversized,
unparseable, unrecognised or partial answer is an ERROR verdict, which blocks —
a broken external reviewer holds a snapshot, it never lets one through. And the
recorded state is the worse of what the endpoint declared and what its own
findings imply, so an endpoint cannot pass content its evidence condemns. An
endpoint that needs to answer later returns PENDING, which blocks until it is
resolved.
See Adding an external vetter for the wire contract and a minimal working example.
Coverage gaps are reported, not hidden¶
A file larger than the configured size limit, or one that is not valid UTF-8, is
not silently skipped. secret-scan and prompt-injection record one
informational file-not-scanned finding per reason (over the size limit, or
binary / not UTF-8). The finding says how many files it covers, names the first
twenty and counts the rest. The vetter's coverage summary also states the gap
("scanned 3210 text file(s); 51 file(s) not scanned (22 over the size limit, 29
binary)"). A verdict whose findings are all informational shows that summary on
its row rather than a count, so a pass that skipped files says so where the pass
is read. skill-conformance, which reads only SKILL.md files, records
skill-not-scanned for each skill it could not read. executable-surface
records hook-target-unscanned for a file a hook runs that it could not read,
and that one is medium, not informational: it is code that runs unattended and
that no rule has read. Informational findings do
not change the verdict, but they are visible, so "the scanner did not look at
this" is never invisible.
The one exception is deliberate: under
conformance enforcement
skill-not-scanned blocks rather than informs, because an operator who has made
conformance a publishing requirement must not have "we could not check" read as
"it conformed".
What lands in the ledger¶
Every chain run writes one entry for the run outcome (vetting-completed) to
the append-only ledger, and the ingestion run adds one entry per vetter verdict
(vetting-verdict). A re-vet adds a verdict entry only for a vetter whose
verdict changed since the snapshot's previous run — a different state, finding
count or worst severity, or a vetter version that was not in that run's chain —
and states how many in the completion entry's changed=; the verdicts
themselves are recorded for every vetter on every run and served in the vetting
report. Both kinds of entry are attributed to the system actor kind — the chain
is the gateway's own automated subsystem, not a person. A verdict entry leads
with vetter=state and then carries the finding count, the worst severity
present, and the id of the chain run, so the entry is auditable on its own; a
clean pass additionally states what the vetter examined — the files it scanned
and the rules it applied — so a pass in the ledger is never indistinguishable
from a vetter that did not run. The completion entry carries the same run id,
so a run's scattered verdict entries reassemble into the one run they came from.
The whole waiver lifecycle lands there too:
| Event | Written when | Detail carries |
|---|---|---|
waiver-created |
a risk is accepted | rule, scope, expiry |
waiver-applied |
a waiver lets an approval through | waiver id, rule, location, approver, expiry |
waiver-revoked |
a waiver is withdrawn | rule, scope |
waiver-expired |
the sweep first notices a lapse | rule, scope, approver, expiry |
An auditor asking "why is a snapshot with a critical finding being served" can answer it from the ledger alone — including what was accepted, by whom, and until when.
Reading further¶
- Approving and rejecting snapshots — the reviewer's task, end to end.
- Waiving a vetting finding — accepting a risk, end to end.
- Admin portal — where the verdicts appear. The chain above is drawn there as a flow per snapshot, with a node per step that opens its own evidence; an administrator gets the same flow per marketplace showing which vetters actually run for it and why.
- Configuration — the knobs.
- Trust boundaries — why approval is the boundary the chain protects.