Skip to content

Deploying on Kubernetes

The chart in helm/skills-gateway/ installs the gateway as a single Deployment, a Service, and — optionally — an Ingress, a ServiceAccount and a ConfigMap of application configuration. This guide covers what the chart needs from you before it will install, what it deliberately refuses to guess, and a worked example on serverless Kubernetes.

For any other runtime — plain Docker, a container platform such as ECS or Cloud Run, Nomad, a systemd unit — see Deploying without Kubernetes.

Prerequisites

What Why Notes
A PostgreSQL database Snapshots, the audit ledger, tokens and grants live there The chart does not bring one. Create the database and a Secret with key password.
An OIDC client The whole web surface authenticates with OIDC Client id, client secret, the three endpoint URIs, and the provider's issuer (oidc.issuer, which the chart requires). See Identity providers.
Access to the image ghcr.io/skillsgateway/skillsgateway, or your own mirror For a private mirror, create a kubernetes.io/dockerconfigjson Secret and name it in imagePullSecrets.
A storage decision The chart will not install without one See Storage — this is the one value that has no default.

Installing

helm install skills-gateway ./helm/skills-gateway -f my-values.yaml

A minimal my-values.yaml:

image:
  repository: ghcr.io/skillsgateway/skillsgateway
  # A released version, or an image digest — see the container image reference.
  tag: "<released-version>"

postgresql:
  host: postgres.example.com
  existingSecret: skills-gateway-db      # key: password

oidc:
  clientId: skills-gateway
  existingSecret: skills-gateway-oidc    # key: client-secret
  authorizationUri: https://idp.example.com/oauth2/v2.0/authorize
  tokenUri: https://idp.example.com/oauth2/v2.0/token
  jwkSetUri: https://idp.example.com/discovery/v2.0/keys
  issuer: https://idp.example.com/v2.0

persistence:
  mode: existingClaim
  existingClaim: skills-gateway-data

Configuring the application

The chart's own keys cover the database, the identity provider and storage. Everything else the gateway reads — see Configuration — is set through two passthroughs, so no chart change is needed to reach a setting.

Name an administrator, or the pod will not start

Authorization is always enforced, and a gateway whose configuration grants the admin role to nobody refuses to start rather than run as an estate nobody can administer. Name one in skills-gateway.roles.admins, in a skills-gateway.roles.mappings entry resolving admin, or in a declared skills-gateway.estate.grants entry.

A grant made later through /api/v1/roles does not satisfy this: it is revocable through the API, so it says nothing about whether the next start will have an administrator. The refusal names every path that resolves it.

Structured configuration: config

Anything under config is rendered verbatim into a ConfigMap, mounted at /etc/skills-gateway/application.yaml, and layered over the application.yaml built into the image through SPRING_CONFIG_ADDITIONAL_LOCATION. The image keeps the defaults; this file overrides them.

config:
  skills-gateway:
    roles:
      admins:
        - platform-admin@example.com
    estate:
      marketplaces:
        - name: internal-skills
          url: https://git.example.com/skills/internal.git
      grants:
        - principal: skills-reviewers
          role: approver
          marketplace: internal-skills

This is the place for the declarative estate, which is nested lists and expresses poorly as environment variables. It matters more than it looks: the admin API requires an interactive OIDC session, so estate YAML is currently the only way to configure a gateway from automation at all (#128 tracks changing that).

The estate is reconciled at startup, before the web surface serves its first request, so the declaration is in force from the first login. The chart hashes config into a pod annotation, which means editing it and upgrading rolls the pods — without a restart a changed ConfigMap would be refreshed on disk and never read.

A ConfigMap is not a Secret

Never put a webhook secret, an audit-sink credential or anything else sensitive in config. Use extraEnv with valueFrom.secretKeyRef, or extraEnvFrom with a secretRef, and reference the environment variable from the estate declaration as described in Declarative estate configuration.

Single settings and secrets: extraEnv, extraEnvFrom

extraEnv:
  - name: SKILLSGATEWAY_RETENTION_ENABLED
    value: "true"
  - name: SGW_ESTATE_CI_BOT_SECRET
    valueFrom:
      secretKeyRef:
        name: skills-gateway-estate
        key: ci-bot-secret

extraEnvFrom:
  - secretRef:
      name: skills-gateway-extra

Environment names are Spring's relaxed-binding form of the property: skills-gateway.roles.claim is SKILLSGATEWAY_ROLES_CLAIM.

Behind the Ingress: forwardHeadersStrategy

The application serves plain HTTP and terminates no TLS, so where the Ingress (or a load balancer in front of it) terminates it, the application learns the scheme and host the outside world sees from the X-Forwarded-* headers that proxy sends. The chart names how far those headers are believed:

forwardHeadersStrategy: native   # the default

native honours the headers only from a peer in a private address range, which is where an ingress controller sits; framework honours them from any peer, for a proxy outside those ranges; none ignores them. A value the gateway does not know fails the render. The choice and its consequences are spelled out in Running behind a proxy.

Without a working strategy the OIDC redirect URI becomes http://<pod>:8080/login/oauth2/code/idp rather than the https:// URI registered with the provider, and every login fails on a redirect-URI mismatch. oidc.redirectUri states the registered URI outright, trusting no header, should you ever need to bypass the derivation.

Remove SERVER_FORWARDHEADERSSTRATEGY: framework from extraEnv

Earlier versions of this guide set it there. On the GraalVM native image the release used to publish, that value was inert — Spring Boot registers the filter behind a condition the image evaluates at build time — and setting it also switched off the native strategy Spring Boot would otherwise have deduced on Kubernetes, so following the advice made logins fail rather than fixing them. The gateway registers the filter itself now, and the release artifact is a JVM container, so both halves of that are gone. The chart value replaces the extraEnv entry either way; an extraEnv entry of the same name would override the chart's.

Storage

The gateway keeps three sets of repositories on disk under /data: the quarantine, the published content the facade serves, and hosted first-party marketplaces. PostgreSQL records which snapshots exist and which are approved.

Those two halves are one estate, and only one of them survives a restart onto an empty volume. A gateway that comes back with an empty /data still reports its snapshots as published, still lists them in the catalog, and can serve none of them — with no rehydration path, because a published snapshot is a pinned history, not something re-fetchable from upstream.

So the chart has no storage default and refuses to render without a choice:

persistence:
  mode: existingClaim          # durable
  existingClaim: skills-gateway-data
persistence:
  mode: ephemeral              # emptyDir; everything is lost on restart

Any other value — including leaving mode empty — stops helm install with a message explaining what is lost. Choosing ephemeral is fine for a demo or a throwaway environment; it is a deliberate act either way.

The claim must already exist in the release namespace. The chart does not create a PersistentVolumeClaim, because the interesting decisions — storage class, access mode, size, static or dynamic provisioning — belong to the cluster, not to this chart. Which of them are even available to you is constrained by the platform — see Storage options on serverless Kubernetes.

replicaCount depends on the storage backend

On the default filesystem backend, storage is a plain filesystem written by the ingestion pipeline, the approval step and the facade's own repacking. There is no cross-pod locking, so two replicas writing the same volume can interleave a fetch and a publish into a corrupt repository. The chart refuses to render replicaCount > 1 there rather than trusting you to know that.

Keep replicaCount: 1. A rolling update briefly runs two pods, so prefer the Recreate strategy — or accept a moment's overlap — on volumes that allow only one writer anyway.

On the object-store backend concurrent writers are safe by construction and more than one replica is supported as it renders: the scheduled background passes take a lease apiece, so each one still runs once per interval however many replicas there are. See Running more than one replica.

Ingress and TLS

The Ingress is off by default:

ingress:
  enabled: true
  className: nginx
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt
  hosts:
    - host: skills-gateway.example.com
      paths:
        - path: /
          pathType: Prefix
  tls:
    - secretName: skills-gateway-tls
      hosts:
        - skills-gateway.example.com

TLS is not optional in production

The git facade authenticates clients with personal access tokens over HTTP Basic — see the facade reference. Every clone and every fetch puts a long-lived credential on the wire. Without TLS, anything on the path can read it and then act as that client indefinitely.

Terminate TLS at the Ingress or at a load balancer in front of it, and never expose the container's HTTP listener directly.

Both the portal and /git/** are served by the same container on the same port, so one host and one / path is the normal configuration.

Raise the read timeout. A first ingest of a repository with a long history runs inside one request for minutes, and the ingress-nginx default of 60 seconds cuts the response off while the ingest carries on (why). With ingress-nginx:

ingress:
  annotations:
    nginx.ingress.kubernetes.io/proxy-read-timeout: "600"
    nginx.ingress.kubernetes.io/proxy-send-timeout: "600"

Other controllers have an equivalent; on the AWS Load Balancer Controller it is alb.ingress.kubernetes.io/load-balancer-attributes: idle_timeout.timeout_seconds=600.

Identity and security context

The chart creates a ServiceAccount for the release and runs the pod under it, because workload identity binds to a named account:

serviceAccount:
  create: true
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::000000000000:role/skills-gateway

By default the container runs as the image's non-root user (uid 65532), with all capabilities dropped, privilege escalation disabled and a read-only root filesystem; the chart mounts an emptyDir at /tmp so the runtime keeps somewhere to write. podSecurityContext.fsGroup is what makes the mounted data volume writable by that user.

The embedded git library (JGit) needs to know about that scratch space too: it caches a filesystem-timestamp attribute in a user-level configuration file under $XDG_CONFIG_HOME, and with neither that variable nor HOME otherwise set, it would resolve to a home directory this user does not have on a sealed root filesystem — failing on every fetch with a caught (non-fatal) error that still drowns out real ones. The chart sets XDG_CONFIG_HOME=/tmp/xdg-config in the Deployment for exactly this reason.

Restricting egress

The gateway makes outbound connections to URLs an administrator or an upstream chooses: a marketplace's upstream, an external plugin source, a webhook subscriber, an audit sink, the forge mirror. The in-application address policy (trust boundaries) refuses private and link-local addresses for external plugin sources only. A marketplace upstream, a webhook subscriber URL and the mirror URL are checked for their scheme and nothing else. ADR 0011 names network topology as the primary control, and in a cluster that control is an egress NetworkPolicy. The chart does not ship one, because the right allowlist depends on where your database, object store and identity provider live.

This policy selects the gateway's pods and permits DNS, PostgreSQL, and HTTPS to public addresses. It refuses the cloud metadata endpoint, RFC 1918, carrier-grade NAT and their IPv6 counterparts:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: skills-gateway-egress
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: skills-gateway
      app.kubernetes.io/instance: skills-gateway   # the Helm release name
  policyTypes:
    - Egress
  egress:
    # DNS, to the cluster resolver only.
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns
      ports:
        - { protocol: UDP, port: 53 }
        - { protocol: TCP, port: 53 }
    # PostgreSQL. Name its address; it is usually private, so the rule below
    # would refuse it.
    - to:
        - ipBlock:
            cidr: 10.20.30.40/32
      ports:
        - { protocol: TCP, port: 5432 }
    # HTTPS to public addresses: forges, the identity provider, the object
    # store's public endpoint, webhook subscribers.
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
            except:
              - 10.0.0.0/8
              - 172.16.0.0/12
              - 192.168.0.0/16
              - 100.64.0.0/10
              - 169.254.0.0/16    # link-local, the cloud metadata endpoint
        - ipBlock:
            cidr: ::/0
            except:
              - fc00::/7
              - fe80::/10
      ports:
        - { protocol: TCP, port: 443 }

Adapt it to the estate:

  • Anything private the gateway must reach gets its own rule with the narrowest block that holds it: an in-cluster database (a podSelector or namespaceSelector rule instead of an ipBlock), an object store reached through a private endpoint, an internal forge, an OpenTelemetry collector. Never widen the public rule to cover them.
  • Plain http. skills-gateway.allowed-url-schemes permits http by default. The policy above allows port 443 only, so an http upstream fails to connect. Set allowed-url-schemes: [https] as well, and registration refuses such a URL with a reason instead.
  • An egress proxy. Where egress already goes through a proxy, allow the proxy's address and nothing else, and let the proxy hold the destination allowlist.

A policy the network does not enforce does nothing

A NetworkPolicy is enforced by the cluster's network plugin, and a plugin that does not implement it accepts the object and ignores it. Some platforms do not enforce it for every node type, serverless node pools included on some providers. Use the platform's own control there, such as security groups or a firewall on the egress path. Either way, check the result with the probe below.

Verify it from inside the pod. The image has no shell, so attach an ephemeral debug container: it shares the pod's network namespace, and the policy applies to it.

kubectl debug -it <gateway-pod> --image=curlimages/curl -- \
  curl -sS -m 5 http://169.254.169.254/

A timeout is the expected answer. Any response means the metadata endpoint is reachable from the gateway.

Storage options on serverless Kubernetes

Serverless node pools — AWS Fargate and its equivalents — constrain storage sharply, and the constraints interact. This is the short version of the choice, for the case where the platform is already fixed:

Option Available? Verdict for this workload
Block storage (RWO) No Not a trade-off — there is no configuration that gets you one.
Network filesystem (RWX, e.g. NFS/EFS) Yes, static provisioning only Works. Good for a proof of concept; the substrate git is worst on.
Parallel filesystem (e.g. FSx for Lustre) No Unavailable on serverless pods.
Object storage Supported The production answer here: storage.backend: object-store, credentials from workload identity, and no volume required at all — see the storage guide.
Block storage on managed nodes Yes, if you can choose the platform What today's storage was designed for.

Block storage is unavailable, not merely awkward. On EKS, AWS states plainly that "You can't mount Amazon EBS volumes to Fargate Pods", and the reason is structural: the EBS CSI node component is a DaemonSet, and "Daemonsets aren't supported on Fargate" — the CSI controller can run there, the node component cannot. There is no ReadWriteOnce volume to be had, so the single-writer filesystem this gateway was built around has nowhere to live.

A network filesystem works, with static provisioning only. A Fargate pod mounts EFS without any driver installation, but "You can't use dynamic persistent volume provisioning with Fargate nodes, but you can use static provisioning" — so the PersistentVolume and its access point must exist before the pod does, which is what the worked example below creates by hand.

Two caveats worth understanding before choosing it:

  • Git is a bad fit for a network filesystem. Git's on-disk format is a random walk over large packfiles: object lookup seeks into a pack, follows a delta chain, and does it again — plus constant small metadata operations (stat on loose objects, lock files, fsync on refs). On a local disk the page cache absorbs all of that. Over NFS each one is a network round trip. Expect clone and repack times in multiples, not percentages.
  • RWX does not enforce the single-writer assumption. A network filesystem is ReadWriteMany by nature, so nothing stops a second pod from mounting the same volume and writing to it. The filesystem backend has no cross-pod locking, so replicaCount: 1 holds by convention here rather than by construction — an RWO block volume would have refused the second writer for you.

A parallel filesystem is not an option either: FSx for Lustre is listed as unavailable to Fargate pods.

Object storage is the answer on serverless platforms specifically. Not because of scale — because the alternatives here are "impossible" and "the substrate git is worst on". It is implemented: set storage.backend: object-store, and the repositories live in a bucket with local disk as a cache. Fargate has no instance metadata service, so credentials come from workload identity — see Choosing and migrating the storage backend.

If the platform choice is still open, an ordinary managed node group with an RWO block volume is what the current storage implementation was designed for, and it enforces the single-writer property structurally rather than by convention. Serverless is worth its constraints for plenty of workloads; a git server on a network filesystem is not the case it is best at.

Two operational notes that are easy to miss

  • No instance metadata service. Serverless pods typically cannot reach IMDS, so cloud credentials come from workload identity bound to a named service account (IRSA on EKS). That is why the chart creates a ServiceAccount and lets you annotate it; on EKS the documented remedy for a Fargate pod that needs IAM credentials is exactly that.
  • Egress needs a NAT path. Fargate pods run in private subnets only, with no direct route to an internet gateway. Ingestion fetches from upstream git over the network, so without a NAT gateway (or equivalent egress) registration and sync will hang rather than fail quickly. Inbound traffic likewise arrives through a load balancer you place in front of the pods. Restrict that egress as Restricting egress describes.

Worked example: a statically provisioned network filesystem

Create the access point and the PersistentVolume by hand (ids below are placeholders):

apiVersion: v1
kind: PersistentVolume
metadata:
  name: skills-gateway-data
spec:
  capacity:
    storage: 100Gi           # ignored by the EFS driver; required by the API
  volumeMode: Filesystem
  accessModes:
    - ReadWriteMany
  persistentVolumeReclaimPolicy: Retain
  storageClassName: ""       # static binding: no class, no provisioner
  csi:
    driver: efs.csi.aws.com
    volumeHandle: fs-0123456789abcdef0::fsap-0123456789abcdef0
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: skills-gateway-data
  namespace: skills-gateway
spec:
  accessModes:
    - ReadWriteMany
  storageClassName: ""
  resources:
    requests:
      storage: 100Gi
  volumeName: skills-gateway-data

Give the access point a POSIX owner of 65532:65532 so the non-root container can write to it, then install with persistence.mode: existingClaim and persistence.existingClaim: skills-gateway-data.

This is a proof-of-concept substrate

It works, and it is not what you want in production — see the comparison above. It is a reasonable substrate for a proof of concept, a pilot, or a small internal estate. The production answer is the object-store backend, which stops treating a POSIX filesystem as the source of truth; moving to it is an offline, verified, reversible copy described in Choosing and migrating the storage backend.

Verifying the install

kubectl -n skills-gateway rollout status deploy/skills-gateway
kubectl -n skills-gateway logs deploy/skills-gateway | grep estate

Then log in to the portal and confirm the estate reconciled — GET /api/v1/estate reports the last run, and the audit ledger carries its entries under the config-reconciler principal. See Declarative estate configuration.

Metrics and traces

A default install records its metrics and exports none of them. The chart's otel values switch on OTLP export to your collector:

otel:
  enabled: true
  endpoint: http://otel-collector.observability:4318

There is no metrics endpoint to scrape. The metric names, and the two worth alerting on, are in Observability.

Backups and upgrades

The database and the git storage are one estate and have to be backed up and restored together — see Backing up, restoring and upgrading, which also covers what startup does to the schema and what a rollout does to an in-flight fetch.