CI Runner
Every job in .github/workflows/ci.yml and .github/workflows/release-images.yml runs on the self-hosted runner labelled zcloud-runner-ex44 (a Hetzner EX44: Intel i5-13500, 14 cores / 20 threads, 64 GB RAM, 2× 512 GB NVMe, Ubuntu 26.04 with glibc 2.43). zaflun/lumio is private, so jobs on a self-hosted runner bill no GitHub Actions minutes. Every other workflow (docs deploys, npm publish, runtime-base build, security audit, …) still runs on ubuntu-latest.
Runner selection and rollback
Each ci.yml job carries
runs-on: ${{ vars.CI_RUNNER || 'zcloud-runner-ex44' }}
The job targets the runner's custom label alone. It does not also require self-hosted / linux / x64, because a runner registered with --no-default-labels would never match those.
Rollback without a PR: set the repository variable CI_RUNNER to ubuntu-latest (Settings → Secrets and variables → Actions → Variables). The next run uses GitHub-hosted VMs again. Delete the variable to return to the EX44.
release-images.yml has its own variable, RELEASE_RUNNER, with the same default and the same rollback (ubuntu-latest). It is separate from CI_RUNNER so that releases can move back to hosted VMs without touching PR CI. Releases use the same zcloud-runner-ex44 instances as PR CI; there is no dedicated release instance. See Release images.
Hosted vs. self-hosted behaviour
.github/actions/setup-bazel branches on runner.environment. The hosted path behaves as it always has. The self-hosted path is built for a persistent host:
| Concern | ubuntu-latest (hosted) | zcloud-runner-ex44 (self-hosted) |
|---|---|---|
| Swap | 4 GB /mnt/swapfile per job | none (64 GB RAM) |
| System packages | apt-get install libssl-dev clang lld per job | none; a prerequisite check fails with one ::error:: listing everything missing |
| Bazelisk | preinstalled on the image | bazel-contrib/setup-bazel downloads Bazelisk 1.29.0 into the tool cache |
--jobs / --local_resources | --jobs=2, ram=6144 | --jobs=HOST_CPUS*.3, cpu=HOST_CPUS*.3, ram=HOST_RAM*.25 (6 jobs, ~16 GB per instance on the EX44, sized for 6 instances) |
| Local disk cache | off (--disk_cache=) | ~/.cache/bazel-disk, capped at 150 GB (pruned to 120 GB, least recently used first) |
| Repository cache | restored from the Actions cache | ~/.cache/bazel-repo, persistent |
| BuildBuddy remote cache | read in every job; Rust Build / Rust Tests upload in-org | the same, plus Schema Check and Clippy upload in-org, so the six instances share their results |
| API test output root | moved to /mnt via ~/.bazelrc | unchanged (no sudo, no ~/.bazelrc append) |
| DB-backed API tests | --jobs=2 limits them to two binaries | API Integration Tests raises the service's max_connections to 400 and turns durability off (fsync, synchronous_commit, full_page_writes — as on the host test stack; with fsync on, each per-test template clone fsyncs a whole database and the binaries time out) via ALTER SYSTEM + container restart, then re-reads the new host port, and runs 12 binaries × 4 threads; API tsdb Migration Test stays at 4 × 4 under the default 100 connections |
On both paths:
- The BuildBuddy key never lands in
$HOME. It is written to$RUNNER_TEMP/buildbuddy-auth.bazelrc(mode0600), which~/.bazelrconlyimports by path. The runner empties$RUNNER_TEMPat the start and end of every job. - Every setup-bazel job ends with a teardown step (
.github/actions/setup-bazel/teardown.sh,if: always()). It stops the Bazel server, redacts the key from Bazel'sjava.log(the Bazel client hands every rc option to the server, and the server logs it), deletes the auth rc and, on the self-hosted path, caps the disk cache. Then it fails the job if the key is still in any file under$HOME, or if~/.bazelrcgained extrastartuplines. - Scratch files live under
$RUNNER_TEMP, never under fixed/tmppaths. - The postgres / TimescaleDB service containers publish
5432on a dynamic host port;DATABASE_URLreads it fromjob.services.<name>.ports['5432'].
The ZAF-1693 split (api_lib, then api_test, then the rest, as separate Bazel invocations) stays on both paths.
BuildBuddy keys include the C/C++ toolchain, which Bazel configures from the host's clang. Results built on ubuntu-latest therefore do not match the EX44's, and the reverse is also true. The remote cache is shared between the EX44 instances, not between the EX44 and hosted runners.
Because the instances run as different OS users, a cached action must not write to an absolute path baked in by whoever built it first. sailfish's template derive did exactly that (its proc-macro writes to the build-time OUT_DIR, inside the building user's output base), so vendor/BUILD.bazel builds sailfish-compiler with its hermetic feature under Bazel. That feature compiles templates in memory. If a crate fails on the EX44 with PermissionDenied under /home/<other runner user>/, fix that crate the same way.
Release images
release-images.yml signs and pushes production images from the same persistent host that runs PR CI, as the same OS users. A PR job can write to those users' $HOME and holds a writable BuildBuddy key, so the release builds with nothing inherited from earlier jobs:
- Run-scoped homes. The first step of both jobs points
DOCKER_CONFIG,RUNNER_TOOL_CACHEandBAZELISK_HOMEat$RUNNER_TEMP, and cosign installs to$RUNNER_TEMP/cosign. Docker CLI plugins and credential helpers, the GHCR login, Bazelisk, Node, pnpm and cosign therefore never come from, or land in, the persistent$HOME. Build & Push asserts thatdocker buildxresolves to a system plugin directory and that Bazel's install base is under$RUNNER_TEMP. - Bazel runs isolated. Both jobs call
setup-bazelwithisolated: "true": no local disk cache (--disk_cache=), and an output base, install base (startup --install_base) and repository cache under$RUNNER_TEMP, with--repo_env=DOCKER_CONFIGso the@runtime_basepull uses the run's login. Their teardown (BAZEL_ISOLATED=true) runsbazel clean --expungeand deletes all three directories. - BuildBuddy stays on. The release reads the remote cache like every CI job and uploads its own
--config=release(opt) results (--remote_upload_local_results). In-org PR jobs hold a writable key, so a remote action result is trusted the same way as on PR CI (founder decision, ZAF-1788). - Web images build in a BuildKit builder of their own. The job creates
lumio-release-<run id>-<attempt>(docker-containerdriver, BuildKit on a fixed version tag), pushes from it directly with--provenance=false --sbom=false(a plain image manifest; provenance comes from cosign), prunes it after each app and removes it in anif: always()step. - Sign the built digest, then promote. Each image is pushed under its full-SHA tag only. cosign signs and attests the digest the build reported (buildx
--metadata-file; for Rust the rules_imgimage_manifestdigestoutput), after checking that the registry's:<full sha>agrees. Only then, as the last action of the step, do the other tags move. A promote fetches the manifest by digest andPUTs the same bytes under each tag through the registry API, so the tag points at exactly the signed digest.docker buildx imagetools createis not used for this, because it wraps a single manifest in a new index with a different digest. A run that stops before the promote leaves every floating tag on the previous signed digest. - One release at a time. Build & Push is in the concurrency group
release-imageswithcancel-in-progress: false.all,all-rustandapipush the same images, so a per-app group would not serialise them. GitHub keeps one pending run per group. - Least privilege. Verify runs with
contents: readandpackages: read; only Build & Push getspackages: write,id-token: writeandcontents: write. Both check out withpersist-credentials: false. - Every job has
timeout-minutes(Verify 90, Build & Push 180).
The tag set (full release: version, YYYY.M, YYYY, latest, short SHA, full SHA; pre-release: version, beta, short SHA, full SHA), the keyless cosign signature, the SLSA provenance attestation (both by digest) and the schema upload are the same on both runner kinds.
No runner user holds host root. Each one runs its own rootless Docker daemon instead of being in group docker, so a PR job cannot reach the host or another instance through Docker (docker run -v /:/h alpine ls /h/root fails). The measures above protect a release from leftovers of earlier jobs on the same user. A PR job on the same instance can still read that user's runner credentials; the founder accepted that residual risk (ZAF-1791).
glibc guard
Bazel links the Rust binaries against the build host's glibc, but they run on runtime-base's glibc. If the host's is newer, the container dies at start with version `GLIBC_2.xx' not found, and neither Bazel nor the push notices. Before it pushes anything, Build & Push therefore builds every Rust _push target, reads the newest GLIBC_ symbol version the packed binary needs (objdump -T) and compares it with getconf GNU_LIBC_VERSION in ghcr.io/zaflun/lumio/runtime-base:latest. If the base is older, the job fails with a glibc guard error and pushes nothing.
runtime-base is Ubuntu 26.04 (glibc 2.43), the same release as the EX44; an api built there needs GLIBC_2.43 (atan2f). When the runner's OS moves, move docker/runtime-base.Dockerfile to the same release in the same change. latest is only re-tagged from main, so the base must reach main and build-runtime-base.yml must run there before releases from the new runner can pass the guard. A rollback to ubuntu-latest (glibc 2.39) always passes.
Runner prerequisites
The runner is provisioned once, by hand. CI installs nothing with sudo, and no step on the self-hosted path calls sudo, so the runner user does not need passwordless sudo.
-
GitHub Actions runner as a systemd service (
./svc.sh install <user>) under a dedicated, unprivileged user, registered with the labelzcloud-runner-ex44. -
Rootless Docker per runner user. The user is not in group
docker(that group is host root) and has no sudo. Each user runs its own daemon (dockerd-rootless-setuptool.sh install,loginctl enable-linger <user>, a subuid/subgid range), and the runner's.envsetsDOCKER_HOST=unix:///run/user/<uid>/docker.sock. Packages:docker-ce-rootless-extras,uidmap,dbus-user-session,slirp4netns. A systemd drop-in on the runner unit waits for that socket after a reboot. The API test jobs startpostgres:18-alpine/timescale/timescaledb:latest-pg18service containers, and setup-bazel runsdocker/login-actionfor the@runtime_basepull; both work unchanged under rootless.Each user gets its own host port block for published container ports:
61000 + 600 × Nto+599, where N is the instance number (zaflun0,zaflun11, …). Every rootless daemon allocates dynamic ports on its own, so without disjoint blocks two service jobs on different instances pick the same port and one fails withRootlessKit PortManager.AddPort(): … bind: address already in use. The block is set by a drop-in on the user'sdocker.service(DOCKERD_ROOTLESS_ROOTLESSKIT_DETACH_NETNS=false, plus aDOCKERDwrapper that setsnet.ipv4.ip_local_port_rangein the daemon's network namespace beforedockerdstarts). The blocks sit above the host's ephemeral range (32768–60999). -
Packages (Debian / Ubuntu):
sudo apt-get install -y \git curl ca-certificates jq unzip tar zstd gzip xz-utils \build-essential clang lld libssl-dev pkg-config python3The Docker Engine must include the
buildxplugin (docker-buildx-plugin);release-images.ymlcreates a BuildKit builder with it.objdump(frombinutils, part ofbuild-essential) is used by the release glibc guard..bazelrcforcesCC=clang, links with-fuse-ld=lldand takes OpenSSL from/usr/include+/usr/lib/x86_64-linux-gnu.
Not needed on the host:
- Bazelisk —
bazel-contrib/setup-bazeldownloads it (bazelisk-version). - rustup —
dtolnay/rust-toolchain(Rustfmt job) bootstraps rustup viash.rustup.rswhen it is missing. It needscurl. - Node 24 / pnpm —
actions/setup-nodeandpnpm/action-setupdownload them.
These actions write to the runner's tool cache, RUNNER_TOOL_CACHE (default <runner dir>/_work/_tool), which must stay writable by the runner user.
Disk: plan for ~200 GB free on the runner user's home filesystem: the Bazel output base ~/.bazel (~25 GB with the API integration test binaries), the disk cache (≤ 150 GB), the repository cache, plus the pnpm store and Docker images.
The prerequisite check in setup-bazel tests exactly this list (commands, OpenSSL headers and library, and that the Docker daemon is reachable and rootless) and names the missing apt package for each.
More than one runner instance
One runner instance runs one job at a time. The EX44 runs 6 instances (zcloud-runner-ex44, zcloud-runner-ex44-1 … -5), all with the same label, so GitHub spreads the ~25 ci.yml jobs across them. Give each instance its own OS user: setup-bazel rewrites ~/.bazelrc on every job, and two instances sharing a $HOME would overwrite each other's rc and share one Bazel output base. Separate users also get separate disk caches.
The per-instance Bazel fractions in setup-bazel (HOST_CPUS*.3, HOST_RAM*.25) assume 6 instances, of which up to 4 run Bazel at once next to the Node jobs. At HOST_CPUS*.5 the host ran 30–40 Bazel actions on 20 threads and starved Vitest. Change the instance count and you need to change those fractions too. On the EX44 the TypeScript Tests job also runs Vitest with --testTimeout=20000 instead of the 5 s default.
Disk housekeeping
Cleaned automatically:
$RUNNER_TEMP: the runner empties it before and after every job.- The checkout:
actions/checkoutrunsgit clean -ffdxin it. - Service containers and their network: the runner removes them at job end.
- The Bazel disk cache:
teardown.shcaps it per user.
Not cleaned automatically, and growing per runner user: ~/.bazel (output base), ~/.cache/bazel-repo, the pnpm store, _work/_tool, and old Docker images. Prune them with a root cron job on the host (/usr/local/sbin/ci-runner-cleanup.sh, five times a day on the EX44), not from a CI job, because a job would hit the running jobs of the other instances. The cron job should:
- find the runner users as the owners of
/home/*/*/.runner(the file./config.shcreates), - skip every user whose
Runner.Workerprocess is running, - delete the whole
~/.cache/bazel-repotogether with~/.bazelwhen the repository cache is larger than 15 GB. Never delete individual files in it by age: since Bazel 8 it also holds the extracted contents of external repositories (contents/), whose files keep the archive's old timestamps. An age-basedfind -deletestripsBUILDfiles out of those repositories, and every later build fails with'…' is not a package. - run
pnpm store prune, - delete it the same way when
~/.bazelis larger than 35 GB, - prune each user's own Docker daemon (its store is
~/.local/share/docker, so image storage exists once per instance): stopped containers, volumes, and images and build cache older than 3 days; when the store is larger than 15 GB,docker system prune -af --volumesfor that user.