Infrastructure Management
Quick Reference
Full-fleet deployment: required mass-deploy skill
REQUIRED SUB-SKILL: Any "deploy everything" request (including
/update-all-remote) MUST use the project-local mass-deploy skill. It is the
authoritative full-fleet coordinator contract for normal Pi generalist
subagents; do not replace it with ad-hoc Colmena calls.
The skill enforces this deployment contract:
- Hosts are applied one machine at a time, in canonical order: stormwind → ironforge → orgrimmar → anton → gnomeregan → headscale ("gateway").
- After each clean switch, every documented web endpoint in the fleet (the
tables in
references/host-mapping.md) is probed and must return its expected status (default: 2xx after redirects). The legacy-namedheadscalegateway has zero web endpoints; successful private Hetzner application probes indirectly exercise its Tailscale subnet route and DNS. - Any repaired deploy/build/activation problem, repository or remote-state change, or endpoint heal restarts the sequence from stormwind (max 3 full restarts). The run never advances directly after a repair.
- A pre-flight worker blocks the run on untracked
*.nixfiles (invisible to the git+file flake). Unreachable machines never block: a down host is recorded as skipped, its endpoints are excluded from health checks, and the result ispartial. A sleeping laptop can be deployed later with an ad-hoc single-host run. - Advisory container-upgrade preview and workaround audit workers persist
their results for the final report. The audit follows "Workaround Hygiene"
below: only
WORKAROUND(-tagged overrides underoverlays/are stock-tested at the pinnednixpkgs-unstablerevision; outside markers are reported, not modified.
If the run aborts, surface its stop point, timeline, and next manual action rather than silently retrying or claiming partial work succeeded.
Ad-hoc single-host deploys: the colmena-deployer subagent
For one-off work on a single host (build checks, a quick iteration loop on one
machine), run the colmena call through the colmena-deployer subagent — one
subagent per colmena call. A deploy emits thousands of lines of flake-lock diff
and per-host build output; running it in the main context buries everything
else. The subagent runs the command, watches it to completion, and returns just
a pass/fail summary (plus the root-cause lines on failure). Do not call
colmena directly from the main context.
Spawn it with the Task tool, for example:
Use the colmena-deployer subagent: "Run
colmena apply --on gnomeregan --impurefrom/Users/fdrake/nix. Return the host's activation result. Note: anton can exit 4 on a spurious user dbus-broker reload timeout even when the switch succeeded — verify its current generation against the built path rather than trusting the exit code."
The underlying commands the subagent runs:
colmena apply --on <hostname> --impure # single host
colmena build --on <hostname> --impure # build only, no deploy
After an ad-hoc apply, verify the host's web endpoints from the tables in
references/host-mapping.md before declaring success.
Server Inventory
Hetzner Servers (Colmena-managed, root user)
| Host | Type | Services | |------|------|----------| | headscale (aka "gateway") | Hetzner VPS | Official Tailscale SaaS subnet/DNS gateway advertising 10.1.0.0/16 (all tailnet access to the other Hetzner boxes rides through it); despite the legacy name, it runs no Headscale control plane or web endpoint | | ironforge | Hetzner dedicated | media stack, all podman: jellyfin, seerr (+ jellyseerr redirect), sonarr, radarr, lidarr, prowlarr, sabnzbd, bazarr | | orgrimmar | Hetzner dedicated | gitea (+ gitea-status), woodpecker, paperless (+ paperless-ai), calibre-web, resume, filebrowser | | stormwind | Hetzner dedicated | traceway (observability stack), gatus (internal uptime dashboard) |
LAN NixOS Hosts (Colmena-managed, fdrake user with sudo)
| Host | Type | Services |
|------|------|----------|
| gnomeregan | Home LAN x86_64 box (Wi-Fi) | Borg backups, glance dashboard, personal automation jobs (process-daily, archive-email) under fdrake's systemd-user timers. Runs full workstation home-manager stack. See references/gnomeregan.md. |
WSL Hosts (Colmena-managed, nixos user with sudo)
| Host | Type | Purpose | |------|------|---------| | anton | WSL NixOS on Windows laptop | Gaming and AI processing |
Troubleshooting Workflows
Many *.internal.freddrake.com names unreachable at once
If "service X is down" but other internal sites are also unreachable,
check local Tailscale first: a disconnected probing machine makes the private
fleet look down. Internal service names are resolved by the legacy-named
headscale gateway on official Tailscale SaaS. If local Tailscale is ready,
diagnose that gateway's Tailscale client and dnsmasq; do not add a Headscale
control plane or synthetic web health endpoint. Successful probes of the
private Hetzner application URLs exercise both its DNS and advertised
10.1.0.0/16 route.
tailscale status
tailscale dns status
Service Not Working
- Check service status:
ssh <hostname> "systemctl status <service>" - Check logs:
ssh <hostname> "journalctl -u <service> -n 100" - Restart service:
ssh <hostname> "systemctl restart <service>"
Note: For Hetzner servers, SSH as root. For anton, SSH as nixos and use sudo.
Podman/Container Issues
Check socket status:
ssh <hostname> "systemctl status podman.socket"
List running containers:
ssh <hostname> "podman ps -a"
SSH Connection Issues
If colmena fails with SSH errors:
- Verify the host is reachable:
ping <hostname> - Check if SSH is listening:
ssh <hostname> "ss -tlnp | grep 22" - For Hetzner servers, check via Hetzner console if needed
Common Colmena Patterns
Deploy All Hosts
Use the project-local mass-deploy skill (see "Full-fleet deployment" above).
Do NOT hand a comma-separated all-hosts colmena apply to a subagent — that
deploys in parallel with no health gating between machines.
Update Secrets Before Deploy
just update-secrets
Run this first when secrets changed, then deploy (mass-deploy or single-host).
Workaround Hygiene (unstable hosts)
The unstable hosts (anton, gnomeregan, and the macbook) periodically hit a
package that is broken in nixpkgs-unstable — a redundant patch, a flaky test,
a build failure. The fix is a per-package override in overlays/. These are
temporary: once upstream fixes the package, the override becomes dead
weight and can subtly mask later regressions.
Each temporary override carries a greppable marker comment:
# WORKAROUND(<pkg>): <reason>; remove when <condition>.
The mass-deploy skill runs this audit automatically between pre-flight and
the first apply. Run it manually before ad-hoc single-host unstable deploys,
macbook rebuilds, or when bumping nixpkgs-unstable outside a fleet deploy:
grep -rn 'WORKAROUND(' overlays/
For each marker, test whether the workaround is still needed by building the
stock package (override absent) from the pinned unstable rev on an
x86_64-linux unstable host. Do not just rebuild the whole host: a host build
exercises the overridden package (standalone overlays like highlight.nix
are applied via nodeNixpkgs in colmena/default.nix), so it cannot tell you
whether the stock package is fixed upstream. Building the stock package
directly is the accurate check.
REV=$(jq -r '.nodes["nixpkgs-unstable"].locked.rev' flake.lock) # the pinned rev
# on an unstable x86_64-linux host (anton or gnomeregan):
ssh anton "NIXPKGS_ALLOW_UNFREE=1 nix build --no-link -L --impure \
github:nixos/nixpkgs/$REV#<pkg>" # e.g. tailscale, python313Packages.uvloop
- Builds or substitutes clean → upstream is green at this rev → delete the
override and its marker. A substitute from
cache.nixos.orgis itself a strong signal: Hydra built that stock derivation with its tests enabled. - Fails → keep it; leave the marker. (e.g.
highlight's patch still double-applies → not fixed, never cached.)
For flaky-test / timeout workarounds (a test that flakes or times out rather than failing deterministically), one green is necessary but not sufficient — yet removal is still low-risk: with the override gone, normal builds just substitute Hydra's green binary and only re-run the test on a cache miss, like any other package. Re-adding the one-line override later is trivial.
After deciding to remove one, confirm the consuming config still resolves with
it gone — e.g. nix build --dry-run \ .#darwinConfigurations.macbook-pro.config.system.build.toplevel: the package
should appear under "will be fetched" (substituted), not "will be
built" (which would run its tests locally).
Notes:
- Only override packages tagged
WORKAROUND(are candidates for removal. Other entries inoverlays/default.nixare intentional pins, not staleness-driven — leave them alone:woodpecker-agent(locked to the server image; governed by the woodpecker-upgrade skill),spotifydarwin src override.
- When you add a new workaround, give it a
WORKAROUND(<pkg>)marker with a concrete removal condition so the next audit can retire it. - A new overlay file must be
git add-ed beforecolmena ... --impure— agit+fileflake only sees tracked files, so an untracked overlay fails withpath '.../overlays/<name>.nix' does not exist.
File Locations
| Purpose | Path |
|---------|------|
| Colmena host configs | colmena/hosts/<hostname>.nix |
| Hetzner common modules | colmena/hetzner-common/ |
| WSL common modules | colmena/wsl-common/ |
| NixOS host configs | modules/nixos/host/<hostname>/configuration.nix |
| Application configs | apps/<appname>.nix |
| Service modules (incl. secrets) | modules/services/<service>.nix |
| Container image SHAs | apps/fetcher/containers-sha.nix |
| Container definitions | apps/fetcher/containers.toml |
Related Skills
- provision-nixos-server: Create new Hetzner servers from scratch
References
references/host-mapping.md— Inventory of every managed host (SSH user, port, role).references/gnomeregan.md— Gnomeregan-specific setup, sops identity model, and disaster-recovery procedure. Read first before changing its config or rebuilding it.