cleanup talos

This commit is contained in:
2026-10-06 15:48:50 +02:00
parent d78e570b44
commit 3769647b77
11 changed files with 140 additions and 348 deletions
+132 -33
View File
@@ -1,8 +1,8 @@
# cirrus
Self-hosted edge: a host running only [sish](https://github.com/antoniomika/sish). Clusters
without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the
edge relays public traffic back through them:
Self-hosted edge: a Fedora CoreOS VPS running only [sish](https://github.com/antoniomika/sish).
Clusters without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`,
and the edge relays public traffic back through them:
```
client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client ──▶ envoy gateway ──▶ app
@@ -11,46 +11,145 @@ client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client
- :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header.
- :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port.
Production runs on Fedora CoreOS (the VPS is too small for Talos), the dev edge on Talos +
Kubernetes. Both use the same sish configuration.
Production serves everything from `cumulus`: `traberph.de` and `*.traberph.de` on :80/:443 (IPv4
and IPv6), Gitea SSH on :22. The cluster side (connectors, Envoy gateways, per-app routes) is
documented in the `cumulus` README.
## Layout
```
coreos/ production edge: Butane config (sish quadlet, firewall, updates) → coreos/README.md
talos/ dev edge node config: generated base + patches → talos/README.md
kubernetes/ dev edge sish deployment, kustomize base + overlay → kubernetes/README.md
**/.secrets/ host and connector private keys (gitignored)
| File | |
|---|---|
| `cirrus.yaml` | Butane config: users, sshd, network, firewall, sish, updates |
| `.secrets/ssh_host_ed25519_key` | Edge SSH host key (gitignored). Connectors pin its public half (`SISH_HOST_KEY`) |
| `.secrets/connector-cumulus` | Private key of the `cumulus` connector (gitignored), goes into a Secret in `cumulus` |
| `cirrus.ign` | Build output, embeds the host key (gitignored) |
Secrets only live in the gitignored `.secrets/` and `*.ign`; nothing secret is committed. Back up
`.secrets/` outside this folder (password manager), it is the only copy.
## Build and install
```sh
butane --strict --files-dir . cirrus.yaml > cirrus.ign
coreos-installer install /dev/<disk> --ignition-file cirrus.ign # or the provider's user-data
```
Everything is applied by hand (`butane` + Ignition, `talosctl`, `kubectl apply -k`). Secrets never
leave the gitignored files (`coreos/.secrets/`, `coreos/*.ign`, `talos/controlplane.yaml`,
`talos/talosconfig`, `**/.secrets/`).
Ignition runs only on first boot. Changing `cirrus.yaml` later does nothing to a running host:
either reinstall, or make the same change on the host by hand (and keep the file in sync).
## Environments
Admin access: `ssh -p 5001 core@tunnel.traberph.de` (public key only, user `core` only).
| | Edge | Domains |
|---|---|---|
| `cirrus` | `tunnel.traberph.de`, Fedora CoreOS (stable) | any hostname pointed at the VPS |
| `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS |
### First boot checklist
Production (`cirrus`) serves everything from `cumulus` (since 2026-10-06): `traberph.de` and
`*.traberph.de` on :80/:443 (IPv4 and IPv6), Gitea SSH on :22. The cluster side (connectors, Envoy
gateways, per-app routes) is documented in the `cumulus` README. The dev edge only serves
`.test`/`.sto` names.
- `ss -tlnp`: sish on 22, 80, 443, 5002; sshd on 5001.
- `journalctl -u sish`: `Loading ssh_host_ed25519_key as ssh-ed25519 host key` (the provisioned
key, not a generated one).
- After the first OS update: SSH on 5001 still works, `semodule -l | grep sshd_port_5001`.
## Production rollout (done 2026-10-06)
## Network
Rolled out as planned: CoreOS VPS with sish, second connector on `cumulus` for TLS passthrough +
PROXY v2, Envoy HTTPS gateway with cert-manager (Let's Encrypt HTTP-01 over the :80 route), services
moved from the Cloudflare tunnel one hostname at a time by switching DNS.
Hostname `cirrus.traberph.de` on netcup. netcup gives IPv4 via DHCP but no IPv6 router
advertisements with a usable prefix: the IPv6 address from the netcup panel (/64) is set statically
in `/etc/NetworkManager/system-connections/ens3.nmconnection`, gateway `fe80::1`.
**Still open**
- **Backups.** `coreos/.secrets/` (edge host key, `cumulus` connector key) and the Talos dev
credentials exist only in this folder. Store them in a password manager.
| | |
|---|---|
| IPv4 | `46.38.234.119` (DHCP) |
| IPv6 | `2a03:4000:2:83c::1/64` (static), gateway `fe80::1` |
Only the global address (`scope global`) goes into DNS, never the `fe80::` link-local one.
## Ports
nftables (`/etc/sysconfig/nftables.conf`), default drop. Loopback, ICMP, DHCP replies and
replies to outgoing connections are allowed.
| Port (tcp) | |
|---|---|
| 5001 | Admin sshd |
| 5002 | sish SSH endpoint for connectors (public key auth) |
| 80 | sish HTTP, routed by `Host` header |
| 443 | sish TLS passthrough, routed by SNI |
| 22, 20000-20099 | Raw TCP forwards (22 = Gitea SSH) |
The forward ports must match `port-bind-range` in the sish config, otherwise a claimed port is
silently unreachable.
sshd on 5001 needs an SELinux exception (5001 is labelled `commplex_link_port_t`):
`sshd-port-selinux.service` installs `/etc/cirrus/sshd_port_5001.cil` once and again whenever an
OS update dropped it.
## sish
| Path on the host | |
|---|---|
| `/etc/containers/systemd/sish.container` | Quadlet unit (`sish.service`), rootful Podman with host networking: non-root (uid 65532), only `CAP_NET_BIND_SERVICE`, read-only root, `MemoryMax=256M`. Own writable `/tmp` tmpfs (`Tmpfs=…,mode=1777,notmpcopyup`): sish creates a temp file per forward, and podman's automatic read-only `/tmp` (copied from the image, root 755) would make every forward fail with "remote port forwarding failed" |
| `/etc/sish/config.yml` | sish config |
| `/var/lib/sish/keys/` | Host key (read-only in the container) |
| `/var/lib/sish/pubkeys/clients` | Authorized connector keys, `authorized_keys` format |
Notable config values:
| | |
|---|---|
| `ssh-address: ":5002"` | Connector SSH endpoint |
| `domain: tunnel.traberph.de` | The edge's own name. A requested name without a dot becomes `<name>.tunnel.traberph.de` |
| `bind-any-host: true` | Single tenant: connectors may claim any hostname containing a dot, wildcards included |
| `verify-dns: false` | No `_sish` TXT ownership checks (pointless with `bind-any-host`) |
| `sni-proxy`, `*-load-balancer: true` | TLS passthrough on :443, several connectors may serve the same name |
| `proxy-protocol-version: "2"` | PROXY header for connectors that request it |
| `idle-connection-timeout: 1h` | Default 5s kills websockets, SSE and slow uploads |
| `service-console-max-content-length: 0` | Default -1 buffers every body in memory, large uploads OOM-kill sish |
Every authorized key can claim every hostname and port, and with the load balancers on it can join
an existing one. So only add keys of connectors you control (sish has no per-key permissions).
**Add or remove a connector:** edit `/var/lib/sish/pubkeys/clients` on the host (sish watches the
directory, no restart needed) and the same block in `cirrus.yaml`.
**Config change:** edit `/etc/sish/config.yml`, `systemctl restart sish`, mirror it in
`cirrus.yaml`. Connectors drop for a few seconds and reconnect on their own.
```sh
systemctl status sish
journalctl -u sish -f
```
## DNS
| Record | |
|---|---|
| `tunnel.traberph.de` A/AAAA → VPS | Connectors (:5002) and admin SSH (:5001) |
| `<host>` or `*.<domain>` A/AAAA → VPS | Every hostname a connector serves; unclaimed names get a 404 (:80) or no answer (:443) |
| `*.tunnel.traberph.de` A/AAAA → VPS | Optional, only if fallback names should be reachable |
A wildcard claim (`*.example.com`) does not cover the apex `example.com`, neither in DNS nor in sish:
connectors claim the apex separately. Records that point elsewhere (e.g. still proxied through
Cloudflare) take precedence over the wildcard; deleting such a record silently moves the name to the
edge, where it only works if a connector serves it.
## Updates
Both are automatic:
- **OS:** Zincati stages new Fedora CoreOS releases (stable stream) and reboots only in the window
03:00-04:00 UTC (`/etc/zincati/config.d/55-updates-strategy.toml`).
- **sish:** upstream publishes only exact tags (`v2.24.0`), so `podman auto-update` can't follow a
version line. `sish-update.timer` (daily ~05:00 UTC) runs `/usr/local/bin/sish-update`, which
sets `Image=` in the quadlet to the newest tag matching `TRACK=v2.` (minor + patch, never a new
major), restarts sish and rolls back if nothing listens on :5002 after 30s. `TRACK=v2.24.` limits
it to patch releases.
```sh
journalctl -u zincati -u sish-update
systemctl start sish-update # check now
```
Both restart sish (tunnels drop for a few seconds). A major sish release (`v3`) is a manual change
of `TRACK` and `Image=`. The `sish-client` tags in `cumulus` are updated by hand.
## Open
- **Backups.** `.secrets/` exists only in this folder. Store it in a password manager.
- **Uptime check.** External check on the edge (sish :5002 and one route per protocol): the edge is
a single point of failure for everything behind it.
- **Updates by hand.** Production updates are automatic (OS in a nightly reboot window, sish within
v2, see `coreos/README.md`). sish-client tags in `cumulus` and the dev edge (`talosctl upgrade` /
`upgrade-k8s`, sish image tag) are updated by hand; the `talosconfig` admin certificate expires
after one year.