This commit is contained in:
2026-10-06 15:31:20 +02:00
parent 7714efd9ea
commit d78e570b44
3 changed files with 183 additions and 47 deletions
+29 -47
View File
@@ -1,6 +1,6 @@
# cirrus
Self-hosted edge: a Talos node running only [sish](https://github.com/antoniomika/sish). Clusters
Self-hosted edge: a host running only [sish](https://github.com/antoniomika/sish). Clusters
without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the
edge relays public traffic back through them:
@@ -11,64 +11,46 @@ client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client
- :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header.
- :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port.
Production runs on Fedora CoreOS (the VPS is too small for Talos), the dev edge on Talos +
Kubernetes. Both use the same sish configuration.
## Layout
```
talos/ node config: generated base + patches (sysctl, scheduling, firewall) → talos/README.md
kubernetes/ sish deployment, kustomize base + per-edge overlay → kubernetes/README.md
.secrets/ connector private keys (gitignored)
coreos/ production edge: Butane config (sish quadlet, firewall, updates) → coreos/README.md
talos/ dev edge node config: generated base + patches → talos/README.md
kubernetes/ dev edge sish deployment, kustomize base + overlay → kubernetes/README.md
**/.secrets/ host and connector private keys (gitignored)
```
Everything is applied by hand (`talosctl`, `kubectl apply -k`). Secrets never leave the
gitignored files (`talos/controlplane.yaml`, `talos/talosconfig`, `**/.secrets/`).
Everything is applied by hand (`butane` + Ignition, `talosctl`, `kubectl apply -k`). Secrets never
leave the gitignored files (`coreos/.secrets/`, `coreos/*.ign`, `talos/controlplane.yaml`,
`talos/talosconfig`, `**/.secrets/`).
## Environments
| | Edge | Domains |
|---|---|---|
| `cirrus` | `tunnel.traberph.de`, Fedora CoreOS (stable) | any hostname pointed at the VPS |
| `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS |
Currently served through the dev edge from `cumulus`: Gitea SSH on :22, `http://cirrus.sto` (hello).
Production (`cirrus`) serves everything from `cumulus` (since 2026-10-06): `traberph.de` and
`*.traberph.de` on :80/:443 (IPv4 and IPv6), Gitea SSH on :22. The cluster side (connectors, Envoy
gateways, per-app routes) is documented in the `cumulus` README. The dev edge only serves
`.test`/`.sto` names.
## Production rollout plan
## Production rollout (done 2026-10-06)
The dev edge is verified end to end: firewall, port range, tunnels, and the Talos config
reproduces exactly from the files here. What is still missing for production:
Rolled out as planned: CoreOS VPS with sish, second connector on `cumulus` for TLS passthrough +
PROXY v2, Envoy HTTPS gateway with cert-manager (Let's Encrypt HTTP-01 over the :80 route), services
moved from the Cloudflare tunnel one hostname at a time by switching DNS.
**Before the rollout**
1. **Backups.** `talos/controlplane.yaml`, `talos/talosconfig`, the edge host key and the
connector private keys exist only in this folder. Store them in a password manager (and put
the folder under git, secrets stay gitignored).
2. **Domains and DNS.** Pick the production domains; create DNS records for them (wildcards
where needed) pointing to the VPS IPv4/IPv6.
3. **HTTPS on the cluster side.** `cumulus` only has the plain connector (SSH, HTTP). For
TLS passthrough it needs a second connector (SNI, PROXY v2), an Envoy HTTPS listener that
accepts the PROXY header only from sish-client, and certificates (cert-manager with DNS-01, or
HTTP-01 over the :80 route).
4. **Remove test access.** Leave `connector-hello.pub` (local test stack) out of the
production overlay; give each production connector its own key.
**Rollout**
1. VPS: boot the Talos image (same factory schematic as dev) and check the provider's disk name
and network (DHCP vs static, IPv6).
2. Generate a new base config into its own folder (`talos/README.md`, "New node") and apply it
with the existing patches plus one that disables the discovery service (single node, no
external dependency needed).
3. Bootstrap, fetch the kubeconfig. Verify against the firewall table: only the listed ports
answer from outside.
4. Kubernetes: new overlay `kubernetes/cirrus-prod` (copy of `cirrus-dev`) with the production
`SISH_DOMAIN`/`SISH_BIND_HOSTS`, production connector keys and a new host key. Set
`dnsPolicy: Default` for sish so the tunnel does not depend on CoreDNS. `diff`, then `apply -k`.
5. Cluster side (`cumulus`): sish-client deployment(s) for the production edge (edge IP, new host
key, production routes), matching Gateway listeners/routes and network policies. Keep the dev
connectors until production is verified.
6. Verify from outside: `ssh-keyscan` for SSH routes, `curl` for HTTP/HTTPS routes, the backend
sees the real client IP, closed ports stay closed.
7. Move existing services (e.g. from the Cloudflare tunnel) one hostname at a time by switching
their DNS records.
**After the rollout**
- External uptime check on the edge (sish :2222 and one route per protocol): the edge is a
single point of failure for everything behind it.
- Updates by hand: `talosctl upgrade` / `upgrade-k8s`, the sish image tag, sish-client tags in
`cumulus`. The `talosconfig` admin certificate expires after one year.
**Still open**
- **Backups.** `coreos/.secrets/` (edge host key, `cumulus` connector key) and the Talos dev
credentials exist only in this folder. Store them in a password manager.
- **Uptime check.** External check on the edge (sish :5002 and one route per protocol): the edge is
a single point of failure for everything behind it.
- **Updates by hand.** Production updates are automatic (OS in a nightly reboot window, sish within
v2, see `coreos/README.md`). sish-client tags in `cumulus` and the dev edge (`talosctl upgrade` /
`upgrade-k8s`, sish image tag) are updated by hand; the `talosconfig` admin certificate expires
after one year.
+130
View File
@@ -0,0 +1,130 @@
# Fedora CoreOS: cirrus edge (production)
Single Fedora CoreOS host that runs only sish, as a rootful Podman quadlet with host networking.
Everything is defined in `cirrus.yaml` (Butane) and applied once at install time by Ignition.
The VPS is too small for Talos, so production runs on CoreOS; the Talos setup in `talos/` and
`kubernetes/` stays the dev edge.
| File | |
|---|---|
| `cirrus.yaml` | Butane config: users, sshd, firewall, sish, updates |
| `.secrets/ssh_host_ed25519_key` | Edge SSH host key (gitignored). Connectors pin its public half (`SISH_HOST_KEY`) |
| `.secrets/connector-cumulus` | Private key of the `cumulus` connector (gitignored), goes into a Secret in `cumulus` |
| `cirrus.ign` | Build output, embeds the host key (gitignored) |
## Build and install
```sh
butane --strict --files-dir coreos coreos/cirrus.yaml > coreos/cirrus.ign
coreos-installer install /dev/<disk> --ignition-file coreos/cirrus.ign # or the provider's user-data
```
Ignition runs only on first boot. Changing `cirrus.yaml` later does nothing to a running host:
either reinstall, or make the same change on the host by hand (and keep the file in sync).
Admin access: `ssh -p 5001 core@tunnel.traberph.de` (public key only, user `core` only).
## Network
Hostname `cirrus.traberph.de`. netcup gives IPv4 via DHCP (`46.38.234.119`) but no IPv6 router
advertisements with a usable prefix: the IPv6 address from the netcup panel (/64) is set statically
in `/etc/NetworkManager/system-connections/ens3.nmconnection`, gateway `fe80::1`.
| | |
|---|---|
| IPv4 | `46.38.234.119` (DHCP) |
| IPv6 | `2a03:4000:2:83c::1/64` (static), gateway `fe80::1` |
Only the global address (`scope global`) goes into DNS, never the `fe80::` link-local one.
## Ports
nftables (`/etc/sysconfig/nftables.conf`), default drop. Loopback, ICMP, DHCP replies and
replies to outgoing connections are allowed.
| Port (tcp) | |
|---|---|
| 5001 | Admin sshd |
| 5002 | sish SSH endpoint for connectors (public key auth) |
| 80 | sish HTTP, routed by `Host` header |
| 443 | sish TLS passthrough, routed by SNI |
| 22, 20000-20099 | Raw TCP forwards (22 = Gitea SSH) |
The forward ports must match `port-bind-range` in the sish config, otherwise a claimed port is
silently unreachable.
sshd on 5001 needs an SELinux exception (5001 is labelled `commplex_link_port_t`):
`sshd-port-selinux.service` installs `/etc/cirrus/sshd_port_5001.cil` once and again whenever an
OS update dropped it.
## sish
| Path on the host | |
|---|---|
| `/etc/containers/systemd/sish.container` | Quadlet unit (`sish.service`): non-root (uid 65532), only `CAP_NET_BIND_SERVICE`, read-only root, `MemoryMax=256M`. Own writable `/tmp` tmpfs (`Tmpfs=…,mode=1777,notmpcopyup`): sish creates a temp file per forward, and podman's automatic read-only `/tmp` (copied from the image, root 755) would make every forward fail with "remote port forwarding failed" |
| `/etc/sish/config.yml` | sish config, same as `kubernetes/base/sish/config.yml` except the values below |
| `/var/lib/sish/keys/` | Host key (read-only in the container) |
| `/var/lib/sish/pubkeys/clients` | Authorized connector keys, `authorized_keys` format |
Differences to the k8s config:
| | |
|---|---|
| `ssh-address: ":5002"` | 2222 in k8s |
| `domain: tunnel.traberph.de` | The edge's own name. A requested name without a dot becomes `<name>.tunnel.traberph.de` |
| `bind-any-host: true` | Single tenant: connectors may claim any hostname containing a dot, wildcards included. Replaces `bind-hosts` |
| `verify-dns: false` | No `_sish` TXT ownership checks (pointless with `bind-any-host`) |
Every authorized key can claim every hostname and port, and with the load balancers on it can join
an existing one. So only add keys of connectors you control (sish has no per-key permissions).
**Add or remove a connector:** edit `/var/lib/sish/pubkeys/clients` on the host (sish watches the
directory, no restart needed) and the same block in `cirrus.yaml`.
**Config change:** edit `/etc/sish/config.yml`, `systemctl restart sish`, mirror it in
`cirrus.yaml`. Connectors drop for a few seconds and reconnect on their own.
```sh
systemctl status sish
journalctl -u sish -f
```
## DNS
| Record | |
|---|---|
| `tunnel.traberph.de` A/AAAA → VPS | Connectors (:5002) and admin SSH (:5001) |
| `<host>` or `*.<domain>` A/AAAA → VPS | Every hostname a connector serves; unclaimed names get a 404 (:80) or no answer (:443) |
| `*.tunnel.traberph.de` A/AAAA → VPS | Optional, only if fallback names should be reachable |
A wildcard claim (`*.example.com`) does not cover the apex `example.com`, neither in DNS nor in sish:
connectors claim the apex separately. Records that point elsewhere (e.g. still proxied through
Cloudflare) take precedence over the wildcard; deleting such a record silently moves the name to the
edge, where it only works if a connector serves it.
## Updates
Both are automatic:
- **OS:** Zincati stages new Fedora CoreOS releases (stable stream) and reboots only in the window
03:00-04:00 UTC (`/etc/zincati/config.d/55-updates-strategy.toml`).
- **sish:** upstream publishes only exact tags (`v2.24.0`), so `podman auto-update` can't follow a
version line. `sish-update.timer` (daily ~05:00 UTC) runs `/usr/local/bin/sish-update`, which
sets `Image=` in the quadlet to the newest tag matching `TRACK=v2.` (minor + patch, never a new
major), restarts sish and rolls back if nothing listens on :5002 after 30s. `TRACK=v2.24.` limits
it to patch releases.
```sh
journalctl -u zincati -u sish-update
systemctl start sish-update # check now
```
Both restart sish (tunnels drop for a few seconds). A major sish release (`v3`) is a manual change
of `TRACK` and `Image=`.
## First boot checklist
- `ss -tlnp`: sish on 22, 80, 443, 5002; sshd on 5001.
- `journalctl -u sish`: `Loading ssh_host_ed25519_key as ssh-ed25519 host key` (the provisioned
key, not a generated one).
- After the first OS update: SSH on 5001 still works, `semodule -l | grep sshd_port_5001`.
+24
View File
@@ -16,6 +16,27 @@ storage:
user: { id: 65532 }
group: { id: 65532 }
files:
- path: /etc/hostname
mode: 0644
contents:
inline: cirrus.traberph.de
# netcup: IPv4 via DHCP, IPv6 static (no router advertisements, gateway is always fe80::1)
- path: /etc/NetworkManager/system-connections/ens3.nmconnection
mode: 0600
contents:
inline: |
[connection]
id=ens3
type=ethernet
interface-name=ens3
[ipv4]
method=auto
[ipv6]
method=manual
address1=2a03:4000:2:83c::1/64
gateway=fe80::1
# Edge SSH host key (connectors pin its public half). Kept only locally in coreos/.secrets/
# (gitignored), so back it up outside this folder. Build: butane --files-dir coreos ...
- path: /var/lib/sish/keys/ssh_host_ed25519_key
@@ -118,6 +139,9 @@ storage:
AddCapability=CAP_NET_BIND_SERVICE
NoNewPrivileges=true
ReadOnly=true
# sish creates a temp file per forward. The automatic read-only /tmp tmpfs copies the
# image's /tmp (root, 755), so uid 65532 can't write there: give it a plain sticky /tmp.
Tmpfs=/tmp:rw,nosuid,nodev,noexec,size=16m,mode=1777,notmpcopyup
Volume=/etc/sish:/config:ro,Z
Volume=/var/lib/sish/keys:/keys:ro,Z
Volume=/var/lib/sish/pubkeys:/pubkeys:ro,Z