This commit is contained in:
2026-10-06 15:31:20 +02:00
parent 7714efd9ea
commit d78e570b44
3 changed files with 183 additions and 47 deletions
+29 -47
View File
@@ -1,6 +1,6 @@
# cirrus
Self-hosted edge: a Talos node running only [sish](https://github.com/antoniomika/sish). Clusters
Self-hosted edge: a host running only [sish](https://github.com/antoniomika/sish). Clusters
without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the
edge relays public traffic back through them:
@@ -11,64 +11,46 @@ client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client
- :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header.
- :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port.
Production runs on Fedora CoreOS (the VPS is too small for Talos), the dev edge on Talos +
Kubernetes. Both use the same sish configuration.
## Layout
```
talos/ node config: generated base + patches (sysctl, scheduling, firewall) → talos/README.md
kubernetes/ sish deployment, kustomize base + per-edge overlay → kubernetes/README.md
.secrets/ connector private keys (gitignored)
coreos/ production edge: Butane config (sish quadlet, firewall, updates) → coreos/README.md
talos/ dev edge node config: generated base + patches → talos/README.md
kubernetes/ dev edge sish deployment, kustomize base + overlay → kubernetes/README.md
**/.secrets/ host and connector private keys (gitignored)
```
Everything is applied by hand (`talosctl`, `kubectl apply -k`). Secrets never leave the
gitignored files (`talos/controlplane.yaml`, `talos/talosconfig`, `**/.secrets/`).
Everything is applied by hand (`butane` + Ignition, `talosctl`, `kubectl apply -k`). Secrets never
leave the gitignored files (`coreos/.secrets/`, `coreos/*.ign`, `talos/controlplane.yaml`,
`talos/talosconfig`, `**/.secrets/`).
## Environments
| | Edge | Domains |
|---|---|---|
| `cirrus` | `tunnel.traberph.de`, Fedora CoreOS (stable) | any hostname pointed at the VPS |
| `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS |
Currently served through the dev edge from `cumulus`: Gitea SSH on :22, `http://cirrus.sto` (hello).
Production (`cirrus`) serves everything from `cumulus` (since 2026-10-06): `traberph.de` and
`*.traberph.de` on :80/:443 (IPv4 and IPv6), Gitea SSH on :22. The cluster side (connectors, Envoy
gateways, per-app routes) is documented in the `cumulus` README. The dev edge only serves
`.test`/`.sto` names.
## Production rollout plan
## Production rollout (done 2026-10-06)
The dev edge is verified end to end: firewall, port range, tunnels, and the Talos config
reproduces exactly from the files here. What is still missing for production:
Rolled out as planned: CoreOS VPS with sish, second connector on `cumulus` for TLS passthrough +
PROXY v2, Envoy HTTPS gateway with cert-manager (Let's Encrypt HTTP-01 over the :80 route), services
moved from the Cloudflare tunnel one hostname at a time by switching DNS.
**Before the rollout**
1. **Backups.** `talos/controlplane.yaml`, `talos/talosconfig`, the edge host key and the
connector private keys exist only in this folder. Store them in a password manager (and put
the folder under git, secrets stay gitignored).
2. **Domains and DNS.** Pick the production domains; create DNS records for them (wildcards
where needed) pointing to the VPS IPv4/IPv6.
3. **HTTPS on the cluster side.** `cumulus` only has the plain connector (SSH, HTTP). For
TLS passthrough it needs a second connector (SNI, PROXY v2), an Envoy HTTPS listener that
accepts the PROXY header only from sish-client, and certificates (cert-manager with DNS-01, or
HTTP-01 over the :80 route).
4. **Remove test access.** Leave `connector-hello.pub` (local test stack) out of the
production overlay; give each production connector its own key.
**Rollout**
1. VPS: boot the Talos image (same factory schematic as dev) and check the provider's disk name
and network (DHCP vs static, IPv6).
2. Generate a new base config into its own folder (`talos/README.md`, "New node") and apply it
with the existing patches plus one that disables the discovery service (single node, no
external dependency needed).
3. Bootstrap, fetch the kubeconfig. Verify against the firewall table: only the listed ports
answer from outside.
4. Kubernetes: new overlay `kubernetes/cirrus-prod` (copy of `cirrus-dev`) with the production
`SISH_DOMAIN`/`SISH_BIND_HOSTS`, production connector keys and a new host key. Set
`dnsPolicy: Default` for sish so the tunnel does not depend on CoreDNS. `diff`, then `apply -k`.
5. Cluster side (`cumulus`): sish-client deployment(s) for the production edge (edge IP, new host
key, production routes), matching Gateway listeners/routes and network policies. Keep the dev
connectors until production is verified.
6. Verify from outside: `ssh-keyscan` for SSH routes, `curl` for HTTP/HTTPS routes, the backend
sees the real client IP, closed ports stay closed.
7. Move existing services (e.g. from the Cloudflare tunnel) one hostname at a time by switching
their DNS records.
**After the rollout**
- External uptime check on the edge (sish :2222 and one route per protocol): the edge is a
single point of failure for everything behind it.
- Updates by hand: `talosctl upgrade` / `upgrade-k8s`, the sish image tag, sish-client tags in
`cumulus`. The `talosconfig` admin certificate expires after one year.
**Still open**
- **Backups.** `coreos/.secrets/` (edge host key, `cumulus` connector key) and the Talos dev
credentials exist only in this folder. Store them in a password manager.
- **Uptime check.** External check on the edge (sish :5002 and one route per protocol): the edge is
a single point of failure for everything behind it.
- **Updates by hand.** Production updates are automatic (OS in a nightly reboot window, sish within
v2, see `coreos/README.md`). sish-client tags in `cumulus` and the dev edge (`talosctl upgrade` /
`upgrade-k8s`, sish image tag) are updated by hand; the `talosconfig` admin certificate expires
after one year.