# cirrus Self-hosted edge: a Talos node running only [sish](https://github.com/antoniomika/sish). Clusters without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the edge relays public traffic back through them: ``` client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client ──▶ envoy gateway ──▶ app ``` - :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header. - :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port. ## Layout ``` talos/ node config: generated base + patches (sysctl, scheduling, firewall) → talos/README.md kubernetes/ sish deployment, kustomize base + per-edge overlay → kubernetes/README.md .secrets/ connector private keys (gitignored) ``` Everything is applied by hand (`talosctl`, `kubectl apply -k`). Secrets never leave the gitignored files (`talos/controlplane.yaml`, `talos/talosconfig`, `**/.secrets/`). ## Environments | | Edge | Domains | |---|---|---| | `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS | Currently served through the dev edge from `cumulus`: Gitea SSH on :22, `http://cirrus.sto` (hello). ## Production rollout plan The dev edge is verified end to end: firewall, port range, tunnels, and the Talos config reproduces exactly from the files here. What is still missing for production: **Before the rollout** 1. **Backups.** `talos/controlplane.yaml`, `talos/talosconfig`, the edge host key and the connector private keys exist only in this folder. Store them in a password manager (and put the folder under git, secrets stay gitignored). 2. **Domains and DNS.** Pick the production domains; create DNS records for them (wildcards where needed) pointing to the VPS IPv4/IPv6. 3. **HTTPS on the cluster side.** `cumulus` only has the plain connector (SSH, HTTP). For TLS passthrough it needs a second connector (SNI, PROXY v2), an Envoy HTTPS listener that accepts the PROXY header only from sish-client, and certificates (cert-manager with DNS-01, or HTTP-01 over the :80 route). 4. **Remove test access.** Leave `connector-hello.pub` (local test stack) out of the production overlay; give each production connector its own key. **Rollout** 1. VPS: boot the Talos image (same factory schematic as dev) and check the provider's disk name and network (DHCP vs static, IPv6). 2. Generate a new base config into its own folder (`talos/README.md`, "New node") and apply it with the existing patches plus one that disables the discovery service (single node, no external dependency needed). 3. Bootstrap, fetch the kubeconfig. Verify against the firewall table: only the listed ports answer from outside. 4. Kubernetes: new overlay `kubernetes/cirrus-prod` (copy of `cirrus-dev`) with the production `SISH_DOMAIN`/`SISH_BIND_HOSTS`, production connector keys and a new host key. Set `dnsPolicy: Default` for sish so the tunnel does not depend on CoreDNS. `diff`, then `apply -k`. 5. Cluster side (`cumulus`): sish-client deployment(s) for the production edge (edge IP, new host key, production routes), matching Gateway listeners/routes and network policies. Keep the dev connectors until production is verified. 6. Verify from outside: `ssh-keyscan` for SSH routes, `curl` for HTTP/HTTPS routes, the backend sees the real client IP, closed ports stay closed. 7. Move existing services (e.g. from the Cloudflare tunnel) one hostname at a time by switching their DNS records. **After the rollout** - External uptime check on the edge (sish :2222 and one route per protocol): the edge is a single point of failure for everything behind it. - Updates by hand: `talosctl upgrade` / `upgrade-k8s`, the sish image tag, sish-client tags in `cumulus`. The `talosconfig` admin certificate expires after one year.