- talos/: generated base config (gitignored) + patches for control-plane scheduling, unprivileged ports and the ingress firewall - kubernetes/: sish base and cirrus-dev overlay, applied with kubectl - READMEs incl. production rollout plan Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
75 lines
3.8 KiB
Markdown
75 lines
3.8 KiB
Markdown
# cirrus
|
|
|
|
Self-hosted edge: a Talos node running only [sish](https://github.com/antoniomika/sish). Clusters
|
|
without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the
|
|
edge relays public traffic back through them:
|
|
|
|
```
|
|
client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client ──▶ envoy gateway ──▶ app
|
|
```
|
|
|
|
- :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header.
|
|
- :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port.
|
|
|
|
## Layout
|
|
|
|
```
|
|
talos/ node config: generated base + patches (sysctl, scheduling, firewall) → talos/README.md
|
|
kubernetes/ sish deployment, kustomize base + per-edge overlay → kubernetes/README.md
|
|
.secrets/ connector private keys (gitignored)
|
|
```
|
|
|
|
Everything is applied by hand (`talosctl`, `kubectl apply -k`). Secrets never leave the
|
|
gitignored files (`talos/controlplane.yaml`, `talos/talosconfig`, `**/.secrets/`).
|
|
|
|
## Environments
|
|
|
|
| | Edge | Domains |
|
|
|---|---|---|
|
|
| `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS |
|
|
|
|
Currently served through the dev edge from `cumulus`: Gitea SSH on :22, `http://cirrus.sto` (hello).
|
|
|
|
## Production rollout plan
|
|
|
|
The dev edge is verified end to end: firewall, port range, tunnels, and the Talos config
|
|
reproduces exactly from the files here. What is still missing for production:
|
|
|
|
**Before the rollout**
|
|
1. **Backups.** `talos/controlplane.yaml`, `talos/talosconfig`, the edge host key and the
|
|
connector private keys exist only in this folder. Store them in a password manager (and put
|
|
the folder under git, secrets stay gitignored).
|
|
2. **Domains and DNS.** Pick the production domains; create DNS records for them (wildcards
|
|
where needed) pointing to the VPS IPv4/IPv6.
|
|
3. **HTTPS on the cluster side.** `cumulus` only has the plain connector (SSH, HTTP). For
|
|
TLS passthrough it needs a second connector (SNI, PROXY v2), an Envoy HTTPS listener that
|
|
accepts the PROXY header only from sish-client, and certificates (cert-manager with DNS-01, or
|
|
HTTP-01 over the :80 route).
|
|
4. **Remove test access.** Leave `connector-hello.pub` (local test stack) out of the
|
|
production overlay; give each production connector its own key.
|
|
|
|
**Rollout**
|
|
1. VPS: boot the Talos image (same factory schematic as dev) and check the provider's disk name
|
|
and network (DHCP vs static, IPv6).
|
|
2. Generate a new base config into its own folder (`talos/README.md`, "New node") and apply it
|
|
with the existing patches plus one that disables the discovery service (single node, no
|
|
external dependency needed).
|
|
3. Bootstrap, fetch the kubeconfig. Verify against the firewall table: only the listed ports
|
|
answer from outside.
|
|
4. Kubernetes: new overlay `kubernetes/cirrus-prod` (copy of `cirrus-dev`) with the production
|
|
`SISH_DOMAIN`/`SISH_BIND_HOSTS`, production connector keys and a new host key. Set
|
|
`dnsPolicy: Default` for sish so the tunnel does not depend on CoreDNS. `diff`, then `apply -k`.
|
|
5. Cluster side (`cumulus`): sish-client deployment(s) for the production edge (edge IP, new host
|
|
key, production routes), matching Gateway listeners/routes and network policies. Keep the dev
|
|
connectors until production is verified.
|
|
6. Verify from outside: `ssh-keyscan` for SSH routes, `curl` for HTTP/HTTPS routes, the backend
|
|
sees the real client IP, closed ports stay closed.
|
|
7. Move existing services (e.g. from the Cloudflare tunnel) one hostname at a time by switching
|
|
their DNS records.
|
|
|
|
**After the rollout**
|
|
- External uptime check on the edge (sish :2222 and one route per protocol): the edge is a
|
|
single point of failure for everything behind it.
|
|
- Updates by hand: `talosctl upgrade` / `upgrade-k8s`, the sish image tag, sish-client tags in
|
|
`cumulus`. The `talosconfig` admin certificate expires after one year.
|