ipv6
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# cirrus
|
||||
|
||||
Self-hosted edge: a Talos node running only [sish](https://github.com/antoniomika/sish). Clusters
|
||||
Self-hosted edge: a host running only [sish](https://github.com/antoniomika/sish). Clusters
|
||||
without inbound ports (e.g. `cumulus`) open outbound SSH tunnels to it with `sish-client`, and the
|
||||
edge relays public traffic back through them:
|
||||
|
||||
@@ -11,64 +11,46 @@ client ──▶ edge :22/:80/:443/:200xx (sish) ══ssh══▶ sish-client
|
||||
- :443 is routed by SNI without decrypting (TLS passthrough), optionally with a PROXY v2 header.
|
||||
- :80 is routed by `Host` header, raw TCP ports (e.g. :22 for Gitea SSH) by port.
|
||||
|
||||
Production runs on Fedora CoreOS (the VPS is too small for Talos), the dev edge on Talos +
|
||||
Kubernetes. Both use the same sish configuration.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
talos/ node config: generated base + patches (sysctl, scheduling, firewall) → talos/README.md
|
||||
kubernetes/ sish deployment, kustomize base + per-edge overlay → kubernetes/README.md
|
||||
.secrets/ connector private keys (gitignored)
|
||||
coreos/ production edge: Butane config (sish quadlet, firewall, updates) → coreos/README.md
|
||||
talos/ dev edge node config: generated base + patches → talos/README.md
|
||||
kubernetes/ dev edge sish deployment, kustomize base + overlay → kubernetes/README.md
|
||||
**/.secrets/ host and connector private keys (gitignored)
|
||||
```
|
||||
|
||||
Everything is applied by hand (`talosctl`, `kubectl apply -k`). Secrets never leave the
|
||||
gitignored files (`talos/controlplane.yaml`, `talos/talosconfig`, `**/.secrets/`).
|
||||
Everything is applied by hand (`butane` + Ignition, `talosctl`, `kubectl apply -k`). Secrets never
|
||||
leave the gitignored files (`coreos/.secrets/`, `coreos/*.ign`, `talos/controlplane.yaml`,
|
||||
`talos/talosconfig`, `**/.secrets/`).
|
||||
|
||||
## Environments
|
||||
|
||||
| | Edge | Domains |
|
||||
|---|---|---|
|
||||
| `cirrus` | `tunnel.traberph.de`, Fedora CoreOS (stable) | any hostname pointed at the VPS |
|
||||
| `cirrus-dev` | `10.20.5.130` (LAN), Talos v1.14.1, Kubernetes v1.37.0 | `.test` / `.sto` via local DNS |
|
||||
|
||||
Currently served through the dev edge from `cumulus`: Gitea SSH on :22, `http://cirrus.sto` (hello).
|
||||
Production (`cirrus`) serves everything from `cumulus` (since 2026-10-06): `traberph.de` and
|
||||
`*.traberph.de` on :80/:443 (IPv4 and IPv6), Gitea SSH on :22. The cluster side (connectors, Envoy
|
||||
gateways, per-app routes) is documented in the `cumulus` README. The dev edge only serves
|
||||
`.test`/`.sto` names.
|
||||
|
||||
## Production rollout plan
|
||||
## Production rollout (done 2026-10-06)
|
||||
|
||||
The dev edge is verified end to end: firewall, port range, tunnels, and the Talos config
|
||||
reproduces exactly from the files here. What is still missing for production:
|
||||
Rolled out as planned: CoreOS VPS with sish, second connector on `cumulus` for TLS passthrough +
|
||||
PROXY v2, Envoy HTTPS gateway with cert-manager (Let's Encrypt HTTP-01 over the :80 route), services
|
||||
moved from the Cloudflare tunnel one hostname at a time by switching DNS.
|
||||
|
||||
**Before the rollout**
|
||||
1. **Backups.** `talos/controlplane.yaml`, `talos/talosconfig`, the edge host key and the
|
||||
connector private keys exist only in this folder. Store them in a password manager (and put
|
||||
the folder under git, secrets stay gitignored).
|
||||
2. **Domains and DNS.** Pick the production domains; create DNS records for them (wildcards
|
||||
where needed) pointing to the VPS IPv4/IPv6.
|
||||
3. **HTTPS on the cluster side.** `cumulus` only has the plain connector (SSH, HTTP). For
|
||||
TLS passthrough it needs a second connector (SNI, PROXY v2), an Envoy HTTPS listener that
|
||||
accepts the PROXY header only from sish-client, and certificates (cert-manager with DNS-01, or
|
||||
HTTP-01 over the :80 route).
|
||||
4. **Remove test access.** Leave `connector-hello.pub` (local test stack) out of the
|
||||
production overlay; give each production connector its own key.
|
||||
|
||||
**Rollout**
|
||||
1. VPS: boot the Talos image (same factory schematic as dev) and check the provider's disk name
|
||||
and network (DHCP vs static, IPv6).
|
||||
2. Generate a new base config into its own folder (`talos/README.md`, "New node") and apply it
|
||||
with the existing patches plus one that disables the discovery service (single node, no
|
||||
external dependency needed).
|
||||
3. Bootstrap, fetch the kubeconfig. Verify against the firewall table: only the listed ports
|
||||
answer from outside.
|
||||
4. Kubernetes: new overlay `kubernetes/cirrus-prod` (copy of `cirrus-dev`) with the production
|
||||
`SISH_DOMAIN`/`SISH_BIND_HOSTS`, production connector keys and a new host key. Set
|
||||
`dnsPolicy: Default` for sish so the tunnel does not depend on CoreDNS. `diff`, then `apply -k`.
|
||||
5. Cluster side (`cumulus`): sish-client deployment(s) for the production edge (edge IP, new host
|
||||
key, production routes), matching Gateway listeners/routes and network policies. Keep the dev
|
||||
connectors until production is verified.
|
||||
6. Verify from outside: `ssh-keyscan` for SSH routes, `curl` for HTTP/HTTPS routes, the backend
|
||||
sees the real client IP, closed ports stay closed.
|
||||
7. Move existing services (e.g. from the Cloudflare tunnel) one hostname at a time by switching
|
||||
their DNS records.
|
||||
|
||||
**After the rollout**
|
||||
- External uptime check on the edge (sish :2222 and one route per protocol): the edge is a
|
||||
single point of failure for everything behind it.
|
||||
- Updates by hand: `talosctl upgrade` / `upgrade-k8s`, the sish image tag, sish-client tags in
|
||||
`cumulus`. The `talosconfig` admin certificate expires after one year.
|
||||
**Still open**
|
||||
- **Backups.** `coreos/.secrets/` (edge host key, `cumulus` connector key) and the Talos dev
|
||||
credentials exist only in this folder. Store them in a password manager.
|
||||
- **Uptime check.** External check on the edge (sish :5002 and one route per protocol): the edge is
|
||||
a single point of failure for everything behind it.
|
||||
- **Updates by hand.** Production updates are automatic (OS in a nightly reboot window, sish within
|
||||
v2, see `coreos/README.md`). sish-client tags in `cumulus` and the dev edge (`talosctl upgrade` /
|
||||
`upgrade-k8s`, sish image tag) are updated by hand; the `talosconfig` admin certificate expires
|
||||
after one year.
|
||||
|
||||
Reference in New Issue
Block a user