Skip to content

Overhaul K3S-Deploy (k3sup + kube-vip + MetalLB) - #170

Merged
DefNotJeffrey merged 13 commits into
mainfrom
k3s-deploy-overhaul
Aug 22, 2026
Merged

DefNotJeffrey merged 13 commits into
mainfrom
k3s-deploy-overhaul

Conversation

@cyberops7

@cyberops7 cyberops7 commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Overhaul the K3S-Deploy script (k3sup + kube-vip + MetalLB)

Modernizes Kubernetes/K3S-Deploy/ to current versions and best practices, fixes long-standing bugs, hardens the script, and rewrites the README.

Versions

  • k3s v1.35.6+k3s1, kube-vip v1.2.1, MetalLB v0.16.0.
  • The kube-vip manifest is regenerated from the pinned image with --inCluster (its env schema changed at 1.x, and --inCluster makes it authenticate via the kube-vip ServiceAccount instead of the kubeadm /etc/kubernetes/admin.conf path, which does not exist on k3s).

Highlights

  • kube-vip runs on all control-plane nodes; the local kubeconfig points at the VIP (fixes the localhost:8080 / API-unreachable class of errors).
  • Removed the redundant kube-vip cloud-provider (MetalLB owns type: LoadBalancer); single MetalLB native manifest (self-creates its namespace) — no more mismatched v0.12.1 / v0.13.12 applies.
  • k3sup node-token fetched once and reused; k3sup ready for cluster wait.
  • Post-join control-plane readiness gate so a half-succeeded server join on a degraded etcd cannot pass as success; MetalLB CR applies retried to ride out the validating-webhook startup race.
  • set -euo pipefail; ssh-keyscan host keys (no more clobbering ~/.ssh/config); per-node time sync + apt-guarded prerequisites with a clear message on non-apt distros; arch-aware kubectl (fixes the hardcoded amd64 download); re-run safe.
  • Pre-flight config summary + [y/N] gate (ASSUME_YES=1 to bypass); cleaner non-blinking step output; RAW_BASE override for testing.

Testing

  • Static: shellcheck clean, bash -n clean, kubeconform-valid manifests.
  • MetalLB layer validated on a throwaway k3d cluster.
  • Full multi-node e2e passed (re-run on the latest commit) on a 6-VM Multipass cluster (3 control-plane + 2 workers + admin, Ubuntu 24.04, arm64): all 5 nodes Ready on v1.35.6+k3s1, the kube-vip VIP serves the API, and MetalLB assigns a reachable LoadBalancer IP (curl to the nginx LB succeeds). Full deploy output: https://gist.github.com/cyberops7/5dcaabf65478f6a6375bd9fca75bb511
How this was tested (full methodology)

Environment: Apple Silicon MacBook (arm64), Multipass 1.16.3, Ubuntu 24.04 images.

Topology: 6 VMs on the Multipass subnet — 1 admin box + 3 control-plane + 2 workers, each 2 vCPU / 2 GB RAM / 8 GB disk (admin 1 GB). The unmodified k3s.sh runs on the admin VM and deploys to the other 5 over SSH. Assertions run from the admin VM, which sits on the same L2 as the nodes, so the kube-vip VIP and the MetalLB LoadBalancer IP are actually reachable — not just allocated.

1. Generate an SSH key and launch the VMs. A cloud-init injects the public key into each node's ubuntu user (which has passwordless sudo on the Multipass image).

ssh-keygen -t ed25519 -N '' -f ./id_e2e -C k3s-e2e

# cloud-init.yaml:
#   #cloud-config
#   ssh_authorized_keys:
#     - <contents of id_e2e.pub>
#   ssh_pwauth: false

for n in admin master1 master2 master3 worker1 worker2; do
  multipass launch 24.04 --name "$n" --cpus 2 --memory 2G --disk 8G \
    --cloud-init cloud-init.yaml
done

2. Stage the repo onto the admin VM and fill in the config block. Node IPs come from multipass info; the VIP and LB range are taken from that subnet (e.g. 192.168.252.0/24 → VIP .240, lbrange .241-.250); the interface comes from the node's default route.

# Copy this repo's Kubernetes/K3S-Deploy/ and the e2e key to admin:~ , then
# edit the "YOU SHOULD ONLY NEED TO EDIT THIS SECTION" block in k3s.sh:
#   master1..3 / worker1..2 = the multipass IPs
#   user=ubuntu
#   interface=<node NIC, from: ip -o -4 route show default>
#   vip=<subnet>.240
#   lbrange=<subnet>.241-<subnet>.250

3. Run the script unmodified. RAW_BASE=file:// sources the sibling manifests from the local checkout, so the run does not depend on the branch being merged first. ASSUME_YES=1 skips the prompt; NO_COLOR=1 keeps the captured log clean.

multipass exec admin -- bash -c '
  export ASSUME_YES=1 NO_COLOR=1 RAW_BASE=file:///home/ubuntu/JimsGarage
  bash ~/JimsGarage/Kubernetes/K3S-Deploy/k3s.sh 2>&1 | tee ~/k3s-install.log'

4. Assert from the admin VM (same L2 as the nodes, so the VIP and LB IP are reachable).

multipass exec admin -- bash -c '
  export KUBECONFIG=~/.kube/config
  kubectl get nodes -o wide
  kubectl --request-timeout=10s get --raw=/readyz            # VIP serves the API
  lb=$(kubectl get svc nginx-1 -n default \
        -o jsonpath="{.status.loadBalancer.ingress[0].ip}")
  curl -fsS "http://$lb"                                     # MetalLB LB reachable'

Teardown: multipass delete <name> for each VM, then multipass purge.

Result: PASS, ~5 minutes wall-clock (VM launch ~2m with the image cached; install + assertions ~3m). 5 nodes Ready on v1.35.6+k3s1, the VIP serves the API, and curl to the MetalLB-assigned nginx LoadBalancer IP succeeds.

Closes / supersedes

Closes #40
Closes #68
Closes #78
Closes #80
Closes #162

Refs #62 — this fixes the dangerous ~/.ssh/config clobber in k3s.sh, but the same pattern also lives in the RKE2 / Docker-Swarm / Kubernetes-Lite scripts (not touched here), so #62 stays open as the repo-wide tracking issue.

Supersedes #66, #89, #145 (please close in favor of this).

cyberops7 added 12 commits July 5, 2026 21:13
Regenerated from ghcr.io/kube-vip/kube-vip:v1.2.1 (env schema changed
since 0.x: vip_cidr->vip_subnet, dropped vip_ddns/svc_*, added
dhcp_mode/dns_mode/vip_nodename). Uses REPLACE_INTERFACE/REPLACE_VIP tokens.
The kubeadm default mounts /etc/kubernetes/admin.conf, which does not
exist on k3s. --inCluster uses the kube-vip ServiceAccount (provided by
the applied rbac.yaml) instead, which is the correct mode for k3s.
- pin k3s v1.35.6+k3s1 (channel alternative documented), MetalLB v0.16.0
- kube-vip v1.2.1 control-plane VIP placed on all masters; kubeconfig -> VIP
- drop redundant kube-vip cloud-provider; single MetalLB native manifest
- k3sup node-token (fetched once) + k3sup ready
- set -euo pipefail, ssh-keyscan known_hosts (no ~/.ssh/config clobber),
  per-node NTP + apt-guarded prereqs, arch-aware kubectl, re-run safe
- pre-flight config summary + [y/N] gate (ASSUME_YES bypass)
- non-blinking step/info/ok/warn/err output; RAW_BASE override for testing

Closes #40 #62 #68 #78 #80 #159 #162; supersedes #66 #89 #145
- fall back ready_nodes=0 if kubectl fails, so the summary still prints
- tolerate EOF on the pre-flight prompt so it aborts cleanly
- rewrite only the server URL (https://master1:6443 -> VIP) instead of a
  global regex substitution whose dots could over-match cert data
- poll the VIP /readyz after repointing the kubeconfig, so a broken
  kube-vip fails early with a clear message instead of downstream
Found running the deploy end-to-end on a multi-node cluster:
- k3sup: get.k3sup.dev leaves an arch-suffixed binary (k3sup-arm64) when
  run unprivileged; install whichever binary it produced, not a literal
  'k3sup' that may not exist.
- kube-vip: scp landed the manifest in the ssh user's home, but the mv
  ran as root (sudo bash -s) where ~ is /root; scp to /tmp and mv from
  the absolute path instead.
…sup node-token/ready, non-blinking output, VIP readiness)
- match the global README.md convention; update the reference in k3s.sh
- pad the options table so it reads cleanly in raw and rendered views,
  kept under the 80-col markdownlint limit
- join all nodes to --server-ip $vip (the readiness-gated VIP) so no single
  master is the registration SPOF; drop the dead --server-user on joins
  (a --node-token is already supplied, so k3sup doesn't SSH the server)
- add --tls-san=$vip to every server's args (not just master1's install) so
  master2/3 declare the VIP SAN uniformly
- pin --flannel-iface/--node-ip on workers too (correct on multi-NIC nodes)
- filter k3sup's repeated 'slicervm.com' promo tip on install/join (sed, so
  pipefail still surfaces a real failure)
- guard 'kubectl expose' with || true so a re-run doesn't abort on AlreadyExists

Re-validated on the 6-VM Multipass e2e: 5 nodes Ready, VIP serves the API,
MetalLB LoadBalancer reachable.
Add a post-join control-plane readiness gate so a half-succeeded server
join on a degraded etcd can no longer pass as success. Retry the MetalLB
IPAddressPool/L2Advertisement applies to ride out the validating-webhook
startup race. Narrow the k3sup cleanup glob, quiet the final k3sup ready
tip, and clarify two comments.

README: add networking caveats (VIP/LB outside the DHCP pool, same-subnet
ARP/L2, interface naming), certName key placement, a post-run verification
section, machine-count/admin-box notes, and fix the RAW_BASE description.
@cyberops7

Copy link
Copy Markdown
Collaborator Author

Pushed a round of hardening from an independent review pass over the script and README (c436d29).

Script robustness

  • Added a post-join control-plane readiness gate. The VIP check in Step 5 only proves master1 is up (a single-member etcd), so a control-plane join that half-succeeds could previously sail through to the end on a degraded cluster. After the join loop the script now waits for all control-plane nodes to be Ready and asserts the expected count before continuing.
  • Wrapped the MetalLB IPAddressPool / L2Advertisement applies in a small retry. MetalLB's validating webhook has a brief window where the controller Deployment reports Available but the webhook endpoints aren't serving yet, so the first apply can intermittently fail with no endpoints available for service metallb-webhook-service and abort the run under set -e.
  • Minor: narrowed the k3sup cleanup to the binary the installer actually produced (instead of a k3sup-* glob in the CWD), quieted the last k3sup promo tip on k3sup ready, and tidied two comments.

README

  • Added an "Important caveats" section covering the most common failure modes: the VIP and MetalLB range must be outside the DHCP pool and non-overlapping, kube-vip (ARP) and MetalLB (L2) require the same subnet, and interface must match the node's real NIC (often ens18/enp0s3, not eth0).
  • Clarified certName (filename only, where to place the key), flagged the example IPs as placeholders, and added a post-run verification section (kubeconfig merge + k3s-ha context, worker labels, per-node recovery with the k3s uninstall scripts).
  • Fixed the RAW_BASE description and the machine-count / admin-box notes.

Still a draft; I'll re-run the multi-node e2e to exercise the readiness gate and the MetalLB retry live before marking it ready.

The end-of-file-fixer pre-commit hook (CI) requires a single trailing
newline; the kube-vip manifest had a blank line at EOF.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants