Repository navigation
Overhaul K3S-Deploy (k3sup + kube-vip + MetalLB) - #170
Merged
Merged
Conversation
Regenerated from ghcr.io/kube-vip/kube-vip:v1.2.1 (env schema changed since 0.x: vip_cidr->vip_subnet, dropped vip_ddns/svc_*, added dhcp_mode/dns_mode/vip_nodename). Uses REPLACE_INTERFACE/REPLACE_VIP tokens.
The kubeadm default mounts /etc/kubernetes/admin.conf, which does not exist on k3s. --inCluster uses the kube-vip ServiceAccount (provided by the applied rbac.yaml) instead, which is the correct mode for k3s.
- pin k3s v1.35.6+k3s1 (channel alternative documented), MetalLB v0.16.0 - kube-vip v1.2.1 control-plane VIP placed on all masters; kubeconfig -> VIP - drop redundant kube-vip cloud-provider; single MetalLB native manifest - k3sup node-token (fetched once) + k3sup ready - set -euo pipefail, ssh-keyscan known_hosts (no ~/.ssh/config clobber), per-node NTP + apt-guarded prereqs, arch-aware kubectl, re-run safe - pre-flight config summary + [y/N] gate (ASSUME_YES bypass) - non-blinking step/info/ok/warn/err output; RAW_BASE override for testing Closes #40 #62 #68 #78 #80 #159 #162; supersedes #66 #89 #145
- fall back ready_nodes=0 if kubectl fails, so the summary still prints - tolerate EOF on the pre-flight prompt so it aborts cleanly
- rewrite only the server URL (https://master1:6443 -> VIP) instead of a global regex substitution whose dots could over-match cert data - poll the VIP /readyz after repointing the kubeconfig, so a broken kube-vip fails early with a clear message instead of downstream
Found running the deploy end-to-end on a multi-node cluster: - k3sup: get.k3sup.dev leaves an arch-suffixed binary (k3sup-arm64) when run unprivileged; install whichever binary it produced, not a literal 'k3sup' that may not exist. - kube-vip: scp landed the manifest in the ssh user's home, but the mv ran as root (sudo bash -s) where ~ is /root; scp to /tmp and mv from the absolute path instead.
…sup node-token/ready, non-blinking output, VIP readiness)
- match the global README.md convention; update the reference in k3s.sh - pad the options table so it reads cleanly in raw and rendered views, kept under the 80-col markdownlint limit
- join all nodes to --server-ip $vip (the readiness-gated VIP) so no single master is the registration SPOF; drop the dead --server-user on joins (a --node-token is already supplied, so k3sup doesn't SSH the server) - add --tls-san=$vip to every server's args (not just master1's install) so master2/3 declare the VIP SAN uniformly - pin --flannel-iface/--node-ip on workers too (correct on multi-NIC nodes) - filter k3sup's repeated 'slicervm.com' promo tip on install/join (sed, so pipefail still surfaces a real failure) - guard 'kubectl expose' with || true so a re-run doesn't abort on AlreadyExists Re-validated on the 6-VM Multipass e2e: 5 nodes Ready, VIP serves the API, MetalLB LoadBalancer reachable.
Add a post-join control-plane readiness gate so a half-succeeded server join on a degraded etcd can no longer pass as success. Retry the MetalLB IPAddressPool/L2Advertisement applies to ride out the validating-webhook startup race. Narrow the k3sup cleanup glob, quiet the final k3sup ready tip, and clarify two comments. README: add networking caveats (VIP/LB outside the DHCP pool, same-subnet ARP/L2, interface naming), certName key placement, a post-run verification section, machine-count/admin-box notes, and fix the RAW_BASE description.
Collaborator
Author
|
Pushed a round of hardening from an independent review pass over the script and README (c436d29). Script robustness
README
Still a draft; I'll re-run the multi-node e2e to exercise the readiness gate and the MetalLB retry live before marking it ready. |
The end-of-file-fixer pre-commit hook (CI) requires a single trailing newline; the kube-vip manifest had a blank line at EOF.
cyberops7
marked this pull request as ready for review
July 16, 2026 05:07
This was referenced Aug 22, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overhaul the K3S-Deploy script (k3sup + kube-vip + MetalLB)
Modernizes
Kubernetes/K3S-Deploy/to current versions and best practices, fixes long-standing bugs, hardens the script, and rewrites the README.Versions
v1.35.6+k3s1, kube-vipv1.2.1, MetalLBv0.16.0.kube-vipmanifest is regenerated from the pinned image with--inCluster(its env schema changed at 1.x, and--inClustermakes it authenticate via thekube-vipServiceAccount instead of the kubeadm/etc/kubernetes/admin.confpath, which does not exist on k3s).Highlights
localhost:8080/ API-unreachable class of errors).type: LoadBalancer); single MetalLB native manifest (self-creates its namespace) — no more mismatchedv0.12.1/v0.13.12applies.k3sup node-tokenfetched once and reused;k3sup readyfor cluster wait.set -euo pipefail;ssh-keyscanhost keys (no more clobbering~/.ssh/config); per-node time sync + apt-guarded prerequisites with a clear message on non-apt distros; arch-awarekubectl(fixes the hardcoded amd64 download); re-run safe.[y/N]gate (ASSUME_YES=1to bypass); cleaner non-blinking step output;RAW_BASEoverride for testing.Testing
bash -nclean, kubeconform-valid manifests.v1.35.6+k3s1, the kube-vip VIP serves the API, and MetalLB assigns a reachable LoadBalancer IP (curlto the nginx LB succeeds). Full deploy output: https://gist.github.com/cyberops7/5dcaabf65478f6a6375bd9fca75bb511How this was tested (full methodology)
Environment: Apple Silicon MacBook (arm64), Multipass 1.16.3, Ubuntu 24.04 images.
Topology: 6 VMs on the Multipass subnet — 1 admin box + 3 control-plane + 2 workers, each 2 vCPU / 2 GB RAM / 8 GB disk (admin 1 GB). The unmodified
k3s.shruns on the admin VM and deploys to the other 5 over SSH. Assertions run from the admin VM, which sits on the same L2 as the nodes, so the kube-vip VIP and the MetalLB LoadBalancer IP are actually reachable — not just allocated.1. Generate an SSH key and launch the VMs. A cloud-init injects the public key into each node's
ubuntuuser (which has passwordless sudo on the Multipass image).2. Stage the repo onto the admin VM and fill in the config block. Node IPs come from
multipass info; the VIP and LB range are taken from that subnet (e.g.192.168.252.0/24→ VIP.240,lbrange .241-.250); the interface comes from the node's default route.3. Run the script unmodified.
RAW_BASE=file://sources the sibling manifests from the local checkout, so the run does not depend on the branch being merged first.ASSUME_YES=1skips the prompt;NO_COLOR=1keeps the captured log clean.4. Assert from the admin VM (same L2 as the nodes, so the VIP and LB IP are reachable).
Teardown:
multipass delete <name>for each VM, thenmultipass purge.Result: PASS, ~5 minutes wall-clock (VM launch ~2m with the image cached; install + assertions ~3m). 5 nodes
Readyonv1.35.6+k3s1, the VIP serves the API, andcurlto the MetalLB-assigned nginx LoadBalancer IP succeeds.Closes / supersedes
Closes #40
Closes #68
Closes #78
Closes #80
Closes #162
Refs #62 — this fixes the dangerous
~/.ssh/configclobber ink3s.sh, but the same pattern also lives in the RKE2 / Docker-Swarm / Kubernetes-Lite scripts (not touched here), so #62 stays open as the repo-wide tracking issue.Supersedes #66, #89, #145 (please close in favor of this).