fix(hp-nodes): skip userdata reboot in kernel mode — fix EKS join failure - #56
Merged
Merged
Conversation
…lure
In kernel mode (default), the userdata's `nohup bash -c 'sleep 60; reboot' &`
fires concurrently with the EKS bootstrap script. Kubelet starts registering
the node, then reboot interrupts mid-registration. The nodegroup ends up
ACTIVE in AWS but the nodes never appear in `kubectl get nodes`, ultimately
flipping to NodeCreationFailure.
The reboot exists to make grub-set kernel cmdline params take effect
(hugepagesz, isolcpus, nohz_full, rcu_nocbs). But:
- Hugepages are already allocated at runtime via sysfs writes immediately
below this block, and persisted across future operator-triggered reboots
by the systemd-enabled dpdk-hugepages.service.
- isolcpus / nohz_full / rcu_nocbs are TMM jitter optimizations — TMM
still works correctly without them via cpuset pinning in the pod spec.
In sriov mode the reboot path is preserved unchanged: vfio-pci binding
genuinely requires a fresh kernel cmdline, and the dpdk-continuation service
coordinates a kubelet restart post-reboot.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes the HP-nodes EKS join failure where the nodegroup goes ACTIVE but nodes never appear in
kubectl get nodes, eventually flipping toNodeCreationFailure.Root cause: in kernel mode (the default since 2026-04), the userdata's
nohup bash -c 'sleep 60; reboot' &(line 218) fires concurrently with the EKS bootstrap script. Kubelet starts its first registration, then reboot interrupts mid-flight.Fix: in kernel mode, skip the grub mutation that motivates the reboot, and set
NEEDS_REBOOT=falseso the existing Phase 6 reboot path is bypassed. Sriov mode is untouched.Why this is safe in kernel mode
sysfswrites (lines 68–73 in the existing script).dpdk-hugepages.serviceis still installed andsystemctl enabled, so future operator-triggered reboots restore hugepages.isolcpus/nohz_full/rcu_nocbswould require a kernel cmdline + reboot to take effect, but they're TMM jitter optimizations only — TMM still functions via cpuset pinning in its pod spec.dpdk-continuation.servicecoordinates a post-rebootsystemctl restart kubelet).Diff scope
11 lines changed in
infra/aws/high-performance-nodes/scripts/compact_userdata.sh. No other files touched, no module.json / pack manifest changes — drop-in.Test plan
bnkforge.pack.jsonvalidation gate (PR ci: validate every bnkforge.pack.json against forge contract #51) passeskubectl get nodeswithin ~5 min of nodegroup ACTIVEaws eks describe-nodegrouphealth/var/log/dpdk-setup.logwe see "kernel mode: skipping grub mutation, no reboot will be scheduled"