Skip to content

fix(hp-nodes): skip userdata reboot in kernel mode — fix EKS join failure - #56

Merged
JLCode-tech merged 1 commit into
release/2.2from
fix/hp-nodes-kernel-mode-no-reboot
May 1, 2026
Merged

JLCode-tech merged 1 commit into
release/2.2from
fix/hp-nodes-kernel-mode-no-reboot

Conversation

@JLCode-tech

Copy link
Copy Markdown
Owner

Summary

Fixes the HP-nodes EKS join failure where the nodegroup goes ACTIVE but nodes never appear in kubectl get nodes, eventually flipping to NodeCreationFailure.

Root cause: in kernel mode (the default since 2026-04), the userdata's nohup bash -c 'sleep 60; reboot' & (line 218) fires concurrently with the EKS bootstrap script. Kubelet starts its first registration, then reboot interrupts mid-flight.

Fix: in kernel mode, skip the grub mutation that motivates the reboot, and set NEEDS_REBOOT=false so the existing Phase 6 reboot path is bypassed. Sriov mode is untouched.

Why this is safe in kernel mode

  • Hugepages are already allocated at runtime via sysfs writes (lines 68–73 in the existing script).
  • The dpdk-hugepages.service is still installed and systemctl enabled, so future operator-triggered reboots restore hugepages.
  • isolcpus / nohz_full / rcu_nocbs would require a kernel cmdline + reboot to take effect, but they're TMM jitter optimizations only — TMM still functions via cpuset pinning in its pod spec.
  • In sriov mode the original reboot path is preserved verbatim (vfio-pci binding genuinely requires a fresh cmdline, and dpdk-continuation.service coordinates a post-reboot systemctl restart kubelet).

Diff scope

11 lines changed in infra/aws/high-performance-nodes/scripts/compact_userdata.sh. No other files touched, no module.json / pack manifest changes — drop-in.

Test plan

  • CI bnkforge.pack.json validation gate (PR ci: validate every bnkforge.pack.json against forge contract #51) passes
  • After merge: forge resyncs catalog, picks up new HP-nodes commit hash
  • Fresh project, AWS EKS Foundation blueprint, click Deploy All
  • HP nodegroup goes ACTIVE and nodes appear in kubectl get nodes within ~5 min of nodegroup ACTIVE
  • No NodeCreationFailure in aws eks describe-nodegroup health
  • Reboot is NOT scheduled — verify in /var/log/dpdk-setup.log we see "kernel mode: skipping grub mutation, no reboot will be scheduled"

…lure

In kernel mode (default), the userdata's `nohup bash -c 'sleep 60; reboot' &`
fires concurrently with the EKS bootstrap script. Kubelet starts registering
the node, then reboot interrupts mid-registration. The nodegroup ends up
ACTIVE in AWS but the nodes never appear in `kubectl get nodes`, ultimately
flipping to NodeCreationFailure.

The reboot exists to make grub-set kernel cmdline params take effect
(hugepagesz, isolcpus, nohz_full, rcu_nocbs). But:

  - Hugepages are already allocated at runtime via sysfs writes immediately
    below this block, and persisted across future operator-triggered reboots
    by the systemd-enabled dpdk-hugepages.service.
  - isolcpus / nohz_full / rcu_nocbs are TMM jitter optimizations — TMM
    still works correctly without them via cpuset pinning in the pod spec.

In sriov mode the reboot path is preserved unchanged: vfio-pci binding
genuinely requires a fresh kernel cmdline, and the dpdk-continuation service
coordinates a kubelet restart post-reboot.
@JLCode-tech
JLCode-tech merged commit 98ed98d into release/2.2 May 1, 2026
2 checks passed
@JLCode-tech
JLCode-tech deleted the fix/hp-nodes-kernel-mode-no-reboot branch May 1, 2026 05:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant