Skip to content

fix(bnk-vlans): retry kubectl apply on webhook-transient errors - #65

Merged
JLCode-tech merged 1 commit into
release/2.2from
fix/vlans-webhook-wait
Jun 29, 2026
Merged

JLCode-tech merged 1 commit into
release/2.2from
fix/vlans-webhook-wait

Conversation

@bonnyr-f5

Copy link
Copy Markdown
Collaborator

Summary

Replaces the fixed 3x-sleep retry (introduced in #54) with a bounded retry loop (300s budget, 15s interval) that distinguishes webhook-transient errors from real failures.

  • Retries only on: connection refused, no endpoints available, failed to call webhook, i/o timeout, context deadline exceeded.
  • Fails fast on anything else (validation rejection, kubeconfig errors) so real bugs aren't masked.

Motivation

During a dual_dpu_obmc deploy on mgx-19 (2026-05-19), the 3-attempt approach from #54 still wasn't enough on a cold-start f5-cne-controller pod. The endpoint registers (kubelet readiness probe passes) before the in-container TLS server binds port 3340, and the 15s + 30s + 30s schedule lands inside that gap. Manual kubectl apply worked after another ~60s; the deploy succeeded on UI retry.

The fixed retry schedule has two failure modes:

  1. The TLS port can take longer than 75s to bind on a cold cluster
  2. Any retry on a non-transient error wastes budget and obscures real failures

The bounded retry with error-class matching addresses both.

Test plan

  • Re-deploy on a cold cluster where f5-cne-controller pod cold-starts; verify the loop reports "Webhook transient at Xs:" log lines and then succeeds
  • Verify that a webhook 503 surfaces an error within ~300s rather than hanging
  • Verify a validation error (e.g. bad VLAN spec) fails fast rather than burning the full budget

Note for reviewer

This branch is currently behind release/2.2 (merge-base predates #52 and #54). Net diff against current release/2.2 is the single bounded-retry replacement in bnk/bnk-vlans/main.tf. Rebase before merging if you want a linear history.

🤖 Generated with Claude Code

Endpoint registration fires on kubelet's readiness probe, which on a
cold-start cne-controller pod precedes the in-container TLS server
binding port 3340 by ~2-3 minutes. The previous 150s endpoint wait +
single apply attempt routinely failed against "connection refused"
during that gap. Replace the single attempt with a bounded retry loop
that only retries on webhook-transient errors (connection refused,
no endpoints available, failed to call webhook, i/o timeout,
context deadline exceeded) and fails fast on validation rejections or
kubeconfig errors so real bugs aren't masked. Budget: 300s @ 15s
intervals.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@bonnyr-f5
bonnyr-f5 force-pushed the fix/vlans-webhook-wait branch from 7919cf0 to 2b36e07 Compare May 19, 2026 13:24
@bonnyr-f5
bonnyr-f5 requested a review from JLCode-tech May 19, 2026 13:25

@JLCode-tech JLCode-tech left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Lgtm

@JLCode-tech
JLCode-tech merged commit db63b5c into release/2.2 Jun 29, 2026
2 checks passed
@JLCode-tech
JLCode-tech deleted the fix/vlans-webhook-wait branch June 29, 2026 06:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants