Skip to content

feat(scheduling): workload- and capacity-aware routing for heterogeneous rollout buckets - #13

Open
shiinakkk wants to merge 5 commits into
NetX-lab:mainfrom
shiinakkk:main
Open

shiinakkk wants to merge 5 commits into
NetX-lab:mainfrom
shiinakkk:main

Conversation

@shiinakkk

Copy link
Copy Markdown

Summary

Upgrades heterogeneous rollout scheduling from request-count load balancing to a three-layer signal system, and fixes correctness issues found in an audit of the live rollout call chain.

Problems solved

  • Blind load balancing: routing by min(active_requests) ignored request weight and bucket capacity — a 20K-token agentic trajectory counted the same as a 200-token answer, and small TP-1 buckets were compared against TP-4 buckets on raw counts.
  • No feedback from vLLM: schedulers never saw real KV occupancy, waiting queues, or preemptions; estimates drifted silently.
  • Multi-rank herding: each training rank ranked instances by its own partial view, so N ranks stampeded the same "idle" instance.
  • Wasted prefill from scattering: n_samples repeats and multi-turn replays were spread across instances, defeating vLLM's prefix cache.
  • Accounting/state bugs: scheduler-failure fallback double-decremented unrelated requests (fooling wait_until_idle), weight-sync rebinding reset learned state every step, kv-ratio selection mixed units, and partial capacity tables silently starved large buckets.

Approach

  1. Token-metered accounting — per-request reservation (prompt + EMA-expected output, clamped by max_new_tokens) retired by reservation ID, so out-of-order completions can't steal load from in-flight requests; counters conserve to exactly zero.
  2. KV-capacity normalization (load_metric: kv_tokens) — occupancy ratios per instance, with capacity resolved from vLLM's own profiled kv_cache_size_tokens (via /metrics) first, then last profiled value, then analytic/config estimate.
  3. Closed-loop /metrics feedback — engine-layer staggered poller (measured <1% latency impact even at 5Hz); additive bias fusion (local + EMA(observed − local)) instead of max(); hysteresis admission (block >0.90, unblock <0.75) with preemption-event TTL (counters are cumulative — events expire, no permanent blacklisting); graceful degradation to open-loop on feed outage.
  4. Cross-rank aggregation — each rank publishes per-instance load via atomic-rename files + heartbeat TTL; ranking uses cluster-wide totals (owner-token guarded against rebind races).
  5. Prefix affinity — bounded prompt_id → instance LRU so repeated prompts hit warm prefix caches; strictly subordinate to readiness/queue/capacity/admission filters (blocked or full sticky instances are bypassed and remapped). Live validation on a TP1+TP2 cluster: cache hits concentrated +2304/0 on the sticky instance vs +1152/+1152 scattered with affinity off.
  6. Exact token counts — all workflows now pass counts they already computed; unknown callers fall back to chars/4.

All features are config-gated and default-off (load_metric: requests preserves legacy behavior). Verified by behavioral regression tests (tests/test_scheduling_audit.py), the scheduler smoke suite, and a reproducible live-HTTP validator (examples/validate_rollout_scheduling.py) with a JSON report of routing decisions and real vLLM cache-counter deltas.

shiinakkk and others added 5 commits September 29, 2026 15:13
- load_metric: requests (legacy default) / tokens / kv_tokens (per-instance
  occupancy normalized by measured or configured KV capacity)
- Per-request reservation accounting survives out-of-order completions;
  output-length EMA clamped by max_new_tokens budget
- VLLMMetricsPoller: staggered /metrics scraping, lenient Prometheus
  parsing, additive bias fusion, hysteresis admission, preemption TTL
- SharedTokenLoadState: cross-rank load publication with heartbeat TTL
  to prevent multi-rank herding
- Prefix affinity (prompt_id -> instance LRU) subordinate to
  readiness/queue/capacity/admission filters
- Learned state (EMA, affinity, LA-MLFQ history) export/import API
- Engine-owned /metrics poller and shared load state survive scheduler
  rebinding; stopped and re-created on whole-engine replacement
- Per-instance KV capacity calibration from fresh profiling snapshots
- RR fallback routes debit their reservation (ghost-dec fix) and skip
  blocked C-MLFQ state; fallback picks only selectable ready instances
- reconfigure_from_plan refuses to run with in-flight requests or
  pending scout futures; trainer rebind carries scheduler learned state
- wait_until_idle covers scout-follower futures; generate_batch
  propagates BaseException (cancelled samples)
- Every workflow now passes input_tokens from segments it already
  tokenized (engine falls back to chars/4 only for unknown callers)
- prompt_id is forwarded unconditionally so prefix affinity works with
  non-C-MLFQ schedulers (agentic falls back to a prompt hash)
- Tool-wait cancellation releases C-MLFQ state (CancelledError is a
  BaseException; asyncio import was missing from the handlers)
- tests/test_scheduling_audit.py: behavioral regressions for reservation
  accounting, C-MLFQ budget compat, per-instance capacities, metrics
  feed robustness (counter resets, NaN, stale admission), shared-state
  close/cross-process visibility, affinity interactions, scout
  cancellation, workflow id forwarding
- examples/validate_rollout_scheduling.py: end-to-end HTTP validator
  (deployment injected at runtime, server-side tokenize, JSON report
  with routing decisions and vLLM prefix-cache deltas)
- configs/scheduling_adaptive.yaml: environment-agnostic preset
  (kv_tokens + metrics feedback + prefix affinity)
- Document rollout scheduling options in configuration_reference
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant