Conversation
- load_metric: requests (legacy default) / tokens / kv_tokens (per-instance occupancy normalized by measured or configured KV capacity) - Per-request reservation accounting survives out-of-order completions; output-length EMA clamped by max_new_tokens budget - VLLMMetricsPoller: staggered /metrics scraping, lenient Prometheus parsing, additive bias fusion, hysteresis admission, preemption TTL - SharedTokenLoadState: cross-rank load publication with heartbeat TTL to prevent multi-rank herding - Prefix affinity (prompt_id -> instance LRU) subordinate to readiness/queue/capacity/admission filters - Learned state (EMA, affinity, LA-MLFQ history) export/import API
- Engine-owned /metrics poller and shared load state survive scheduler rebinding; stopped and re-created on whole-engine replacement - Per-instance KV capacity calibration from fresh profiling snapshots - RR fallback routes debit their reservation (ghost-dec fix) and skip blocked C-MLFQ state; fallback picks only selectable ready instances - reconfigure_from_plan refuses to run with in-flight requests or pending scout futures; trainer rebind carries scheduler learned state - wait_until_idle covers scout-follower futures; generate_batch propagates BaseException (cancelled samples)
- Every workflow now passes input_tokens from segments it already tokenized (engine falls back to chars/4 only for unknown callers) - prompt_id is forwarded unconditionally so prefix affinity works with non-C-MLFQ schedulers (agentic falls back to a prompt hash) - Tool-wait cancellation releases C-MLFQ state (CancelledError is a BaseException; asyncio import was missing from the handlers)
- tests/test_scheduling_audit.py: behavioral regressions for reservation accounting, C-MLFQ budget compat, per-instance capacities, metrics feed robustness (counter resets, NaN, stale admission), shared-state close/cross-process visibility, affinity interactions, scout cancellation, workflow id forwarding - examples/validate_rollout_scheduling.py: end-to-end HTTP validator (deployment injected at runtime, server-side tokenize, JSON report with routing decisions and vLLM prefix-cache deltas) - configs/scheduling_adaptive.yaml: environment-agnostic preset (kv_tokens + metrics feedback + prefix affinity) - Document rollout scheduling options in configuration_reference
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Upgrades heterogeneous rollout scheduling from request-count load balancing to a three-layer signal system, and fixes correctness issues found in an audit of the live rollout call chain.
Problems solved
min(active_requests)ignored request weight and bucket capacity — a 20K-token agentic trajectory counted the same as a 200-token answer, and small TP-1 buckets were compared against TP-4 buckets on raw counts.wait_until_idle), weight-sync rebinding reset learned state every step, kv-ratio selection mixed units, and partial capacity tables silently starved large buckets.Approach
max_new_tokens) retired by reservation ID, so out-of-order completions can't steal load from in-flight requests; counters conserve to exactly zero.load_metric: kv_tokens) — occupancy ratios per instance, with capacity resolved from vLLM's own profiledkv_cache_size_tokens(via/metrics) first, then last profiled value, then analytic/config estimate./metricsfeedback — engine-layer staggered poller (measured <1% latency impact even at 5Hz); additive bias fusion (local + EMA(observed − local)) instead of max(); hysteresis admission (block >0.90, unblock <0.75) with preemption-event TTL (counters are cumulative — events expire, no permanent blacklisting); graceful degradation to open-loop on feed outage.prompt_id → instanceLRU so repeated prompts hit warm prefix caches; strictly subordinate to readiness/queue/capacity/admission filters (blocked or full sticky instances are bypassed and remapped). Live validation on a TP1+TP2 cluster: cache hits concentrated +2304/0 on the sticky instance vs +1152/+1152 scattered with affinity off.All features are config-gated and default-off (
load_metric: requestspreserves legacy behavior). Verified by behavioral regression tests (tests/test_scheduling_audit.py), the scheduler smoke suite, and a reproducible live-HTTP validator (examples/validate_rollout_scheduling.py) with a JSON report of routing decisions and real vLLM cache-counter deltas.