Skip to content

Frontier attn_prefill prediction shows large TTFT deviation at long prefill, blocks 64K scenarios #26

Description

@D-Shirley

Context

We are using Frontier (pre-release-v0.2, co-location mode) to predict the performance of Qwen3-30B-A3B-Instruct-2507 on H20 GPUs, cross-validating against vLLM 0.10.2 benchmarks.

Our target application range: prefill ~64K tokens, decode ~800 tokens, QPS=1. Currently in baseline accuracy validation (prefill 64->4096), we observe deviation direction flipping with prefill scale.


Setup

  • Model: Qwen3-30B-A3B-Instruct-2507
  • GPU: H20 (8x per node)
  • Parallelism: ATTN_TP=4, MOE_EP=8, DP=2
  • NCCL: intra_node_bandwidth_gbps=3600
  • Scheduler: enable_chunked_prefill=false, max_num_batched_tokens=16384
  • Frontier: co-location, sequential simulator, MARS predictor for compute, cuda_event profiling
  • vLLM: v0.10.2, v1 engine, MONOLITHIC scheduler

12-Workload Full Comparison (QPS=2.0, 100 requests, Poisson arrival)

Workload pf/dc TTFT mean (F/V) TTFT p95 (F/V) TPOT mean (F/V) TPOT p95 (F/V) RPS (F/V) tok/s (F/V)
pf64_dc1024 64/1024 46.6/33.4 (+39%) 68.7/37.9 (+81%) 7.1/10.6 (-33%) 8.4/13.7 (-39%) 1.99/1.60 2041/1638 (+25%)
pf128_dc1024 128/1024 50.7/33.2 (+53%) 75.0/37.2 (+102%) 7.2/10.6 (-32%) 8.4/13.5 (-38%) 1.99/1.60 2040/1638 (+25%)
decode_light 256/512 58.0/34.8 (+67%) 80.6/39.8 (+103%) 7.1/10.7 (-34%) 8.0/11.8 (-32%) 2.14/1.82 1095/931 (+18%)
decode_medium 256/1024 59.6/37.9 (+57%) 88.3/45.4 (+94%) 8.0/13.4 (-40%) 8.4/14.2 (-40%) 1.99/1.60 2039/1641 (+24%)
decode_heavy 256/2048 59.0/39.9 (+48%) 88.9/49.0 (+81%) 8.4/14.7 (-43%) 8.7/15.4 (-44%) 1.71/1.28 3511/2615 (+34%)
extreme_decode 256/4096 59.8/39.6 (+51%) 89.5/51.7 (+73%) 8.5/16.7 (-49%) 8.8/17.3 (-49%) 1.33/0.85 5431/3493 (+55%)
mild_decode 512/1536 69.2/94.0 (-26%) 108.7/145.0 (-25%) 8.5/15.5 (-45%) 8.9/16.9 (-48%) 1.84/1.42 2825/2184 (+29%)
balanced 1024/1024 74.5/96.0 (-22%) 115.4/138.3 (-17%) 8.4/14.9 (-44%) 8.9/16.8 (-47%) 1.99/1.52 2033/1553 (+31%)
deep_decode 1024/2048 75.0/93.4 (-20%) 116.9/122.9 (-5%) 8.6/16.3 (-47%) 9.1/17.4 (-48%) 1.71/1.25 3498/2570 (+36%)
prefill_heavy 2048/256 84.3/112.8 (-25%) 136.0/155.9 (-13%) 6.1/11.8 (-49%) 9.3/14.1 (-34%) 2.18/1.90 559/487 (+15%)
extreme_prefill 2048/1024 86.2/117.2 (-27%) 134.2/174.8 (-23%) 8.8/15.8 (-45%) 9.3/17.8 (-48%) 1.98/1.58 2026/1616 (+25%)
pf4096_dc1024 4096/1024 131.3/272.4 (-52%) 208.0/375.3 (-45%) 6.8/15.7 (-57%) 9.3/20.9 (-56%) 1.97/1.55 2016/1586 (+27%)

F/V = Frontier / vLLM. Delta = (F-V)/V. TTFT/TPOT in ms. tok/s = decode token throughput.

Key observations:

  • Deviation direction flips: prefill <= 256: Frontier TTFT pessimistic (+39% to +103%); prefill >= 512: optimistic (-5% to -52%)
  • TPOT systematically underestimated: 12 workloads TPOT mean delta -32% to -57%, p95 -32% to -56%
  • 4096 prefill worst for TTFT: TTFT mean -52%, p95 -45%

P0 Low-QPS Validation (QPS=0.2, no queuing)

Workload Frontier p50 vLLM p50 delta
pf128_dc1024 48.1ms 43.5ms +10.6%
pf256_dc1024 53.2ms 36.0ms +47.8%
pf512_dc1536 61.7ms 96.9ms -36.3%
pf4096_dc1024 103.9ms 189.5ms -45.2%

Direction unchanged after eliminating queuing (<=256 pessimistic, >=512 optimistic)
-> confirmed as compute-path issue.


Questions

  1. We observe TTFT deviation direction flipping with prefill scale: pessimistic at <=256 (mean +39%+103%, p95 +73%+103%), optimistic at >=512 (mean -5%-52%, p95 -5%-45%), worst at 4096. Has this pattern been reported before? Any data for >4096 or <256 prefill ranges?

  2. Frontier decomposes the prefill step into a dozen operators (attn_pre_proj, attn_prefill, moe_grouped_gemm, EP alltoall, etc.), profiles each independently, and sums them up. In vLLM production, these kernels are fused via CUDA Graph replay. We suspect this architectural difference is the primary cause of the -52% TTFT deviation at 4096 prefill. What was the design reasoning behind the per-operator decomposition approach, and how was the gap with CUDA Graph fused execution considered? Are there plans to address this discrepancy?

  3. While debugging, we noticed attention uses two features: kv_cache_size and prefill_chunk_size_squared (frontier/attention/profiling_mapping.py:229-232). However, attention computation is query x context = prefill_chunk_size x (kv_cache_size + prefill_chunk_size), not prefill_chunk_size^2. Is there a specific reason for using the squared form rather than prefill_chunk_size x context_length? Is this related to how the profiler measures attention time?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions