Context
We are using Frontier (pre-release-v0.2, co-location mode) to predict the performance of Qwen3-30B-A3B-Instruct-2507 on H20 GPUs, cross-validating against vLLM 0.10.2 benchmarks.
Our target application range: prefill ~64K tokens, decode ~800 tokens, QPS=1. Currently in baseline accuracy validation (prefill 64->4096), we observe deviation direction flipping with prefill scale.
Setup
- Model: Qwen3-30B-A3B-Instruct-2507
- GPU: H20 (8x per node)
- Parallelism: ATTN_TP=4, MOE_EP=8, DP=2
- NCCL: intra_node_bandwidth_gbps=3600
- Scheduler: enable_chunked_prefill=false, max_num_batched_tokens=16384
- Frontier: co-location, sequential simulator, MARS predictor for compute, cuda_event profiling
- vLLM: v0.10.2, v1 engine, MONOLITHIC scheduler
12-Workload Full Comparison (QPS=2.0, 100 requests, Poisson arrival)
| Workload |
pf/dc |
TTFT mean (F/V) |
TTFT p95 (F/V) |
TPOT mean (F/V) |
TPOT p95 (F/V) |
RPS (F/V) |
tok/s (F/V) |
| pf64_dc1024 |
64/1024 |
46.6/33.4 (+39%) |
68.7/37.9 (+81%) |
7.1/10.6 (-33%) |
8.4/13.7 (-39%) |
1.99/1.60 |
2041/1638 (+25%) |
| pf128_dc1024 |
128/1024 |
50.7/33.2 (+53%) |
75.0/37.2 (+102%) |
7.2/10.6 (-32%) |
8.4/13.5 (-38%) |
1.99/1.60 |
2040/1638 (+25%) |
| decode_light |
256/512 |
58.0/34.8 (+67%) |
80.6/39.8 (+103%) |
7.1/10.7 (-34%) |
8.0/11.8 (-32%) |
2.14/1.82 |
1095/931 (+18%) |
| decode_medium |
256/1024 |
59.6/37.9 (+57%) |
88.3/45.4 (+94%) |
8.0/13.4 (-40%) |
8.4/14.2 (-40%) |
1.99/1.60 |
2039/1641 (+24%) |
| decode_heavy |
256/2048 |
59.0/39.9 (+48%) |
88.9/49.0 (+81%) |
8.4/14.7 (-43%) |
8.7/15.4 (-44%) |
1.71/1.28 |
3511/2615 (+34%) |
| extreme_decode |
256/4096 |
59.8/39.6 (+51%) |
89.5/51.7 (+73%) |
8.5/16.7 (-49%) |
8.8/17.3 (-49%) |
1.33/0.85 |
5431/3493 (+55%) |
| mild_decode |
512/1536 |
69.2/94.0 (-26%) |
108.7/145.0 (-25%) |
8.5/15.5 (-45%) |
8.9/16.9 (-48%) |
1.84/1.42 |
2825/2184 (+29%) |
| balanced |
1024/1024 |
74.5/96.0 (-22%) |
115.4/138.3 (-17%) |
8.4/14.9 (-44%) |
8.9/16.8 (-47%) |
1.99/1.52 |
2033/1553 (+31%) |
| deep_decode |
1024/2048 |
75.0/93.4 (-20%) |
116.9/122.9 (-5%) |
8.6/16.3 (-47%) |
9.1/17.4 (-48%) |
1.71/1.25 |
3498/2570 (+36%) |
| prefill_heavy |
2048/256 |
84.3/112.8 (-25%) |
136.0/155.9 (-13%) |
6.1/11.8 (-49%) |
9.3/14.1 (-34%) |
2.18/1.90 |
559/487 (+15%) |
| extreme_prefill |
2048/1024 |
86.2/117.2 (-27%) |
134.2/174.8 (-23%) |
8.8/15.8 (-45%) |
9.3/17.8 (-48%) |
1.98/1.58 |
2026/1616 (+25%) |
| pf4096_dc1024 |
4096/1024 |
131.3/272.4 (-52%) |
208.0/375.3 (-45%) |
6.8/15.7 (-57%) |
9.3/20.9 (-56%) |
1.97/1.55 |
2016/1586 (+27%) |
F/V = Frontier / vLLM. Delta = (F-V)/V. TTFT/TPOT in ms. tok/s = decode token throughput.
Key observations:
- Deviation direction flips: prefill <= 256: Frontier TTFT pessimistic (+39% to +103%); prefill >= 512: optimistic (-5% to -52%)
- TPOT systematically underestimated: 12 workloads TPOT mean delta -32% to -57%, p95 -32% to -56%
- 4096 prefill worst for TTFT: TTFT mean -52%, p95 -45%
P0 Low-QPS Validation (QPS=0.2, no queuing)
| Workload |
Frontier p50 |
vLLM p50 |
delta |
| pf128_dc1024 |
48.1ms |
43.5ms |
+10.6% |
| pf256_dc1024 |
53.2ms |
36.0ms |
+47.8% |
| pf512_dc1536 |
61.7ms |
96.9ms |
-36.3% |
| pf4096_dc1024 |
103.9ms |
189.5ms |
-45.2% |
Direction unchanged after eliminating queuing (<=256 pessimistic, >=512 optimistic)
-> confirmed as compute-path issue.
Questions
-
We observe TTFT deviation direction flipping with prefill scale: pessimistic at <=256 (mean +39%+103%, p95 +73%+103%), optimistic at >=512 (mean -5%-52%, p95 -5%-45%), worst at 4096. Has this pattern been reported before? Any data for >4096 or <256 prefill ranges?
-
Frontier decomposes the prefill step into a dozen operators (attn_pre_proj, attn_prefill, moe_grouped_gemm, EP alltoall, etc.), profiles each independently, and sums them up. In vLLM production, these kernels are fused via CUDA Graph replay. We suspect this architectural difference is the primary cause of the -52% TTFT deviation at 4096 prefill. What was the design reasoning behind the per-operator decomposition approach, and how was the gap with CUDA Graph fused execution considered? Are there plans to address this discrepancy?
-
While debugging, we noticed attention uses two features: kv_cache_size and prefill_chunk_size_squared (frontier/attention/profiling_mapping.py:229-232). However, attention computation is query x context = prefill_chunk_size x (kv_cache_size + prefill_chunk_size), not prefill_chunk_size^2. Is there a specific reason for using the squared form rather than prefill_chunk_size x context_length? Is this related to how the profiler measures attention time?
Context
We are using Frontier (pre-release-v0.2, co-location mode) to predict the performance of Qwen3-30B-A3B-Instruct-2507 on H20 GPUs, cross-validating against vLLM 0.10.2 benchmarks.
Our target application range: prefill ~64K tokens, decode ~800 tokens, QPS=1. Currently in baseline accuracy validation (prefill 64->4096), we observe deviation direction flipping with prefill scale.
Setup
12-Workload Full Comparison (QPS=2.0, 100 requests, Poisson arrival)
Key observations:
P0 Low-QPS Validation (QPS=0.2, no queuing)
Direction unchanged after eliminating queuing (<=256 pessimistic, >=512 optimistic)
-> confirmed as compute-path issue.
Questions
We observe TTFT deviation direction flipping with prefill scale: pessimistic at <=256 (mean +39%
+103%, p95 +73%+103%), optimistic at >=512 (mean -5%-52%, p95 -5%-45%), worst at 4096. Has this pattern been reported before? Any data for >4096 or <256 prefill ranges?Frontier decomposes the prefill step into a dozen operators (attn_pre_proj, attn_prefill, moe_grouped_gemm, EP alltoall, etc.), profiles each independently, and sums them up. In vLLM production, these kernels are fused via CUDA Graph replay. We suspect this architectural difference is the primary cause of the -52% TTFT deviation at 4096 prefill. What was the design reasoning behind the per-operator decomposition approach, and how was the gap with CUDA Graph fused execution considered? Are there plans to address this discrepancy?
While debugging, we noticed attention uses two features:
kv_cache_sizeandprefill_chunk_size_squared(frontier/attention/profiling_mapping.py:229-232). However, attention computation is query x context = prefill_chunk_size x (kv_cache_size + prefill_chunk_size), not prefill_chunk_size^2. Is there a specific reason for using the squared form rather thanprefill_chunk_size x context_length? Is this related to how the profiler measures attention time?