Enable bounded non-windowed GQA workspace estimation - #32696
Chi Lo (chilo-ms) wants to merge 1 commit into
Conversation
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The new kernel-declaration test disables every backend route compatible with its model and will fail.
Get a fresh assessment by requesting another Copilot review.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Adds bounded Level-2 workspace estimation for non-windowed CUDA GroupQueryAttention.
Changes:
- Adds a session-configured total-sequence-length bound.
- Models one-sided cache alias preservation.
- Adds estimator, aggregation, declaration tests, and documentation.
| File | Description |
|---|---|
group_query_attention.h |
Stores the configured bound. |
group_query_attention.cc |
Parses and applies the bound. |
group_query_attention_workspace_estimate.h |
Extends estimator configuration. |
group_query_attention_workspace_estimate.cc |
Builds non-windowed bounds. |
group_query_attention_workspace_bounds.h |
Adds past-capacity and alias metadata. |
group_query_attention_workspace_bounds.cc |
Aggregates partial-alias workspace routes. |
group_query_attention_workspace_estimate_test.cc |
Tests estimation and declaration behavior. |
onnxruntime_session_options_config_keys.h |
Defines the new session option. |
attention_workspace_estimation.md |
Documents non-windowed estimation. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| GTEST_SKIP() << "A CUDA device is required to construct the CUDA kernel."; | ||
| } | ||
|
|
||
| ScopedEnvironmentVariables scoped_env_vars{{{"ORT_ENABLE_XQA", "0"}}}; |
| /// Optional positive upper bound for the total_sequence_length scalar of non-windowed CUDA | ||
| /// GroupQueryAttention nodes. The scalar value is unavailable during workspace declaration, | ||
| /// so callers must provide a sound bound when it may exceed the past-cache capacity. |
Full reviewVerdict: Request changes. The partial-alias and workspace-sizing implementation is generally sound, but there are three confirmed Major issues. The first is a functional integration gap rather than merely a test problem. 1. Major: the new session bound does not reach Level 1, so constrained-memory planning still cannot use it
As a result, every non-windowed GQA node still uses the generic workspace fallback during partitioning. After kernel construction, Level 2 can declare the larger bounded root.
The documentation statement that Level 1 “cannot access the session option” is also inaccurate: the framework already provides a session-config channel; this key has simply not been plumbed through it. Suggested fix: add this key to
2. Major:
|


Description
Follow-up to #32617 for non-windowed CUDA
GroupQueryAttentionworkspace estimation.ep.cuda.gqa_workspace_max_total_sequence_lengthas an explicit positive bound for the runtimetotal_sequence_lengthscalar used by Level-2 declaration;MayInplaceconservatively by including the valid one-sided past/present alias case and its full past-tensor preservation buffer; andThe Level-1 node adapter remains unavailable for non-windowed inputs because it cannot access session configuration. This PR changes estimation and declaration only; it does not consume a planned workspace root or change runtime allocation topology.
Validation
git diff --checkpassed.group_query_attention_workspace_estimate_test.cc.onnxruntime_providers_cuda_utCMake dependency cycle/module-loading setup when internal tests are enabled.Tracks #29775.