fix(gc): fold the post-fp-setup stack allocation into the walker's frame base (#7328) - #7329
Merged
Merged
Conversation
|
Caution Review failedThe pull request is closed. ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughThe AArch64 stack-map walker now includes contiguous ChangesAArch64 stack-map offset fix
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related issues
Possibly related PRs
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
proggeramlug
pushed a commit
that referenced
this pull request
Aug 3, 2026
…echanism
Two corrections and one measurement.
The gap suite re-run against RS4GC in-process, two arms per test
(shadow-stack control + RS4GC), 479/479: 447 pass->pass, 19 pre-existing
diffs unchanged, 13 node_fail, ZERO new regressions, ZERO refusals, ZERO
compile failures. Zero refusals is the load-bearing number -- 128 of the
479 tests contain `try {}` and the bridge cannot compile any of them.
The earlier soak's "13 regressions, do not flip" was measured against
the bridge, before #7329/#7330, on a backend that structurally cannot
compile a quarter of the suite. It should not be carried forward.
And the x86-64 mechanism was wrong. The workflow comment claimed
_Unwind_GetGR(ctx, 7) "does not reliably return the stack pointer".
Measured on x86-64 Linux (glibc 2.39, gcc 13.3.0): it SEGFAULTS. RBX,
RBP and RIP return correctly; RAX and RSP both SIGSEGV, because libgcc
tracks only the columns CFI restores and RSP is derived from the CFA
rather than tracked. The fault is in the call itself, so no address
validation after it can help -- the previous wording pointed at the
wrong fix. Details and a reproducer in #7333.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #7328.
Root cause
The fast x29-chain stack-map walker derived a frame's base from
add x29, sp, #immalone, on the premise — stated in the function's own doc comment — that establishing the frame pointer is a function's last stack adjustment.It is not. LLVM emits a further allocation after the
addwhen a function has a large or separately laid-out local area:So every slot in such a frame was read 368 bytes high.
Reproduced, then fixed
PERRY_STACKMAP_WALKER=verifyonnotry_controlwith collections forced:The first slot agrees — that frame has no trailing
sub— and the other five are each exactly +368, matchingsub sp, sp, #0x170found in the disassembly.After the fix, the same command is clean, with movement asserted:
eligible=true,copied_objects=28per cycle. That distinction matters here — my first attempt reported 5/5 clean whilePERRY_GC_DIAGshowedeligible=false fallback=conservative_stack, i.e. no copying minor ran at all.PERRY_CONSERVATIVE_STACK_SCAN=offwas needed to make copying eligible (#7255).Why this is worse than a crash
It is a silent wrong answer.
verifyis not the default, so an ordinary run has nothing to disagree with the fast walker — it enumerates wrong addresses, the collector treats them as the root set, and live objects are missed. ForcingPERRY_STACKMAP_WALKER=unwindon the same probe raised objects copied from 23 to 110.It affects both statepoint backends, since they share the walker.
The fix
Accumulate the contiguous run of
sub sp, sp, #immimmediately following theadd. Asub spseparated from that run is a body operation — a dynamic alloca, a call-argument area — whose effect the stack map's own slot offsets already carry, and is deliberately not folded in.Verification
Four unit tests: the trailing-
subcase, the unchanged common shape, asubafter the prologue run (must not be counted), and a leaf with no fp setup (must still fail closed so the caller falls back to the unwinder).Sabotage-checked — neutralising the accumulation fails
a_sub_after_the_fp_setup_is_includedwith "the trailingsub sp, sp, #0x170must be added to the fp offset".cargo fmtclean; file-size, GC store-site and addr-class gates green. (ci_public_baseline_checkis expected-red — the artifact is deliberately stale during the current optimization work.)Not verified
aarch64 only — the decoder is
#[cfg(target_arch = "aarch64")]and x86-64 returnsNone(falls back to the unwinder), so there is no x86-64 path to regress. Whether this clears any of the 13 gap regressions the statepoint soak found is not measured here; #7327 (the bridge emitting no statepoint oninvoke) is the more likely cause of the exception-shaped ones.Summary by CodeRabbit
Bug Fixes
Tests