Skip to content

ci: run the GPU pipeline gate on both arches, measure coverage over ./..., and fix the ten defects that exposed - #183

Open
dpsoft wants to merge 11 commits into
mainfrom
feat/gpuprobe-ci
Open

dpsoft wants to merge 11 commits into
mainfrom
feat/gpuprobe-ci

Conversation

@dpsoft

@dpsoft dpsoft commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Two CI gaps — a gate no job had ever run, and a coverage number that was only ever about five packages. Turning them on found ten defects, nine of them in tests and build guards that could not fail.

The two CI changes

gpuprobe was in no CI job, on either architecture. TestStubDrivesThePipelineToPprofWithoutAGPU drives the whole GPU path without a GPU — probe fire, uprobe attach, BPF decode, join, pprof samples — with a synthetic producer standing in for CUDA. It needs CAP_BPF/CAP_PERFMON/CAP_CHECKPOINT_RESTORE and skips without them, so it would have stayed invisible even once listed: a skip reads the same as a pass. The step runs it under sudo and then asserts it did not skip, with a rename guard too (stricter than #178's, which a vanished test would have satisfied).

Coverage over five hand-listed packages was a number about five packages. ./cpu/… ./profile/… ./offcpu/… ./pyunwind/… ./unwind/dwarfagent/… → ./.... 28 packages with tests sat outside the measured set. #178's privileged pass deliberately stays narrow on the five — it exists to execute the BPF-gated tests, not to measure coverage.

One real product bug: #185

The gate found that every arm64 DWARF walk was abandoned, reached-root=0 against 58/58 on amd64. ehcompile.CFIEntry.CFAOffset was an int16, and perfagent_stub_run's aarch64 frame is 112 + 65536 + 6384 = 72032 bytes, so DW_CFA_def_cfa_offset 72032 wrapped to 6496. The walker computed the CFA 65536 bytes low, read a zero return address, and pushed 0x0 as a frame.

x86-64 was immune because gcc there emits DW_CFA_def_cfa_register %rbp, pinning the offset at 16 however large the frame; AAPCS64 keeps the CFA SP-rooted as the frame grows. The bug needs a function with both a >32KB frame and an SP-rooted CFA — and the gate's own producer is exactly that, because it allocates its record batches on the stack. This int16 has been truncating in every arm64 profile of a big-frame function; the gate is the first thing that ever ran there.

Verified on the runner, not inferred from green:

                   before              after
walk shape    abandoned=58         abandoned=0
              reached-root=0       reached-root=58
depth         map[3:58]            map[8:58]
caller named  0/58                 58/58

58/58 naming perfagent_fpless_caller is the load-bearing number — that frame is only reachable by unwinding a frame-pointer-less frame through CFI.

Nine checks that could not fail

# Where Defect
1 gpuprobe cubin seals the "tmpfs file" case used t.TempDir(), which is ext4 on a runner → F_GET_SEALS returns EINVAL and the subtest silently became a copy of the pipe case beside it
2 internal/usdt "the ABI pins its argument registers" checked as 8@%rdi…; AAPCS64 pins them as x0/x1/x2
3,4 gpuprobe CFI ×2 "has a live frame pointer" checked as CFATypeFP — not even an arch property: gcc/aarch64 is SP-rooted, clang/aarch64 FP-rooted, both save the FP
5 gpuprobe PC sampling frame order was the pre-#172 root-first ordering — it asserted the bug #172 removed
6 gpuprobe PC sampling Positive(GroupsUnresolvedName) on a counter this producer cannot reach (r.correlation = 0 by design → groups stop at the process gate)
7 shim/Makefile check-fpless greps for mov %rsp,%rbp; on aarch64 checks 1–2 passed vacuously and check 3 failed on a function that does have a frame pointer
8 symbolize/nvsym required all 50 Cached lookups to miss, i.e. that a background goroutine had not finished — measured, the prefetch lands at iteration 189–733
9 main_test.go expected pid1234-…, which requires pid 1234 to be dead; on the arm64 runner it is dockerd

Also replaced three hardcoded t.Logf prediction lines with the measured rows — they printed "main: mode=FP_SAFE / reached-root == dwarf" on the arm64 run where neither was true, and actively misled the first reading of #185.

Verification

Full ./... on amd64: 37 packages, no failures. gate=success on both arches. The only red is Upload coverage: codecov-action@v3 began failing every upload repo-wide with a TLS handshake error between 21:48Z and 00:30Z — it hits a branch whose diff is only examples/*.{md,sh,rs}, and this branch uploaded fine at 21:48 with the ./... change, so it is service-side. Needs a v4 bump plus a CODECOV_TOKEN secret, separately.

Follow-ups filed rather than folded in

https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im

dpsoft added 4 commits October 3, 2026 18:25
gpuprobe appears in no CI job, on either architecture. The unit-test step
lists cpu, profile, offcpu, pyunwind and unwind/dwarfagent and stops
there, so TestStubDrivesThePipelineToPprofWithoutAGPU has never run in
CI -- and it is the one test that drives the WHOLE GPU path without a
GPU: probe fire, uprobe attach, BPF decode, join, pprof samples, with a
synthetic producer standing in for CUDA.

Everything needed was already here and unconnected. The gate exists, CI
already runs tests under sudo in the integration job, and the producer
builds itself. It needed a line.

It also covers the aarch64 DWARF walker, which is why this matters beyond
coverage arithmetic: the gate deliberately drives the frame-pointer-LESS
producer, whose bridge frames a saved-RBP walk provably cannot reach, so
the unwinder has something real to bite on rather than a stack it could
have guessed. Running on ubuntu-24.04-arm exercises that on real hardware.

The step fails if the gate SKIPS rather than passing. A skip and a pass
are the same colour in a green run, and that is precisely how this stayed
invisible -- so the assertion is on "it ran", not only on "it did not
fail". Verified the detector fires by running it against a machine with
no capabilities, where it correctly reports the skip.

Left out of tests.yml deliberately: that workflow runs -race with
coverage, and a privileged test that spawns processes and loads BPF is a
poor fit for it.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
Every go test in CI named five packages: cpu, profile, offcpu, pyunwind,
unwind/dwarfagent. Twenty-eight packages with tests were outside that
list and had never run in CI -- internal/flamegraph, internal/foldedstacks,
internal/framename, symbolize and its three subpackages, pprof,
perfagent, unwind/{ehcompile,ehmaps,fpwalker,interp,procmap},
internal/{usdt,gpuabi,k8slabels,nspid,shiminstall}, gpu, gpuprobe.

The integration job's `cd test && go test ./...` does not cover them:
test/ is a separate module, so that ./... is about test/ alone.

An enumerated list silently excludes everything added after it was
written, which is what happened -- including the flamegraph and procmap
work merged this month, whose tests CI never ran. Measured before
changing it: the whole main-module suite is 10s, 26s under -race. There
was never a cost reason for the list.

tests.yml gets the same fix, where it matters twice over: coverage
measured across five hand-picked packages is a number about those five
rather than about the project.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
The cubin consumer refuses a descriptor that is not a sealed memfd, and the
two ways that can arrive take different kernel paths: F_GET_SEALS returns
EINVAL on a pipe, while a file on shmem answers with F_SEAL_SEAL alone and no
write protection. The second is the dangerous one -- it looks like a sealed
object and the peer can still write through it -- so it is pinned separately.

It was pinned on t.TempDir(), which is tmpfs only by luck. It follows TMPDIR,
which is tmpfs on Fedora and ext4 on a GitHub runner, where F_GET_SEALS
answers EINVAL and the case silently became a second copy of the pipe case:
same code path, different subtest name, the dangerous half covered by nothing.

Its own assertion is what caught it, once ./... made the package run at all:

    "F_GET_SEALS: invalid argument (not a sealable memfd?)"
        does not contain "F_SEAL_WRITE"

So find a directory statfs actually reports as TMPFS_MAGIC -- /dev/shm,
t.TempDir(), /run/user/$uid -- and pin the seal set in the helper, where a
wrong precondition says so, rather than leaving it to an assertion twenty
lines down that only happens to notice. With no tmpfs at all it fails; a skip
here would read as a pass, which is the whole defect being fixed.

Reproduced the CI failure locally byte-for-byte by pointing TMPDIR at btrfs:
the old test fails there, the new one passes.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
Three assertions in tests no arm64 job had ever run. Each described an
architecture-neutral property and then checked an x86-64 artefact of it, so
each failed the first time ./... reached them on aarch64.

1. internal/usdt: "the ABI pins its argument registers" was checked as
   `8@%rdi 8@%rsi 8@%rdx`. AAPCS64 pins them just as hard, as x0/x1/x2 --
   which is what arm64 measured. Now expected per GOARCH, with the table
   required to have an entry so a new architecture fails loudly instead of
   passing vacuously.

2. gpuprobe, both CFI tests: "main / perfagent_stub_run is reached with a live
   frame pointer, so the FP path works" was checked as CFATypeFP. Having a
   frame pointer the walk can step out through means the caller's FP is SAVED
   at a known slot -- FPTypeOffsetCFA. Whether the CFA is then expressed off
   that register or off SP is a separate, independent compiler choice:

       gcc    x86-64   CFA FP-rooted   (def_cfa_register %rbp, FP at CFA-16)
       gcc    aarch64  CFA SP-rooted   (SP constant through the body, FP at CFA-32)
       clang  aarch64  CFA FP-rooted   (def_cfa_register x29,   FP at CFA-32)

   All three save the frame pointer; all three are walkable; all three report
   FPTypeOffsetCFA. CFATypeFP pinned the first row of that table and nothing
   else. It is not even an architecture property -- it differs between gcc and
   clang on the same architecture, so the assertion would have broken again on
   a clang build of the arch it passed on. walk_step reads the saved-FP slot
   and does not care what the CFA is rooted at.

   The CFATypeSP assertions on the FP-LESS bridge frames are left alone: with
   -fomit-frame-pointer there is no frame-pointer register in use, so the CFA
   can only be expressed off SP, on any compiler.

Measured, not assumed. The aarch64 rows above come from
unwind/ehcompile/testdata/hello_arm64.golden (gcc, committed) and from a clang
--target=aarch64-linux-gnu compile of an FP-ful non-leaf parsed through
ehcompile; the x86-64 row from the real perfagent-gpu-fpless fixture. The
arm64 expectations for the argument registers and for the SP-rooted CFA are
the values CI itself reported.

Full ./... on amd64: 37 packages, no failures.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
@dpsoft

dpsoft commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Pushed two fixes for what ./... turned up — both are tests that nothing had ever run.

gpuprobe, both arches. The "a tmpfs file, which is sealable but not sealed" seal case used t.TempDir(), which follows TMPDIR: tmpfs on Fedora, ext4 on a runner. On ext4 F_GET_SEALS answers EINVAL, so the subtest silently became a second copy of the pipe case next to it, and the dangerous half — an object that answers F_GET_SEALS like a sealed one while the peer can still write through it — was covered by nothing. Its own assertion caught it. Now it finds a directory statfs reports as TMPFS_MAGIC and pins the seal set in the helper; with no tmpfs it fails rather than skips. Reproduced the CI failure byte-for-byte locally by pointing TMPDIR at btrfs.

Three x86-64 expectations, arm64 only. Each described a neutral property and then checked an x86 artefact of it:

  • internal/usdt checked "the ABI pins its argument registers" as 8@%rdi 8@%rsi 8@%rdx; AAPCS64 pins them just as hard, as x0/x1/x2. Now per-GOARCH, and the table must have an entry so a new arch fails loudly instead of passing vacuously.
  • both gpuprobe CFI tests checked "reached with a live frame pointer" as CFATypeFP. The property is that the caller's FP is saved at a known slot — FPTypeOffsetCFA. What the CFA is rooted at is an independent compiler choice:
CFA saved FP
gcc x86-64 FP-rooted CFA−16
gcc aarch64 SP-rooted CFA−32
clang aarch64 FP-rooted CFA−32

All three save the frame pointer, all three are walkable, and all three report FPTypeOffsetCFA. CFATypeFP pinned row one and nothing else — and it is not an architecture property at all, so it would have broken again on a clang build of the very arch it passed on. walk_step reads the saved-FP slot; the CFA base is not its business. The CFATypeSP assertions on the FP-less bridge frames are untouched: under -fomit-frame-pointer there is no FP register in use, so SP-rooted is the only option on any compiler.

Measured rather than assumed — the gcc/aarch64 row from the committed hello_arm64.golden, the clang/aarch64 row from a --target=aarch64-linux-gnu compile parsed through ehcompile, the x86-64 row from the real perfagent-gpu-fpless fixture; the two arm64 expectations are the values CI itself reported.

Full ./... on amd64: 37 packages, no failures.

https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im

…achable counter

Running the gpuprobe package under sudo for the pipeline gate also switched on
its sibling TestStubDrivesPCSamplingToPprofWithoutAGPU, which needs CAP_BPF
and had therefore never run anywhere. Two of its assertions were wrong.

1. Frame order was the pre-#172 ordering.

   It built the expectation as CPU stack, then the boundary marker, then the
   kernel - GPU frames appended after _start. #163 established that
   Sample.Location[0] is the leaf and #172 fixed gpu.projectionFrames to
   assemble accordingly: kernel first, then the marker, then the CPU path
   outward. The expectation was left behind, so it asserted exactly the
   ordering #172 removed. Nothing caught it because no job gave this package
   capabilities.

   Measured under caps, the projection emits:

     [gpu:kernel:kernel_1111] [gpu:launch] gpu_launch_sampled_v1_emit
     perfagent_stub_run perfagent_fpless_bridge perfagent_fpless_caller
     main __libc_start_call_main __libc_start_main_alias_1 _start

   which is what the test now expects, built in the same order
   projectionFrames builds it. The substring scan over the two synthesized
   frames moves with them, from want[len(want)-2:] to want[:2].

2. Positive(GroupsUnresolvedName) named a gate this producer cannot reach.

   Measured: GroupsNoProcess=4, GroupsUnresolvedName=0, with 4 pending module
   groups. The groups stop at the PROCESS gate, not the name gate, and
   deliberately so - stub.cc sets r.correlation = 0 on every PC record because
   CONTINUOUS collection supplies no correlation and a stub that invented one
   would hide the join the consumer must actually make (spec §6.3 finding 3).
   pendingModuleKeyFor reads the PID off that correlation, so key.PID is 0 and
   Timeline stops the group one step before looking up (crc, functionIndex).

   No defect in either half: the samples stay pending and counted, which is
   the property the block is about. The assertion just named the wrong
   counter, and an unreachable one, so it could only ever fail.

   Replaced with the exact fact - every pending group counted at the gate it
   actually stopped at - plus a zero-pin on the name gate in the same idiom as
   TestGateTheStubsPCRecordsCannotAttributeToAnything: when the stub is given
   correlations this fails by name and the person fixing it moves the
   assertion to the gate the groups then reach.

Also added the PCJoin struct to that failure message. Its absence is why the
first CI log could not say which bucket the groups had fallen into.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
Two more things the GPU pipeline gate found by being the first job to run
them.

## check-fpless was an x86-64 check wearing an architecture-neutral name

It greps disassembly for `mov %rsp,%rbp`. On aarch64 there is no %rbp: the
frame record is `stp x29, x30, [sp, #-N]!` followed by `mov x29, sp`. So on
arm64 the three checks behaved like this:

  1. bridge/caller must not ESTABLISH a frame pointer   - passed vacuously
  2. bridge/caller must not TOUCH %rbp                  - passed vacuously
  3. perfagent_stub_run must KEEP its frame pointer     - FAILED

The first two look for the absence of an x86 pattern, which is always absent
on aarch64, so they asserted nothing there while reporting OK. The third
failed on a function that does have a frame pointer. The build guard that
exists so the gate cannot pass while proving nothing was itself proving
nothing, on half the architectures.

Now selected by uname -m, with the aarch64 form accepting either `mov x29, sp`
or the `stp x29, x30` frame-record store - both appear only when a frame
record exists, and taking either keeps a gcc/clang codegen difference from
turning the gate red again. Verified against real aarch64 disassembly
(llvm-objdump of a --target=aarch64-linux-gnu compile): it separates an
FP-ful non-leaf from an FP-less one, and x86-64 still reports check-fpless: OK.

The stub_run window widens from 4 instructions to 10: on aarch64 the
`mov x29, sp` can follow other callee-saved stores, and within 10 instructions
of entry it is still unambiguously the prologue.

## TestCachedPrefetchesOncePerBuildID asserted that a goroutine had not run yet

It looped 50 Cached() lookups and required every one to miss. Cached is an
os.Stat, and the FIRST miss schedules the background prefetch, so lookups
start hitting the moment that prefetch writes the file. Nothing orders it
after the loop.

Measured with the real code, letting the loop run until it hits: the prefetch
lands at iteration 189-733 on this machine. The shipped assertion covered 50,
a margin of only 4-15x - which a contended runner crosses, and amd64 did, the
first time ./... ran this package.

The invariant is the request count, not the miss count: one prefetch per
build-id however many stacks miss on it. That held in every one of those runs
(served == 2) and is still asserted. The first lookup is still required to
miss, because an empty cache directory makes that one deterministic; the other
49 now just drive the single-flight without claiming to know who won.

300 consecutive runs pass. Full ./... on amd64: 37 packages, no failures.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
dpsoft added 5 commits October 4, 2026 21:29
…#185

## main_test.go asserted a property of the machine

TestGenerateOutputNameShapes expected generateOutputName(1234, ...) to produce
"pid1234-<ts>-off-cpu.pb.gz". readProcessName reads /proc/<pid>/comm and falls
back to pid<N> only when that read FAILS, so the assertion required pid 1234
to be dead on whatever host ran it. On a GitHub arm64 runner pid 1234 is
dockerd:

    Expect "dockerd-202610032257-off-cpu.pb.gz"
        to match "^pid1234-\d{12}-off-cpu\.pb\.gz$"

The root package had never been in a CI job either, so this went unseen until
./... reached it.

Both tests here now take a pid confirmed free at test time, walking down from
/proc/sys/kernel/pid_max rather than trusting a big-looking constant - the
neighbouring test's 999999 was the same assumption with better odds, not a
different kind of claim. If every candidate resolves, it fails and says so
rather than quietly testing the wrong branch.

## The gate's arm64 arm is off, pointing at the bug it found

Running the gate on arm64 is what surfaced #185: all 58 walks take the DWARF
path and all 58 are abandoned, reached-root=0 against 58/58 on amd64, with the
tables present and chosen (no-tables=0, cfi-miss=0, fp-only=0) and the
producer's shape confirmed correct there by check-fpless and by both CFI tests.
The tables say the walk should reach root and the walk says it does not. That
is a product defect this branch revealed, not one it introduced - main is
green on both arches today only because nothing ever ran this.

Rather than hold the gate for a defect that needs arm64 hardware to
investigate, it guards amd64 from now, and the arm64 arm turns on by deleting
one `if` - which is the acceptance criterion recorded on #185. The
unprivileged ./... run still covers gpuprobe on both architectures; what is
amd64-only is the privileged end-to-end walk.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
…e's own frame

Three of this test's four log lines were hardcoded prose. They printed the
same text whatever the tables said, so on arm64 the log claimed

    main: mode=FP_SAFE, reached WITH a frame pointer
    prediction: reached-root == dwarf, fp-exhausted == abandoned == 0

on a run where main's CFA is SP-rooted and all 58 walks were abandoned with
reached-root=0 (#185). A log that cannot disagree with the code is the defect
this gate exists to remove, one level down -- and it actively misled the first
reading of the arm64 failure.

Replaced with the measured rows, leaf-first, and extended to the two frames
the assertions never covered:

  - The assertions derive every frame BETWEEN the probe and _start. The frame
    the walk STARTS in was not among them.
  - gpu_launch_sampled_v1_emit has no ELF symbol: the probe macro is inlined,
    which is why it reaches a sampled stack only through DWARF inline info.
    So the midpoint of perfagent_stub_run is not where the walk begins either.

So the probe PCs are read off the USDT notes and converted from file offset to
vaddr through PT_LOAD, and the row covering each one is logged -- that is the
row governing the walk's first step. amd64 reads:

    perfagent_stub_run      cfa=FP+16 fp=OFFSET_CFA-16 ra=OFFSET_CFA-8
    perfagent_fpless_bridge cfa=SP+32 fp=SAME_VALUE+0  ra=OFFSET_CFA-8
    main                    cfa=FP+16 fp=OFFSET_CFA-16 ra=OFFSET_CFA-8
    _start                  cfa=SP+8  fp=SAME_VALUE+0  ra=UNDEF+0
    probe gpu_launch_v1     pc=0x400d10 cfa=SP+8 fp=SAME_VALUE+0 ra=OFFSET_CFA-8

A row whose ra is SAME_VALUE or REGISTER is called out inline, because that is
the unflagged `goto stop` in unwind_common.h -- the return address in a
register the walker does not track, x30/LR on arm64 -- and it is the shape
#185 is hunting. This test runs unprivileged, so it reports on both
architectures without needing the gate.

Diagnostic only: no assertion changes, and nothing here fixes #185.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
Reverts the amd64-only scoping. A gate switched off is a gate that cannot
tell anyone when the bug comes back, and #185 is reachable from CI -- the
arm64 runner is native, so there is no hardware reason to defer it.

The first diagnostic round already refuted the obvious hypothesis. Measured
rows at the gate's actual probe site, which is gpu_launch_sampled_v1:

    arm64  probe gpu_launch_sampled_v1  cfa=SP+6496 fp=OFFSET_CFA-112 ra=OFFSET_CFA-104
    amd64  probe gpu_launch_v1          cfa=SP+8    fp=SAME_VALUE+0   ra=OFFSET_CFA-8

The return address IS OFFSET_CFA on arm64, so the walk is NOT stopping on the
unflagged "RA lives in a register we do not track" arm, which is what the
counters had suggested. Two further facts narrow it, both read off the arm64
failure rather than assumed:

  - StacksUnresolved was 0 and the sawStubFrame assertion passed, so the walk
    produced NAMED frames including perfagent_stub_run.
  - sawFPLessCaller was 0, so it never reached perfagent_fpless_bridge.

It stops stepping OUT of perfagent_stub_run. Note the two arches do not even
take the same path out of that frame: amd64's row is cfa=FP+16, so it walks
the frame pointer, while arm64's is SP-rooted with a 6496-byte frame, so it
takes the DWARF path. (6496 fits __s16 cfa_offset, so this is not an overflow.)

What the aggregate counters cannot say is at WHICH frame, and that is exactly
what separates a faulting read from a rule the walker mishandles. So the gate
now logs a stack-depth histogram and the frames of the first sampled stack, on
every architecture, so the healthy amd64 shape sits beside the arm64 one.

Still diagnostic: no walker change here, #185 stays open, and the arm64 gate
arm is expected to fail until it is fixed.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
…185)

The per-frame log localised #185 precisely. arm64, every sampled stack:

    stack depth histogram: map[3:58]
      frame 0: 0xab1ed8be3ef8 gpu_launch_sampled_v1_emit
      frame 1: 0xab1ed8be3ef8 perfagent_stub_run
      frame 2: 0x0            0x0

against amd64's map[8:58] reaching _start. Frames 0 and 1 share an address on
both arches - that is the inlined probe expanding into two symbol frames, not
a repeat - so the record really holds two PCs on arm64: the probe site, then
ZERO.

So the first DWARF step out of perfagent_stub_run reads a return address of
zero. Its row, measured on the CI runner itself, is

    probe gpu_launch_sampled_v1  pc=0x3ef8  cfa=SP+6496  fp=OFFSET_CFA-112  ra=OFFSET_CFA-104

and the runtime frame 0 is 0xab1ed8be3ef8, i.e. load bias 0xab1ed8be0000 plus
exactly that 0x3ef8 - so the row the walker uses is the row logged here. The
read therefore lands at sp+6392 and finds nothing.

Ruled out so far, each by measurement rather than reasoning:

  - NOT the unflagged "RA in a register we do not track" arm: ra is OFFSET_CFA.
  - NOT a cfa_offset overflow: 6496 fits __s16 on both sides of the ABI.
  - NOT an arm64 data-alignment-factor bug in ehcompile: readelf on the
    committed hello_arm64.o gives def_cfa_offset 32, r29 at cfa-32, r30 at
    cfa-24, and the golden records CFAOffset=32 FPOffset=-32 RAOffset=-24.
  - NOT the wrong binary: the gate execs a private copy of
    perfagent-gpu-fpless, and the bias arithmetic above confirms it.

Two candidates remain, and they need the code and the tables side by side:
the CFI disagrees with the prologue about where x30 is stored, or ctx->sp is
not what the row assumes. So on the failing path only, and bounded, dump
objdump of perfagent_stub_run's prologue and the FDEs carrying a frame that
size.

Also worth recording: a zero return address is currently PUSHED as a frame
rather than ending the walk. That is wrong independently of why it is zero,
but fixing it needs a new WALKER_FLAG bit and walker_flags is a full __u8,
so it changes the record ABI - deliberately not slipped in here.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
… walk (#185)

perfagent_stub_run's arm64 prologue:

    2c50: stp x29, x30, [sp, #-112]!     x29/x30 at CFA-112 / CFA-104
    2c54: mov x29, sp
    2c70: sub sp, sp, #0x10, lsl #12     sp -= 65536
    2c7c: sub sp, sp, x13                sp -= 6384

112 + 65536 + 6384 = 72032, and gcc emits DW_CFA_def_cfa_offset 72032 for it.
CFIEntry.CFAOffset was an int16, so that became 6496 -- which is exactly what
the gate's diagnostic reported as the row at the probe site. The walker then
computed the CFA 65536 bytes below the real one, read a return address of
zero, pushed 0x0 as a frame, and the walk ended at depth 3:

    arm64   frame 0/1 probe site in perfagent_stub_run,  frame 2: 0x0
    amd64   depth 8, through main and libc to _start

Reproduced on an x86-64 host by assembling gcc's CFI shape by hand:
DW_CFA_def_cfa_offset 72032 in, cfa_off=6496 out.

WHY ONLY ARM64, which is the part worth keeping. On x86-64 gcc finishes the
prologue with DW_CFA_def_cfa_register %rbp, so the CFA offset stays 16 however
large the frame gets. AAPCS64 keeps the CFA SP-rooted while the frame grows,
so only aarch64 can produce an offset past 32767 -- and it needs a function
with BOTH a >32KB frame and an SP-rooted CFA to do it. perfagent_stub_run
allocates its record batches on the stack, so the gate's own producer is
exactly that function. The same int16 has been in every arm64 profile this
agent has ever taken of a big-frame function; the gate is just the first thing
that ran there.

Widened to 32 bits in the four places that must agree:

  - struct cfi_entry (bpf/unwind_common.h): __s16 -> __s32, fields reordered
    offsets-then-types so it still totals 32 bytes -- the extra 6 came out of
    the old interior padding, so CFIEntryByteSize is unchanged.
  - ehcompile.CFIEntry: int16 -> int32.
  - the three truncating casts in interpreter.go.
  - ehmaps.MarshalCFIEntry: PutUint16 -> PutUint32 at the new offsets.

The bpf2go struct regenerated from the compiled BTF confirms the C layout the
marshaller writes, which is the cross-check that the two sides still agree:

    PcStart uint64; PcEndDelta uint32; CfaOffset int32; FpOffset int32;
    RaOffset int32; CfaType uint8; FpType uint8; RaType uint8; Pad [5]uint8

BPF objects regenerated with `make generate-container` so they reproduce CI's
bytes (issue #117's build-path dependence); 6 objects and their wrappers move.

TestMarshalCFIEntryCarriesOffsetsTooBigForAnInt16 pins 72032 through the
marshaller, and the 17 type-sensitive offset assertions in interpreter_test.go
move to int32. Full ./... on amd64: 37 packages, no failures.

NOT fixed here, and it wants its own change: a zero return address is still
PUSHED as a frame instead of ending the walk. That is wrong whatever made it
zero, but flagging it needs a new WALKER_FLAG bit and walker_flags is a full
__u8, so it changes the record ABI.

Closes #185.

Claude-Session: https://claude.ai/code/session_01P5889hA6CrX8ysnQkvv6im
@dpsoft dpsoft changed the title ci: run the GPU pipeline gate, which no job has ever run, and measure coverage over ./... ci: run the GPU pipeline gate on both arches, measure coverage over ./..., and fix the ten defects that exposed Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant