Skip to content

Session Replay: race between PixelCopy.request and executor compositing/masking on shared screenshot bitmap #5340

Description

@romtsn

Summary

The PixelCopy-based replay strategy (PixelCopyStrategy) shares a single screenshot Bitmap between two writers:

  1. The system PixelCopy worker on the RenderThread, which writes to screenshot when the next frame's PixelCopy.request(window, screenshot, …) performs its blit.
  2. The replay executor (single-threaded), which reads/writes screenshot inside applyMaskingAndNotify (mask render) and compositeSurfaceViewsAndMask (SurfaceView composite).

If a frame's executor work is still in progress when the next recorder cycle fires PixelCopy.request, the system PixelCopy worker overwrites parts of screenshot while the executor is still touching it. The executor being single-threaded prevents executor↔executor overlap, but does not gate the system PixelCopy worker.

The same shape exists in the no-SurfaceView mask path; the SurfaceView capture feature (#5333) just makes it easier to hit because compositing adds a draw pass to the executor side.

What this looks like in practice

Visual:

  • Torn frames — emitted screenshot ends up with a strip of old composited content (mask + SurfaceView already applied) and a strip of newly captured raw window content. Visually, a horizontal seam in the replay frame.
  • Mask shift / partial mask — if the next PixelCopy.request overwrites window pixels while MaskRenderer.renderMasks is still rendering, masks can land on stale pixels or in the wrong region for one frame.
  • SurfaceView ghosting — composited SurfaceView pixels from frame N can persist into frame N+1 in regions the new PixelCopy didn't clear before the executor's recycle ran.

Native crashes (likely related to #4696):

  • SIGABRT in libhwui.so with the typical __android_log_assert → art::Runtime::Abort chain. PixelCopy runs on the RenderThread on top of libhwui; when its hardware blit lands on a Bitmap whose backing hardware buffer is being concurrently drawn into by a Canvas on the executor, libhwui's LOG_ALWAYS_FATAL/CHECK invariants fail and the process aborts.
  • getFrame() called on a context with no surface! — variant of the above where the Window's surface is torn down (e.g. activity backgrounded, surface recreated) between PixelCopy.request and the actual blit. Reported in #4696.
  • Recycle race on close() — close() recycles screenshot on the executor; we guard with isClosed in the PixelCopy callback, but the guard is checked when the callback runs, not at the RenderThread blit moment. If recycle happens between PixelCopy.request and the blit, libhwui sees a recycled bitmap and aborts.

When it manifests

  • Default session replay (1 fps): rare. Mask render typically takes ~10s of ms; the next cycle is 1s away. Native crashes correlated with this race appear concentrated on Android backgrounding flows where the surface tears down (see #4696).
  • Buffer mode at higher frame rates / heavy view trees / slow devices: more frequent. The smaller the inter-frame gap, the easier it is for executor work to still be in flight when the next PixelCopy.request lands.
  • isCaptureSurfaceViews = true (experimental, opt-in via #5333) widens the executor work window because compositing adds another draw pass before masking.
  • ExoPlayer / SurfaceView-heavy apps backgrounding — the classic #4696 repro.

Proposed fix

A frameInFlight: AtomicBoolean set on the main thread immediately before PixelCopy.request(window, screenshot, …) fires, and cleared at the end of applyMaskingAndNotify on the executor. The next recorder cycle skips its capture if the flag is set (drops the frame, matches the codebase's existing drop a frame rather than fight pattern). Single-threaded executor means we only need to coordinate the one main → executor → main handoff.

Additionally, close() should fence against in-flight PixelCopy by either:

  • Posting the screenshot.recycle() after the executor confirms no work is queued and frameInFlight is clear, or
  • Capturing a strong reference to the bitmap in the PixelCopy callback closure and only recycling once we've observed the callback fire.

Alternatives considered:

  • Synchronize on screenshot — blocks the main thread waiting for the executor.
  • Double-buffer (copy bitmap before handing to executor) — doubles the per-frame allocation footprint.

Out of scope here

This is independent of the windowLocation / svLocation field-vs-local races already addressed in #5333 — those snapshots are already taken into locals before the executor handoff.

Related

  • #4696 — SIGABRT in libhwui.so with PixelCopy strategy; likely a manifestation of this race (especially the surface-teardown variant).

Activity

  1. linear-code commented on Apr 28, 2026

    @linear-code
  2. romtsn commented on Apr 28, 2026

    @romtsn
    ContributorAuthor

    Deeper investigation — additional crash mechanisms beyond the executor↔RenderThread race

    Beyond the executor/PixelCopy race covered in the issue body, here are the additional mechanisms specific to #4696's pattern, and why the CanvasStrategy workaround stops them.

    1. Bitmap recycle vs. in-flight PixelCopy

    The path:

    • Main thread: PixelCopyStrategy.capture() → PixelCopy.request(window, screenshot, callback, mainHandler) returns immediately; the actual blit is queued on the RenderThread.
    • Activity goes to background → ReplayIntegration.pause() / stop() propagates → eventually PixelCopyStrategy.close() submits screenshot.recycle() to the (single-threaded) executor.
    • RenderThread tries to write into screenshot's hardware buffer at the same time the executor is freeing it → libhwui LOG_ALWAYS_FATAL → SIGABRT.

    The current isClosed check inside the PixelCopy callback doesn't help here — by the time the callback fires, the hardware blit has already happened (or already SIGABRTed). Any fix needs to fence at queue time, not at callback time.

    2. Window surface torn down during PixelCopy (the getFrame() called on a context with no surface! variant)

    PixelCopy.request(window, ...) reads from the Window's BLAST / buffer-queue–backed surface. When the activity is being stopped, that surface gets disconnected by SurfaceFlinger / ViewRootImpl. If the request was already queued on the RenderThread, the system PixelCopy worker calls into a context whose surface has already been disconnected → libhwui asserts. On some Android versions the assertion text is exactly getFrame() called on a context with no surface!.

    Today we gate with !root.isShown only in ScreenshotRecorder.onDraw(), not at PixelCopy submission time. The recorder loop runs on its own cadence via the main looper handler — there is no check that the window is still attached immediately before PixelCopy.request.

    3. HardwareBuffer lock contention on the destination bitmap

    Even with Bitmap.Config.ARGB_8888 (software-backed), the system's PixelCopy implementation may stage through a GraphicBuffer and lock the destination's pixels for the blit. If the executor is concurrently doing a Canvas op on the same Bitmap (mask render, the new SurfaceView composite), libhwui can hit a CHECK on buffer ownership.

    The single-threaded executor doesn't help here — the contention is between executor and RenderThread, not executor↔executor.

    4. ExoPlayer / SurfaceView correlation in #4696

    ExoPlayer using SurfaceView with hardware decoding doesn't directly cause the crash, but it correlates because:

    • SurfaceView's surface is connected to MediaCodec's hardware decoder, which holds GraphicBuffers across the BufferQueue.
    • On backgrounding, the SurfaceView's surface goes through a destroy/recreate dance.
    • During that dance the Window's overall buffer state is in a transitional state — exactly when our PixelCopy.request is most likely to hit a torn-down surface.

    So ExoPlayer widens the timing window but plain backgrounding without ExoPlayer can still hit it (matches alesrazym's "few reports without ExoPlayer").

    5. LOW_MEMORY / onTrimMemory cleanups

    The original #4696 reporter noted breadcrumbs with LOW_MEMORY. Some devices' graphics drivers drop GraphicBuffers on TRIM_UI_HIDDEN / TRIM_MEMORY_RUNNING_*. If we have a PixelCopy in flight when this happens, the destination bitmap or source surface can lose its hardware backing → libhwui assertion.

    Why CanvasStrategy doesn't crash

    CanvasStrategy calls View.draw(canvas) on a software canvas backed by a Bitmap allocated as ARGB_8888. This:

    • Doesn't go through PixelCopy / RenderThread async pipeline.
    • Doesn't read from the Window's hardware surface.
    • Runs synchronously on the thread that calls it.

    So all four classes of races above (recycle vs. RenderThread, surface-torn-down, hardware-buffer-lock, OOM trim of GraphicBuffers) don't apply to CanvasStrategy. That's why switching to it in alesrazym's app fixed the crashes — at the cost of fidelity (see #5317 for the M2/M3 nested-theme content gap).

    Concrete additional mitigations beyond the frameInFlight gate

    1. Pre-submission attached/visible check: right before PixelCopy.request(window, ...), verify root.isAttachedToWindow && root.windowVisibility == VISIBLE && root.isShown && root.windowToken != null. If any fail, skip the request. Same idea for surfaceView.holder.surface.isValid (we already do this) plus surfaceView.isAttachedToWindow.
    2. Lifecycle-aware pause (via ProcessLifecycleOwner or Application.ActivityLifecycleCallbacks.onActivityPaused API ≥ 29 / onActivityStopped API < 29): set a recordingPaused flag and gate PixelCopy.request on it; resume on onActivityResumed. This closes most of the surface-torn-down window.
    3. Defer recycle: in PixelCopyStrategy.close(), route screenshot.recycle() so it runs only after the in-flight PixelCopy callback has fired (or the strategy was never used). Avoids freeing the bitmap from under the RenderThread.
    4. onTrimMemory handler: implement ComponentCallbacks2.onTrimMemory and pause recording at TRIM_MEMORY_UI_HIDDEN and above.

    Combined, items 1–3 should close the most reproducible path that #4696 is hitting.

  3. romtsn commented on Apr 28, 2026

    @romtsn
    ContributorAuthor

    Correction / clarification on the lifecycle mitigation

    The "lifecycle-aware pause" mitigation in my previous comment was misleading — we already pause replay on backgrounding via LifecycleWatcher.onBackground() (ProcessLifecycleOwner ON_STOP), which calls ReplayController.pause() → ReplayIntegration.pauseInternal() → WindowRecorder.pause() → ScreenshotRecorder.pause() (sets isCapturing=false, unbinds the root). Future recorder.capture() calls then bail early on the isCapturing flag.

    So the actual gaps are narrower than I described, and the mitigation needs to be reframed:

    Gap 1 — 700ms ProcessLifecycleOwner timer

    ProcessLifecycleOwner.ON_STOP fires ~700ms after the last activity's onStop (intentional, to keep activity-to-activity transitions from churning state). During that 700ms window the foreground activity's window/surface is already being torn down, but our recorder is still running its 1 fps loop and firing PixelCopy.request(window, …). That's exactly when the surface teardown can race the request.

    Gap 2 — pause() only stops future captures, not in-flight ones

    After pause():

    • Any PixelCopy.request queued before pause is still in flight on the RenderThread. Its callback fires later and currently only checks isClosed, not isCapturing / a paused flag.
    • For the new SurfaceView path, the callback can issue additional PixelCopy.request(surfaceView, …) calls after pause() returned. Those are the ones most likely to hit a torn-down surface or get their destination bitmap recycled from under them.

    Refined mitigations

    Replacing item 2 in my previous comment with:

    1. Tighten the entry trigger. Pause replay on Application.ActivityLifecycleCallbacks.onActivityStopped directly (not via ProcessLifecycleOwner) — eliminates the 700ms gap. Re-resume on onActivityResumed. Needs a small re-entry guard so quick activity-to-activity transitions don't churn pause/resume.
    2. Make in-flight callbacks check the paused state too. In the PixelCopy callback (both the window callback and the per-SurfaceView callbacks), early-return if we're paused since the request was issued. Don't issue follow-up SurfaceView captures from a paused state.
    3. Defer screenshot.recycle() past the in-flight blit (unchanged from prior comment). Even with Mavenized, HttpClient removed and proxy support added #1 + Sentry Versions Supported #2, a request already accepted by the RenderThread can still SIGABRT if we recycle the destination during its blit; the fix has to fence the recycle, not just the next capture.

    Items 1 & 2 close most of the surface-teardown window. Item 3 is needed regardless because the RenderThread blit window is opaque to us — once PixelCopy.request returns, we have no signal for "blit started" vs "blit will start later," only "blit finished" (the callback).

  4. BAGADIR commented on Apr 28, 2026

    @BAGADIR
  5. moved this to Waiting for: Product Owner in GitHub Issues with 👀 3on Apr 28, 2026
  6. romtsn commented on May 6, 2026

    @romtsn
    ContributorAuthor

    follow up and also cache bitmaps used for surfaceview capturing after this lands: #5333

    Image
  7. added a commit that references this issue on Jul 24, 2026
    7414e9b
  8. github-actions commented on Jul 29, 2026

    @github-actions
    Contributor

    A PR closing this issue has just been released 🚀

    This issue was referenced by PR #5808, which was included in the 8.51.0 release.

  9. added 7 commits that reference this issue on Jul 30, 2026
    e6726e0
    1a97377
    587630a
    4f2c8a1
    d63d92f
    2684800
    5cb2c36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions