Repository navigation
Session Replay: race between PixelCopy.request and executor compositing/masking on shared screenshot bitmap #5340
Description
Activity
Deeper investigation — additional crash mechanisms beyond the executor↔RenderThread race
Beyond the executor/PixelCopy race covered in the issue body, here are the additional mechanisms specific to #4696's pattern, and why the
CanvasStrategyworkaround stops them.1. Bitmap recycle vs. in-flight PixelCopy
The path:
- Main thread:
PixelCopyStrategy.capture()→PixelCopy.request(window, screenshot, callback, mainHandler)returns immediately; the actual blit is queued on the RenderThread. - Activity goes to background →
ReplayIntegration.pause()/stop()propagates → eventuallyPixelCopyStrategy.close()submitsscreenshot.recycle()to the (single-threaded) executor. - RenderThread tries to write into
screenshot's hardware buffer at the same time the executor is freeing it → libhwuiLOG_ALWAYS_FATAL→ SIGABRT.
The current
isClosedcheck inside the PixelCopy callback doesn't help here — by the time the callback fires, the hardware blit has already happened (or already SIGABRTed). Any fix needs to fence at queue time, not at callback time.2. Window surface torn down during PixelCopy (the
getFrame() called on a context with no surface!variant)PixelCopy.request(window, ...)reads from the Window's BLAST / buffer-queue–backed surface. When the activity is being stopped, that surface gets disconnected by SurfaceFlinger / ViewRootImpl. If the request was already queued on the RenderThread, the system PixelCopy worker calls into a context whose surface has already been disconnected → libhwui asserts. On some Android versions the assertion text is exactlygetFrame() called on a context with no surface!.Today we gate with
!root.isShownonly inScreenshotRecorder.onDraw(), not at PixelCopy submission time. The recorder loop runs on its own cadence via the main looper handler — there is no check that the window is still attached immediately beforePixelCopy.request.3. HardwareBuffer lock contention on the destination bitmap
Even with
Bitmap.Config.ARGB_8888(software-backed), the system's PixelCopy implementation may stage through aGraphicBufferand lock the destination's pixels for the blit. If the executor is concurrently doing a Canvas op on the sameBitmap(mask render, the new SurfaceView composite), libhwui can hit a CHECK on buffer ownership.The single-threaded executor doesn't help here — the contention is between executor and RenderThread, not executor↔executor.
4. ExoPlayer / SurfaceView correlation in #4696
ExoPlayer using
SurfaceViewwith hardware decoding doesn't directly cause the crash, but it correlates because:- SurfaceView's surface is connected to MediaCodec's hardware decoder, which holds GraphicBuffers across the BufferQueue.
- On backgrounding, the SurfaceView's surface goes through a destroy/recreate dance.
- During that dance the Window's overall buffer state is in a transitional state — exactly when our
PixelCopy.requestis most likely to hit a torn-down surface.
So ExoPlayer widens the timing window but plain backgrounding without ExoPlayer can still hit it (matches
alesrazym's "few reports without ExoPlayer").5.
LOW_MEMORY/onTrimMemorycleanupsThe original #4696 reporter noted breadcrumbs with
LOW_MEMORY. Some devices' graphics drivers drop GraphicBuffers onTRIM_UI_HIDDEN/TRIM_MEMORY_RUNNING_*. If we have a PixelCopy in flight when this happens, the destination bitmap or source surface can lose its hardware backing → libhwui assertion.Why
CanvasStrategydoesn't crashCanvasStrategycallsView.draw(canvas)on a software canvas backed by aBitmapallocated asARGB_8888. This:- Doesn't go through PixelCopy / RenderThread async pipeline.
- Doesn't read from the Window's hardware surface.
- Runs synchronously on the thread that calls it.
So all four classes of races above (recycle vs. RenderThread, surface-torn-down, hardware-buffer-lock, OOM trim of GraphicBuffers) don't apply to
CanvasStrategy. That's why switching to it inalesrazym's app fixed the crashes — at the cost of fidelity (see #5317 for the M2/M3 nested-theme content gap).Concrete additional mitigations beyond the
frameInFlightgate- Pre-submission attached/visible check: right before
PixelCopy.request(window, ...), verifyroot.isAttachedToWindow && root.windowVisibility == VISIBLE && root.isShown && root.windowToken != null. If any fail, skip the request. Same idea forsurfaceView.holder.surface.isValid(we already do this) plussurfaceView.isAttachedToWindow. - Lifecycle-aware pause (via
ProcessLifecycleOwnerorApplication.ActivityLifecycleCallbacks.onActivityPausedAPI ≥ 29 /onActivityStoppedAPI < 29): set arecordingPausedflag and gatePixelCopy.requeston it; resume ononActivityResumed. This closes most of the surface-torn-down window. - Defer recycle: in
PixelCopyStrategy.close(), routescreenshot.recycle()so it runs only after the in-flight PixelCopy callback has fired (or the strategy was never used). Avoids freeing the bitmap from under the RenderThread. onTrimMemoryhandler: implementComponentCallbacks2.onTrimMemoryand pause recording atTRIM_MEMORY_UI_HIDDENand above.
Combined, items 1–3 should close the most reproducible path that #4696 is hitting.
- Main thread:
Correction / clarification on the lifecycle mitigation
The "lifecycle-aware pause" mitigation in my previous comment was misleading — we already pause replay on backgrounding via
LifecycleWatcher.onBackground()(ProcessLifecycleOwnerON_STOP), which callsReplayController.pause()→ReplayIntegration.pauseInternal()→WindowRecorder.pause()→ScreenshotRecorder.pause()(setsisCapturing=false, unbinds the root). Futurerecorder.capture()calls then bail early on theisCapturingflag.So the actual gaps are narrower than I described, and the mitigation needs to be reframed:
Gap 1 — 700ms
ProcessLifecycleOwnertimerProcessLifecycleOwner.ON_STOPfires ~700ms after the last activity'sonStop(intentional, to keep activity-to-activity transitions from churning state). During that 700ms window the foreground activity's window/surface is already being torn down, but our recorder is still running its 1 fps loop and firingPixelCopy.request(window, …). That's exactly when the surface teardown can race the request.Gap 2 —
pause()only stops future captures, not in-flight onesAfter
pause():- Any
PixelCopy.requestqueued before pause is still in flight on the RenderThread. Its callback fires later and currently only checksisClosed, notisCapturing/ a paused flag. - For the new SurfaceView path, the callback can issue additional
PixelCopy.request(surfaceView, …)calls afterpause()returned. Those are the ones most likely to hit a torn-down surface or get their destination bitmap recycled from under them.
Refined mitigations
Replacing item 2 in my previous comment with:
- Tighten the entry trigger. Pause replay on
Application.ActivityLifecycleCallbacks.onActivityStoppeddirectly (not viaProcessLifecycleOwner) — eliminates the 700ms gap. Re-resume ononActivityResumed. Needs a small re-entry guard so quick activity-to-activity transitions don't churn pause/resume. - Make in-flight callbacks check the paused state too. In the PixelCopy callback (both the window callback and the per-SurfaceView callbacks), early-return if we're paused since the request was issued. Don't issue follow-up SurfaceView captures from a paused state.
- Defer
screenshot.recycle()past the in-flight blit (unchanged from prior comment). Even with Mavenized, HttpClient removed and proxy support added #1 + Sentry Versions Supported #2, a request already accepted by the RenderThread can still SIGABRT if we recycle the destination during its blit; the fix has to fence the recycle, not just the next capture.
Items 1 & 2 close most of the surface-teardown window. Item 3 is needed regardless because the RenderThread blit window is opaque to us — once
PixelCopy.requestreturns, we have no signal for "blit started" vs "blit will start later," only "blit finished" (the callback).- Any
follow up and also cache bitmaps used for surfaceview capturing after this lands: #5333

- added a commit that references this issue
on Jul 24, 2026 github-actions commented
on Jul 29, 2026 on Jul 29, 2026 – with GitHub ActionsContributorMore actionsA PR closing this issue has just been released 🚀
This issue was referenced by PR #5808, which was included in the 8.51.0 release.
- added 7 commits that reference this issue
on Jul 30, 2026
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsWaiting for: Product Owner
Summary
The PixelCopy-based replay strategy (
PixelCopyStrategy) shares a singlescreenshotBitmapbetween two writers:screenshotwhen the next frame'sPixelCopy.request(window, screenshot, …)performs its blit.screenshotinsideapplyMaskingAndNotify(mask render) andcompositeSurfaceViewsAndMask(SurfaceView composite).If a frame's executor work is still in progress when the next recorder cycle fires
PixelCopy.request, the system PixelCopy worker overwrites parts ofscreenshotwhile the executor is still touching it. The executor being single-threaded prevents executor↔executor overlap, but does not gate the system PixelCopy worker.The same shape exists in the no-SurfaceView mask path; the SurfaceView capture feature (#5333) just makes it easier to hit because compositing adds a draw pass to the executor side.
What this looks like in practice
Visual:
screenshotends up with a strip of old composited content (mask + SurfaceView already applied) and a strip of newly captured raw window content. Visually, a horizontal seam in the replay frame.PixelCopy.requestoverwrites window pixels whileMaskRenderer.renderMasksis still rendering, masks can land on stale pixels or in the wrong region for one frame.Native crashes (likely related to #4696):
libhwui.sowith the typical__android_log_assert→art::Runtime::Abortchain. PixelCopy runs on the RenderThread on top of libhwui; when its hardware blit lands on aBitmapwhose backing hardware buffer is being concurrently drawn into by a Canvas on the executor, libhwui'sLOG_ALWAYS_FATAL/CHECKinvariants fail and the process aborts.getFrame() called on a context with no surface!— variant of the above where the Window's surface is torn down (e.g. activity backgrounded, surface recreated) betweenPixelCopy.requestand the actual blit. Reported in #4696.close()—close()recyclesscreenshoton the executor; we guard withisClosedin the PixelCopy callback, but the guard is checked when the callback runs, not at the RenderThread blit moment. If recycle happens betweenPixelCopy.requestand the blit, libhwui sees a recycled bitmap and aborts.When it manifests
PixelCopy.requestlands.isCaptureSurfaceViews = true(experimental, opt-in via #5333) widens the executor work window because compositing adds another draw pass before masking.Proposed fix
A
frameInFlight: AtomicBooleanset on the main thread immediately beforePixelCopy.request(window, screenshot, …)fires, and cleared at the end ofapplyMaskingAndNotifyon the executor. The next recorder cycle skips its capture if the flag is set (drops the frame, matches the codebase's existing drop a frame rather than fight pattern). Single-threaded executor means we only need to coordinate the one main → executor → main handoff.Additionally,
close()should fence against in-flight PixelCopy by either:screenshot.recycle()after the executor confirms no work is queued andframeInFlightis clear, orAlternatives considered:
screenshot— blocks the main thread waiting for the executor.Out of scope here
This is independent of the
windowLocation/svLocationfield-vs-local races already addressed in #5333 — those snapshots are already taken into locals before the executor handoff.Related
libhwui.sowith PixelCopy strategy; likely a manifestation of this race (especially the surface-teardown variant).