From 9ac3172f5bc412b590b4a34e19e3fc6b24e434b7 Mon Sep 17 00:00:00 2001
From: Christian Schmittel <90287914+christian-wr@users.noreply.github.com>
Date: Tue, 29 Sep 2026 19:04:54 +0200
Subject: [PATCH 1/2] fix(compositor): copy the decoder surface before sampling
it
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
`nv12_srvs` built its Y/UV views straight on the ffmpeg D3D11VA output — a texture
array slice, addressed through `FirstArraySlice`. Two documented rules say that is not
a supported way to read decoded video.
`D3D11_BIND_FLAG` is explicit about the first: "you cannot use texture arrays that are
created with this flag in calls to `ID3D11Device::CreateShaderResourceView`". Our pool
is exactly such an array — `get_hw_format` asks for `initial_pool_size = 32` with
`D3D11_BIND_DECODER | D3D11_BIND_SHADER_RESOURCE`. Drivers are free to let the view
creation succeed anyway, and most do, which is why this held up everywhere else.
The second rule is the one that actually bit. The decoded surface stays the decoder's
reference frame, and between the video engine and the 3D pipeline "there is no automatic
hazard tracking" — so the shader may sample a surface the decoder is concurrently
rewriting. ffmpeg's `ID3D11VideoContext` is the very same object as our immediate
context (it comes out of a `QueryInterface` on it), and `SetMultithreadProtected(TRUE)`
only makes an individual call atomic, never a sequence. Nothing ordered the decode
against our draws.
On a Snapdragon X Elite (Adreno X1-85) the editor preview drew opaque black for most of
playback. Measured on 1.13, same build, same recording, the two paths selected by an
environment variable: 768 black frames out of 792 sampling the decoder surface, 2 out of
545 through the copy — and those two are the transparent frames before the first compose,
not black ones. The black was opaque and total, so not "one frame late" but no frame at
all.
What hid the cause for so long: exporting the same project never produced a single black
frame (28 342 verified), because the export drains the GPU every frame through the
encoder, which lets the decode finish before the read. Everything that slowed the live
loop — one more readback, a lock, a sleep — cut the black proportionally without ever
removing it. Four mitigations along those lines were tried and discarded: a `Flush` and
then a `D3D11_QUERY_EVENT` wait around the readback, holding `ID3D11Multithread::Enter`
across compose and readback (90 % black down to 57 %), and capping the loop period
(69 % at 0 ms, 38 % at 33 ms). They were all treating the symptom.
The fix copies the slice into a private texture — `ArraySize = 1`, `BIND_SHADER_RESOURCE`
only — with `CopySubresourceRegion` on the immediate context, which IS ordered against
the draws that follow, and samples that copy. One allocation per decoder texture, then
one GPU→GPU copy per frame and per source; nothing goes back to system memory.
`clear_srv_cache` keeps its meaning and now also drops the copies, so a new decoder
texture landing on a recycled address cannot inherit one sized for the old one.
Export output is unchanged: the same project re-exported frame for frame identical,
326 frames, mean luminance 168.0, min 84.3.
This also supersedes the diagnosis in `feat/compositor-force-cpu-backend`, which read the
same symptom as a broken D3D11 hardware path on that adapter and added
`OPENSCREEN_FORCE_CPU_BACKEND` to work around it. The compose path was never at fault —
the export proves it on the same GPU — so that override is no longer the answer here.
Not covered: no regression test. The failure only appears where the video engine and the
3D pipeline actually race, so it does not reproduce on the CI adapters, and a test that
passes everywhere would prove nothing. The change is exercised by the existing compose
and export tests; the evidence above is the A/B on the affected hardware.
Refs:
- https://learn.microsoft.com/windows/win32/api/d3d11/ne-d3d11-d3d11_bind_flag
- https://learn.microsoft.com/windows/win32/direct3d11/overviews-direct3d-11-devices-intro#threading-considerations
- https://learn.microsoft.com/windows/win32/api/d3d11_4/nn-d3d11_4-id3d11multithread
- https://learn.microsoft.com/windows/win32/api/dxgiformat/ne-dxgiformat-dxgi_format
---
crates/compositor/src/compositor_windows.rs | 112 ++++++++++++++------
1 file changed, 82 insertions(+), 30 deletions(-)
diff --git a/crates/compositor/src/compositor_windows.rs b/crates/compositor/src/compositor_windows.rs
index 59431e383..ec31b943d 100644
--- a/crates/compositor/src/compositor_windows.rs
+++ b/crates/compositor/src/compositor_windows.rs
@@ -25,7 +25,7 @@ use std::collections::HashMap;
use std::ffi::c_void;
use windows::core::Interface;
use windows::Win32::Graphics::Direct3D::{
- D3D11_SRV_DIMENSION_TEXTURE2DARRAY, D3D_PRIMITIVE_TOPOLOGY_TRIANGLESTRIP,
+ D3D11_SRV_DIMENSION_TEXTURE2D, D3D_PRIMITIVE_TOPOLOGY_TRIANGLESTRIP,
};
use windows::Win32::Graphics::Direct3D11::*;
use windows::Win32::Graphics::Dxgi::Common::*;
@@ -149,9 +149,11 @@ pub struct Compositor {
timeline_t_override: RefCell