Status: tried and rejected (2026-10-07): the two queues overlap, but the panels run about half as fast beside the products; see Tried.
The blocked QR (qr-blocked.md) runs its forward pass in one order: an aggregate’s panels, then its update of every column right of it, then the next aggregate’s panels. The panels are latency-bound (a TSQR’s leaves are a simdgroup each, its top one threadgroup), the trailing update is MPS products that fill the GPU. With a look-ahead the next aggregate’s 128 columns are updated first, so its panels can start while the rest of the trailing update runs:
stream P: update(a-1 -> cols of a) panels(a) merge(a) update(a -> cols of a+1) panels(a+1) ...
stream T: trail(a-1) trail(a) ...
trail(a) (the columns from aggregate a + 2 on) waits for panels(a);
the look-ahead update of aggregate a + 1’s columns waits for trail(a - 1),
which brought them up to date with the aggregates before.
At 4096 x 4096 on an M5 Pro the forward pass takes 42 ms of the call’s 62: the panels about 25, the trailing updates 14, the updates inside aggregates
About a day.
Up to 14 ms of 62 at 4096 (1.2x), less at 1024-2048 where the panels are a larger share.
qr_blocks in src/band_reduce.mm.
Built as planned (A, V and Ta as two MTLBuffers over host memory, a second queue, two events an aggregate). The command buffers’ GPU times show the overlap: each trailing update runs beside the next aggregate’s panels. But an aggregate’s panels then took 1.7-2.0 ms instead of about 1.3, the products taking the GPU’s cores from them: 4096 x 4096 in 62.2 ms against 61.2 without, 1024 in 7.0 against 6.8. As in the band reduction (2.13.0), the panels gain nothing from sharing the GPU with MPS’s products. Reverted.