What every section and number below means: reading-reports.md.
Generated by tuning/tune_qr.py. 173 shapes, 519 shape/backend pairs measured twice.
| Device | Apple M5 Pro |
| GPU cores | 20 |
Resident matrices (qr_unblocked) |
120 |
| Shapes measured | 173 |
| Coverage | near-square 29, square 70, tall 44, wide 30 |
M >= 512 -> qr_streaming_amx_reduced
otherwise -> qr_unblocked
The optimum is flat from 480 to 512 rows (every threshold within 0.3% of the best, 1.0193x). Any value inside that band is equivalent on this hardware; 512 is the middle of it.
Paste into kTuned[] in src/qr.mm:
{"Apple M5 Pro", 20, 512, 512, 16, kQrNoLimit, 512, 1},
| threshold | regret | excess | |
|---|---|---|---|
| 128 | 1.1616x | +13.96% | |
| 192 | 1.1121x | +9.10% | |
| 256 | 1.0956x | +7.49% | |
| 288 | 1.0534x | +3.35% | |
| 320 | 1.0534x | +3.35% | |
| 352 | 1.0365x | +1.69% | |
| 384 | 1.0365x | +1.69% | |
| 416 | 1.0273x | +0.78% | |
| 448 | 1.0273x | +0.78% | |
| 480 | 1.0193x | +0.00% | in band |
| 512 | 1.0193x | +0.00% | in band |
| 576 | 1.0299x | +1.04% | |
| 640 | 1.0299x | +1.04% | |
| 768 | 1.0627x | +4.26% | |
| 1024 | 1.0688x | +4.86% |
The penalty is usually asymmetric. Erring low costs little; erring high degrades specifically on tall inputs. Adding GPU cores makes the grid-parallel backend relatively stronger and pushes the true crossover down, so an untuned device is safer low than high.
GPU iff k <= no limit, batch * k >= 512 and batch >= 1 (k = min(M, N))
otherwise LAPACK on the CPU
Fitted on every shape with a CPU timing; inside the region within 0.5% of the best geometric-mean regret, the candidate with the smallest worst case.
| routing | geomean regret | worst | >10% off |
|---|---|---|---|
| chosen | 1.0820x | 5.28x | 20/173 |
| always the GPU (before the CPU path) | 1.6737x | 61.31x | 71/173 |
| always the CPU | 1.8958x | 14.06x | 98/173 |
Held out (fitted on half the shapes, scored on the other half): 1.1432x geomean, 5.28x worst, against 1.6061x and 51.38x for always the GPU.
The CPU was the fastest backend at 71 of 173 shapes, e.g. 1 x 8x8, 1 x 16x16, 1 x 16x256, 1 x 32x32, 1 x 64x64, 1 x 64x256, 1 x 64x384, 1 x 64x640.
Regret is how much slower the chosen backend is than the best one measured at that shape; 1.0000x would be a perfect oracle. The per-region curves are the point: a grid covering only one aspect ratio can be nearly flat, and will then pick a threshold out of noise.
xychart-beta
title "Regret by threshold (lower is better)"
x-axis [128, 192, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 768, 1024]
y-axis "geometric-mean regret" 1.0 --> 1.30
line [1.1616, 1.1121, 1.0956, 1.0534, 1.0534, 1.0365, 1.0365, 1.0273, 1.0273, 1.0193, 1.0193, 1.0299, 1.0299, 1.0627, 1.0688]
line [1.0465, 1.0338, 1.0338, 1.0240, 1.0240, 1.0240, 1.0240, 1.0453, 1.0453, 1.0453, 1.0453, 1.0993, 1.0993, 1.2048, 1.2048]
line [1.2829, 1.2065, 1.1437, 1.0926, 1.0926, 1.0551, 1.0551, 1.0290, 1.0290, 1.0118, 1.0118, 1.0013, 1.0013, 1.0081, 1.0287]
Series order: pooled, tall only, square only.
| threshold | pooled | square | tall | wide | near-square | pooled worst |
|---|---|---|---|---|---|---|
| 128 | 1.1616 | 1.2829 | 1.0465 | 1.0641 | 1.2993 | 2.11x |
| 192 | 1.1121 | 1.2065 | 1.0338 | 1.0222 | 1.2114 | 1.80x |
| 256 | 1.0956 | 1.1437 | 1.0338 | 1.0222 | 1.2114 | 1.67x |
| 288 | 1.0534 | 1.0926 | 1.0240 | 1.0000 | 1.1035 | 1.58x |
| 320 | 1.0534 | 1.0926 | 1.0240 | 1.0000 | 1.1035 | 1.58x |
| 352 | 1.0365 | 1.0551 | 1.0240 | 1.0000 | 1.0690 | 1.58x |
| 384 | 1.0365 | 1.0551 | 1.0240 | 1.0000 | 1.0690 | 1.58x |
| 416 | 1.0273 | 1.0290 | 1.0453 | 1.0000 | 1.0264 | 1.58x |
| 448 | 1.0273 | 1.0290 | 1.0453 | 1.0000 | 1.0264 | 1.58x |
| 480 | 1.0193 | 1.0118 | 1.0453 | 1.0000 | 1.0109 | 1.58x |
| 512 | 1.0193 | 1.0118 | 1.0453 | 1.0000 | 1.0109 | 1.58x |
| 576 | 1.0299 | 1.0013 | 1.0993 | 1.0000 | 1.0000 | 1.94x |
| 640 | 1.0299 | 1.0013 | 1.0993 | 1.0000 | 1.0000 | 1.94x |
| 768 | 1.0627 | 1.0081 | 1.2048 | 1.0000 | 1.0062 | 2.47x |
| 1024 | 1.0688 | 1.0287 | 1.2048 | 1.0000 | 1.0062 | 2.47x |
qr_unblocked gives each matrix a single threadgroup, which sweeps M rows per Householder reflection while N parallelises across that threadgroup’s threads. Rows and columns are therefore not interchangeable, and a rule on max(M, N) cannot express the difference.
| feature | best threshold | geomean regret | worst |
|---|---|---|---|
| M (rows) | 480 | 1.0193x | 1.58x |
| max(M, N) | 576 | 1.0722x | 2.08x |
| K = min(M, N) | 576 | 1.1414x | 6.67x |
batch 1
N=32 N=64 N=128 N=256 N=512
---------------------------------------------
128 1.36 1.77 1.92 1.87 1.49
256 ~ 1.04 . 1.23 1.56 1.67 1.49
384 ## 0.66 # 0.94 . 1.28 1.30 . 1.24
512 ###0.52 ## 0.78 . 1.09 ~ 1.04 . 1.07 <- M >= 512
640 ###0.41 ## 0.61 # 0.85 ## 0.83 ## 0.84
ratio = reduced / unblocked. ### <0.60 ## <0.85 # <0.95
~ tie (0.95-1.05) . <1.30 blank >1.30 (unblocked wins)
batch 16
N=32 N=64 N=128 N=256 N=512
---------------------------------------------
128 . 1.26 1.49 1.62 1.55 1.39
256 ~ 0.98 . 1.21 1.45 1.54 1.36
384 ## 0.69 # 0.94 . 1.17 . 1.27 . 1.22
512 ###0.59 ## 0.74 ~ 1.01 ~ 1.05 . 1.18 <- M >= 512
640 ###0.43 ###0.58 ## 0.75 # 0.90 ~ 1.00
ratio = reduced / unblocked. ### <0.60 ## <0.85 # <0.95
~ tie (0.95-1.05) . <1.30 blank >1.30 (unblocked wins)
Per-region columns matter more than the overall number: an aggregate can look excellent while one aspect class is badly served.
| rule | geomean | worst | >10% off | square | tall | wide | near-square |
|---|---|---|---|---|---|---|---|
| original 5-rule heuristic | 1.1953x | 2.08x | 62/143 | 1.148 | 1.042 | 1.534 | 1.203 |
| max(M,N) >= 512 | 1.0867x | 2.08x | 29/143 | 1.012 | 1.045 | 1.307 | 1.051 |
| K = min(M,N) >= 128 | 1.2781x | 6.67x | 75/143 | 1.283 | 1.459 | 1.064 | 1.257 |
| M >= 512 (chosen) | 1.0193x | 1.58x | 7/143 | 1.012 | 1.045 | 1.000 | 1.011 |
Both of these were rejected on an 8-core M1, and both are re-tested here rather than assumed. A GPU with many more cores saturates qr_unblocked much later, so the batch split in particular could be justified elsewhere. Each is fitted on half the shapes and scored on the other half.
Baseline for comparison — M >= 512 on the held-out half: 1.0244x geomean, 1.58x worst.
| refinement | fitted form | train | held-out | held-out worst | verdict |
|---|---|---|---|---|---|
| batch-dependent split | M >= (480 if batch < 32 else 416) |
1.0142x | 1.0265x | 1.58x | reject |
| narrow-N special case | M >= 416 or (N <= 64 and M >= 288) |
1.0142x | 1.0217x | 1.58x | adopt |
At least one refinement cleared held-out validation on this GPU.
QrPolicyalready carries the batch-split fields (m_crossover_small_batch/m_crossover_large_batch/batch_threshold) — set them from the fitted form above.
Run-to-run ratio over 519 shape/backend pairs measured twice: median 1.013x, p90 1.097x, max 3.169x.
| runtime | pairs | median | p90 | max |
|---|---|---|---|---|
| <1 ms | 173 | 1.031x | 1.440x | 3.169x |
| 1-3 ms | 135 | 1.008x | 1.038x | 1.219x |
| 3-10 ms | 112 | 1.011x | 1.044x | 1.081x |
| 10-30 ms | 47 | 1.009x | 1.025x | 1.084x |
| 30-100 ms | 27 | 1.011x | 1.029x | 1.138x |
| >100 ms | 25 | 1.012x | 1.036x | 1.114x |
Noise concentrates in fast shapes, where submission jitter dominates. A difference smaller than the p90 figure for its size bucket is not a result.
Raw timings in raw.csv; full analysis including plot-ready series in results.json. Re-render this report without remeasuring: python3 tuning/tune_qr.py --from results.json.