metal-linalg

QR dispatch crossover — Apple M5 Pro

What every section and number below means: reading-reports.md.

Generated by tuning/tune_qr.py. 173 shapes, 519 shape/backend pairs measured twice.

   
Device Apple M5 Pro
GPU cores 20
Resident matrices (qr_unblocked) 120
Shapes measured 173
Coverage near-square 29, square 70, tall 44, wide 30

The band

M >= 512  ->  qr_streaming_amx_reduced
otherwise  ->  qr_unblocked

The optimum is flat from 480 to 512 rows (every threshold within 0.3% of the best, 1.0193x). Any value inside that band is equivalent on this hardware; 512 is the middle of it.

Paste into kTuned[] in src/qr.mm:

    {"Apple M5 Pro", 20, 512, 512, 16,   kQrNoLimit, 512, 1},

Cost of missing the band

threshold regret excess  
128 1.1616x +13.96%  
192 1.1121x +9.10%  
256 1.0956x +7.49%  
288 1.0534x +3.35%  
320 1.0534x +3.35%  
352 1.0365x +1.69%  
384 1.0365x +1.69%  
416 1.0273x +0.78%  
448 1.0273x +0.78%  
480 1.0193x +0.00% in band
512 1.0193x +0.00% in band
576 1.0299x +1.04%  
640 1.0299x +1.04%  
768 1.0627x +4.26%  
1024 1.0688x +4.86%  

The penalty is usually asymmetric. Erring low costs little; erring high degrades specifically on tall inputs. Adding GPU cores makes the grid-parallel backend relatively stronger and pushes the true crossover down, so an untuned device is safer low than high.

GPU or CPU

GPU iff k <= no limit, batch * k >= 512 and batch >= 1   (k = min(M, N))
otherwise LAPACK on the CPU

Fitted on every shape with a CPU timing; inside the region within 0.5% of the best geometric-mean regret, the candidate with the smallest worst case.

routing geomean regret worst >10% off
chosen 1.0820x 5.28x 20/173
always the GPU (before the CPU path) 1.6737x 61.31x 71/173
always the CPU 1.8958x 14.06x 98/173

Held out (fitted on half the shapes, scored on the other half): 1.1432x geomean, 5.28x worst, against 1.6061x and 51.38x for always the GPU.

The CPU was the fastest backend at 71 of 173 shapes, e.g. 1 x 8x8, 1 x 16x16, 1 x 16x256, 1 x 32x32, 1 x 64x64, 1 x 64x256, 1 x 64x384, 1 x 64x640.

Regret by threshold

Regret is how much slower the chosen backend is than the best one measured at that shape; 1.0000x would be a perfect oracle. The per-region curves are the point: a grid covering only one aspect ratio can be nearly flat, and will then pick a threshold out of noise.

xychart-beta
    title "Regret by threshold (lower is better)"
    x-axis [128, 192, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 768, 1024]
    y-axis "geometric-mean regret" 1.0 --> 1.30
    line [1.1616, 1.1121, 1.0956, 1.0534, 1.0534, 1.0365, 1.0365, 1.0273, 1.0273, 1.0193, 1.0193, 1.0299, 1.0299, 1.0627, 1.0688]
    line [1.0465, 1.0338, 1.0338, 1.0240, 1.0240, 1.0240, 1.0240, 1.0453, 1.0453, 1.0453, 1.0453, 1.0993, 1.0993, 1.2048, 1.2048]
    line [1.2829, 1.2065, 1.1437, 1.0926, 1.0926, 1.0551, 1.0551, 1.0290, 1.0290, 1.0118, 1.0118, 1.0013, 1.0013, 1.0081, 1.0287]

Series order: pooled, tall only, square only.

threshold pooled square tall wide near-square pooled worst
128 1.1616 1.2829 1.0465 1.0641 1.2993 2.11x
192 1.1121 1.2065 1.0338 1.0222 1.2114 1.80x
256 1.0956 1.1437 1.0338 1.0222 1.2114 1.67x
288 1.0534 1.0926 1.0240 1.0000 1.1035 1.58x
320 1.0534 1.0926 1.0240 1.0000 1.1035 1.58x
352 1.0365 1.0551 1.0240 1.0000 1.0690 1.58x
384 1.0365 1.0551 1.0240 1.0000 1.0690 1.58x
416 1.0273 1.0290 1.0453 1.0000 1.0264 1.58x
448 1.0273 1.0290 1.0453 1.0000 1.0264 1.58x
480 1.0193 1.0118 1.0453 1.0000 1.0109 1.58x
512 1.0193 1.0118 1.0453 1.0000 1.0109 1.58x
576 1.0299 1.0013 1.0993 1.0000 1.0000 1.94x
640 1.0299 1.0013 1.0993 1.0000 1.0000 1.94x
768 1.0627 1.0081 1.2048 1.0000 1.0062 2.47x
1024 1.0688 1.0287 1.2048 1.0000 1.0062 2.47x

Which feature decides

qr_unblocked gives each matrix a single threadgroup, which sweeps M rows per Householder reflection while N parallelises across that threadgroup’s threads. Rows and columns are therefore not interchangeable, and a rule on max(M, N) cannot express the difference.

feature best threshold geomean regret worst
M (rows) 480 1.0193x 1.58x
max(M, N) 576 1.0722x 2.08x
K = min(M, N) 576 1.1414x 6.67x

Decision surface

batch 1

      N=32      N=64      N=128     N=256     N=512     
      ---------------------------------------------
  128    1.36     1.77     1.92     1.87     1.49
  256 ~  1.04  .  1.23     1.56     1.67     1.49
  384 ## 0.66  #  0.94  .  1.28     1.30  .  1.24
  512 ###0.52  ## 0.78  .  1.09  ~  1.04  .  1.07   <- M >= 512
  640 ###0.41  ## 0.61  #  0.85  ## 0.83  ## 0.84

      ratio = reduced / unblocked.  ### <0.60  ## <0.85  # <0.95
      ~ tie (0.95-1.05)   . <1.30   blank >1.30  (unblocked wins)

batch 16

      N=32      N=64      N=128     N=256     N=512     
      ---------------------------------------------
  128 .  1.26     1.49     1.62     1.55     1.39
  256 ~  0.98  .  1.21     1.45     1.54     1.36
  384 ## 0.69  #  0.94  .  1.17  .  1.27  .  1.22
  512 ###0.59  ## 0.74  ~  1.01  ~  1.05  .  1.18   <- M >= 512
  640 ###0.43  ###0.58  ## 0.75  #  0.90  ~  1.00

      ratio = reduced / unblocked.  ### <0.60  ## <0.85  # <0.95
      ~ tie (0.95-1.05)   . <1.30   blank >1.30  (unblocked wins)

Candidate rules

Per-region columns matter more than the overall number: an aggregate can look excellent while one aspect class is badly served.

rule geomean worst >10% off square tall wide near-square
original 5-rule heuristic 1.1953x 2.08x 62/143 1.148 1.042 1.534 1.203
max(M,N) >= 512 1.0867x 2.08x 29/143 1.012 1.045 1.307 1.051
K = min(M,N) >= 128 1.2781x 6.67x 75/143 1.283 1.459 1.064 1.257
M >= 512 (chosen) 1.0193x 1.58x 7/143 1.012 1.045 1.000 1.011

Refinements tested

Both of these were rejected on an 8-core M1, and both are re-tested here rather than assumed. A GPU with many more cores saturates qr_unblocked much later, so the batch split in particular could be justified elsewhere. Each is fitted on half the shapes and scored on the other half.

Baseline for comparison — M >= 512 on the held-out half: 1.0244x geomean, 1.58x worst.

refinement fitted form train held-out held-out worst verdict
batch-dependent split M >= (480 if batch < 32 else 416) 1.0142x 1.0265x 1.58x reject
narrow-N special case M >= 416 or (N <= 64 and M >= 288) 1.0142x 1.0217x 1.58x adopt

At least one refinement cleared held-out validation on this GPU. QrPolicy already carries the batch-split fields (m_crossover_small_batch / m_crossover_large_batch / batch_threshold) — set them from the fitted form above.

Measurement noise

Run-to-run ratio over 519 shape/backend pairs measured twice: median 1.013x, p90 1.097x, max 3.169x.

runtime pairs median p90 max
<1 ms 173 1.031x 1.440x 3.169x
1-3 ms 135 1.008x 1.038x 1.219x
3-10 ms 112 1.011x 1.044x 1.081x
10-30 ms 47 1.009x 1.025x 1.084x
30-100 ms 27 1.011x 1.029x 1.138x
>100 ms 25 1.012x 1.036x 1.114x

Noise concentrates in fast shapes, where submission jitter dominates. A difference smaller than the p90 figure for its size bucket is not a result.


Raw timings in raw.csv; full analysis including plot-ready series in results.json. Re-render this report without remeasuring: python3 tuning/tune_qr.py --from results.json.