What every section and number below means: reading-reports.md.
Cost model scaled to this device from one probe point: block x0.14, cpu x0.95, whole-matrix x1.00 (1.00 is an M1).
Machine state: load 2.2/18 at the start, load 1.2/18 at the end; power mains.
Probe point after the sweep relative to before it: block x0.96, cpu x1.01, tg x0.96 (stable).
Generated by tuning/tune_eigh.py from 203 (N, batch) points, N in [2, 4, 8, 12, 16, 24, 32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048], batch in [1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096], four backends, two or more passes, min-of-repeats.
Row for kTuned[] in src/eigh.mm:
// device, GPU cores, simd_max_n, block_min_n, block_min_n_batched, block_min_batch, gpu_max_n, gpu_min_batch_times_n, gpu_min_batch, values_gpu_max_n, values_gpu_min_batch_times_n, values_gpu_min_batch
{"Apple M5 Pro", 20, 0, 128, 64, 64, 1024, 512, 16, 256, 2048, 32},
To try it without rebuilding:
EIGH_SIMD_MAX_N=0 EIGH_BLOCK_MIN_N=128 EIGH_BLOCK_MIN_N_BATCHED=64 EIGH_BLOCK_MIN_BATCH=64 EIGH_GPU_MAX_N=1024 EIGH_GPU_MIN_BATCH_TIMES_N=512 EIGH_GPU_MIN_BATCH=16 EIGH_VALUES_GPU_MAX_N=256 EIGH_VALUES_GPU_MIN_BATCH_TIMES_N=2048 EIGH_VALUES_GPU_MIN_BATCH=32
The policy in effect on this device came from tuned:Apple M5 Pro. It differs from the fitted one; see the warnings.
Against the best measured backend at every point the whole rule scores 1.0142 geometric-mean regret, worst 1.75x, 10 of 203 points losing more than 10%, and 1.011x the oracle’s total time. The decision is fitted in two stages, below, because the CPU routing would otherwise hide the GPU backend crossover.
Scored against the best GPU backend at each of the 203 points, as if there were no CPU: this is the rule a forced-GPU call (EIGH_DEVICE=gpu) and the detail entry points follow, and it is what a GPU with more cores will lean on.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| policy in effect (‘0’, ‘96’, ‘none’, ‘none’) | 1.0197 | 1.56x | 12 | 1.004 | 0 |
| fitted (‘0’, ‘128’, ‘64’, ‘64’) | 1.0042 | 1.19x | 3 | 1.000 | 0 |
2 of 112 (simd_max_n, block_min_n) pairs are within 0.5% of the best geomean: simd_max_n 0 .. 2, block_min_n 96 .. 96.
With the batch term adopted (below), 4 combinations are within 0.5% of the best: block_min_n 128 .. 128, block_min_n_batched 64 .. 64, block_min_batch 32 .. 256. The curves below vary one constant around the chosen combination.
xychart-beta
title "Regret by block_min_n"
x-axis "block_min_n" [32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048, none]
y-axis "geometric-mean regret" 1.0 --> 1.76
line [1.1970, 1.1016, 1.0379, 1.0111, 1.0042, 1.0140, 1.0464, 1.0943, 1.1638, 1.2597, 1.3971, 1.5324, 1.6336, 1.7414]
| block_min_n | 32 | 48 | 64 | 96 | 128 | 192 | 256 | 384 | 512 | 768 | 1024 | 1536 | 2048 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.1970 | 1.1016 | 1.0379 | 1.0111 | 1.0042 | 1.0140 | 1.0464 | 1.0943 | 1.1638 | 1.2597 | 1.3971 | 1.5324 | 1.6336 | 1.7414 |
| worst | 8.72x | 4.42x | 2.57x | 1.36x | 1.19x | 1.74x | 3.39x | 5.51x | 10.08x | 15.61x | 15.61x | 15.61x | 15.61x | 15.61x |
xychart-beta
title "Regret by block_min_n_batched"
x-axis "block_min_n_batched" [32, 48, 64, 96]
y-axis "geometric-mean regret" 1.0 --> 1.06
line [1.0463, 1.0229, 1.0042, 1.0127]
| block_min_n_batched | 32 | 48 | 64 | 96 |
|---|---|---|---|---|
| geomean | 1.0463 | 1.0229 | 1.0042 | 1.0127 |
| worst | 4.89x | 2.35x | 1.19x | 1.56x |
xychart-beta
title "Regret by block_min_batch"
x-axis "block_min_batch" [1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096]
y-axis "geometric-mean regret" 1.0 --> 1.05
line [1.0379, 1.0316, 1.0254, 1.0192, 1.0132, 1.0070, 1.0042, 1.0054, 1.0081, 1.0126, 1.0180, 1.0235, 1.0288]
| block_min_batch | 1 | 2 | 4 | 8 | 16 | 32 | 64 | 128 | 256 | 512 | 1024 | 2048 | 4096 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.0379 | 1.0316 | 1.0254 | 1.0192 | 1.0132 | 1.0070 | 1.0042 | 1.0054 | 1.0081 | 1.0126 | 1.0180 | 1.0235 | 1.0288 |
| worst | 2.57x | 2.57x | 2.56x | 2.55x | 2.55x | 1.87x | 1.19x | 1.52x | 1.77x | 1.95x | 1.98x | 1.98x | 1.98x |
xychart-beta
title "Regret by simd_max_n"
x-axis "simd_max_n" [0, 2, 4, 8, 12, 16, 24, 32]
y-axis "geometric-mean regret" 1.0 --> 1.23
line [1.0197, 1.0223, 1.0270, 1.0390, 1.0702, 1.0947, 1.1389, 1.2100]
| simd_max_n | 0 | 2 | 4 | 8 | 12 | 16 | 24 | 32 |
|---|---|---|---|---|---|---|---|---|
| geomean | 1.0197 | 1.0223 | 1.0270 | 1.0390 | 1.0702 | 1.0947 | 1.1389 | 1.2100 |
| worst | 1.56x | 1.56x | 1.85x | 1.85x | 3.26x | 3.26x | 3.26x | 5.00x |
Held-out check of a batch-dependent crossover (block from a lower N once the batch is large enough). Fitted on 110 points, scored on the other 93; the verdict is a bootstrap over the test points.
| rule | fitted on train | train geomean | test geomean | test worst | verdict |
|---|---|---|---|---|---|
| two constants | [0, 96, 1000000000, 1000000000] | 1.0191 | 1.0202 | 1.54x | baseline |
| batch-dependent block crossover | {“block_lo”: 64, “batch_hi”: 64} | 1.0044 | 1.0040 | 1.10x | justified (better in 100% of resamples, median gain 1.6%) |
Best GPU backend per point (s simd, t threadgroup, B block), then what the split picks:
N \ batch 1 2 4 8 16 32 64 128 256 512 1024 2048 4096
2 s t t t t s s s t s t t t
4 s t s s t t s s t s s t s
8 t s t t t s s t t t t t t
12 t t t t t t t t t t t s t
16 t t t t t t t t t t t t t
24 t t t t t t t t t t t t t
32 t t t t t t t t t t t t t
48 t t t t t t t t t t t t t
64 t t t t t t t t B B B B B
96 t t t t t B B B B B B B B
128 B B B B B B B B B B B B B
192 B B B B B B B B B B B B B
256 B B B B B B B B B B B . .
384 B B B B B B B B B B . . .
512 B B B B B B B B . . . . .
768 B B B B B B B . . . . . .
1024 B B B B B . . . . . . . .
1536 B B B . . . . . . . . . .
2048 B B B . . . . . . . . . .
N \ batch 1 2 4 8 16 32 64 128 256 512 1024 2048 4096
2 t t t t t t t t t t t t t
4 t t t t t t t t t t t t t
8 t t t t t t t t t t t t t
12 t t t t t t t t t t t t t
16 t t t t t t t t t t t t t
24 t t t t t t t t t t t t t
32 t t t t t t t t t t t t t
48 t t t t t t t t t t t t t
64 t t t t t t B B B B B B B
96 t t t t t t B B B B B B B
128 B B B B B B B B B B B B B
192 B B B B B B B B B B B B B
256 B B B B B B B B B B B . .
384 B B B B B B B B B B . . .
512 B B B B B B B B . . . . .
768 B B B B B B B . . . . . .
1024 B B B B B . . . . . . . .
1536 B B B . . . . . . . . . .
2048 B B B . . . . . . . . . .
Given the split above, GPU iff N <= gpu_max_n, batch * N >= gpu_min_batch_times_n and batch >= gpu_min_batch, scored against the best of all four backends. worst is over the points where the chosen backend was timed; a pick the cost model had to guess is listed in the warnings instead.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| oracle (best per point) | 1.0000 | 1.00x | 0 | 1.000 | 0 |
| policy in effect (‘1024’, ‘512’, ‘16’) | 1.0142 | 1.75x | 10 | 1.011 | 0 |
| fitted (‘1024’, ‘512’, ‘16’) | 1.0142 | 1.75x | 10 | 1.011 | 0 |
8 of 1056 combinations are within 0.5% of the best geomean: gpu_max_n 1024 .. none, gpu_min_batch_times_n 512 .. 1024, gpu_min_batch 16 .. 16.
xychart-beta
title "Regret by gpu_min_batch_times_n"
x-axis "gpu_min_batch_times_n" [0, 64, 128, 256, 512, 1024, 2048, 4096, 8192, 16384, none]
y-axis "geometric-mean regret" 1.0 --> 1.98
line [1.1147, 1.0979, 1.0721, 1.0377, 1.0142, 1.0191, 1.0492, 1.1102, 1.2041, 1.3313, 1.9629]
| gpu_min_batch_times_n | 0 | 64 | 128 | 256 | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.1147 | 1.0979 | 1.0721 | 1.0377 | 1.0142 | 1.0191 | 1.0492 | 1.1102 | 1.2041 | 1.3313 | 1.9629 |
| worst | 21.69x | 12.86x | 7.56x | 3.51x | 1.75x | 2.16x | 3.46x | 5.25x | 6.81x | 9.69x | 16.82x |
xychart-beta
title "Regret by gpu_min_batch"
x-axis "gpu_min_batch" [1, 2, 4, 8, 16, 32]
y-axis "geometric-mean regret" 1.0 --> 1.09
line [1.0781, 1.0604, 1.0407, 1.0229, 1.0142, 1.0339]
| gpu_min_batch | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| geomean | 1.0781 | 1.0604 | 1.0407 | 1.0229 | 1.0142 | 1.0339 |
| worst | 3.40x | 3.06x | 3.06x | 1.75x | 1.75x | 1.77x |
xychart-beta
title "Regret by gpu_max_n"
x-axis "gpu_max_n" [16, 24, 32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048, none]
y-axis "geometric-mean regret" 1.0 --> 1.58
line [1.5676, 1.4594, 1.3633, 1.2834, 1.2116, 1.1577, 1.1113, 1.0727, 1.0485, 1.0340, 1.0285, 1.0200, 1.0142, 1.0142, 1.0142, 1.0142]
| gpu_max_n | 16 | 24 | 32 | 48 | 64 | 96 | 128 | 192 | 256 | 384 | 512 | 768 | 1024 | 1536 | 2048 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.5676 | 1.4594 | 1.3633 | 1.2834 | 1.2116 | 1.1577 | 1.1113 | 1.0727 | 1.0485 | 1.0340 | 1.0285 | 1.0200 | 1.0142 | 1.0142 | 1.0142 | 1.0142 |
| worst | 10.71x | 8.29x | 5.79x | 5.79x | 3.81x | 3.10x | 2.34x | 2.04x | 1.75x | 1.75x | 1.75x | 1.75x | 1.75x | 1.75x | 1.75x | 1.75x |
Held-out check of a per-N boundary (a lookup table of the smallest batch at which the GPU wins, per N) against the product rule. Fitted on 110 points, scored on the other 93.
| rule | fitted on train | train geomean | test geomean | test worst | verdict |
|---|---|---|---|---|---|
| product rule | [1024, 512, 16] | 1.0118 | 1.0171 | 1.75x | baseline |
| per-N table | {“min_batch_by_n”: {“2”: null, “4”: 512, “8”: 128, “12”: 256, “16”: 32, “24”: 64, “32”: 16, “48”: 16, “64”: 32, “96”: 64, “128”: 64, “192”: 16, “256”: 16, “384”: 8, “512”: 16, “768”: 32, “1024”: null, “1536”: null, “2048”: null}} | 1.0024 | 1.1090 | 6.52x | rejected (better in 0% of resamples, median gain -8.2%) |
Best backend per point (c CPU, s simd, t threadgroup, B block, . not measured), what the whole rule picks, and the speedup of the best GPU backend over the CPU:
N \ batch 1 2 4 8 16 32 64 128 256 512 1024 2048 4096
2 c c c c c c c c c c t t t
4 c c c c c c c c t s s t s
8 c c c c c c c t t t t t t
12 c c c c c c t t t t t s t
16 c c c c c t t t t t t t t
24 c c c c t t t t t t t t t
32 c c c c t t t t t t t t t
48 c c c c t t t t t t t t t
64 c c c c t t t t B B B B B
96 c c c c t B B B B B B B B
128 c c c c B B B B B B B B B
192 c c c c B B B B B B B B B
256 c c c B B B B B B B B . .
384 c c c B B B B B B B . . .
512 c c c B B c c B . . . . .
768 c c c c c B B . . . . . .
1024 c c c c B . . . . . . . .
1536 c c c . . . . . . . . . .
2048 c c c . . . . . . . . . .
N \ batch 1 2 4 8 16 32 64 128 256 512 1024 2048 4096
2 c c c c c c c c t t t t t
4 c c c c c c c t t t t t t
8 c c c c c c t t t t t t t
12 c c c c c c t t t t t t t
16 c c c c c t t t t t t t t
24 c c c c c t t t t t t t t
32 c c c c t t t t t t t t t
48 c c c c t t t t t t t t t
64 c c c c t t B B B B B B B
96 c c c c t t B B B B B B B
128 c c c c B B B B B B B B B
192 c c c c B B B B B B B B B
256 c c c c B B B B B B B . .
384 c c c c B B B B B B . . .
512 c c c c B B B B . . . . .
768 c c c c B B B . . . . . .
1024 c c c c B . . . . . . . .
1536 c c c . . . . . . . . . .
2048 c c c . . . . . . . . . .
N \ batch 1 2 4 8 16 32 64 128 256 512 1024 2048 4096
2 0.01 0.01 0.01 0.02 0.05 0.08 0.13 0.29 0.61 0.89 1.82 3.56 6.52
4 0.01 0.01 0.03 0.04 0.10 0.14 0.32 0.62 1.09 2.33 4.20 7.10 11.43
8 0.02 0.04 0.06 0.11 0.20 0.36 0.85 1.49 2.96 5.32 9.04 12.42 16.82
12 0.03 0.05 0.11 0.19 0.36 0.51 1.41 2.72 4.43 6.81 9.69 11.03 14.36
16 0.04 0.08 0.16 0.31 0.54 1.14 2.26 4.07 5.94 8.61 10.18 12.76 14.59
24 0.08 0.15 0.30 0.56 1.13 2.16 3.46 5.25 6.23 8.10 9.21 10.33 10.71
32 0.10 0.19 0.37 0.73 1.46 2.50 3.71 4.17 5.32 6.88 7.59 7.84 8.29
48 0.10 0.21 0.43 0.85 1.77 2.54 3.15 3.92 4.72 5.04 5.35 5.47 5.43
64 0.11 0.23 0.45 0.87 1.73 2.39 2.66 3.26 4.33 5.25 5.59 5.56 5.79
96 0.08 0.15 0.31 0.62 1.21 1.64 2.31 3.11 3.75 3.80 3.73 3.77 3.81
128 0.09 0.17 0.33 0.65 1.21 1.93 2.52 3.00 3.10 2.94 2.90 2.94 2.95
192 0.13 0.26 0.50 0.93 1.57 1.99 2.34 2.31 2.11 2.08 2.13 gpu gpu
256 0.19 0.37 0.69 1.19 1.68 2.00 2.04 1.77 1.73 1.75 gpu . .
384 0.23 0.42 0.71 1.14 1.33 1.29 1.12 1.05 gpu gpu . . .
512 0.29 0.49 0.81 1.11 1.12 0.99 0.93 gpu . . . . .
768 0.30 0.49 0.70 0.63 0.60 gpu gpu . . . . . .
1024 0.39 0.59 0.67 0.58 gpu . . . . . . . .
1536 0.40 0.42 0.40 . . . . . . . . . .
2048 0.41 0.40 0.39 . . . . . . . . . .
The same rule for eigvalsh, with its own thresholds (values_gpu_max_n, values_gpu_min_batch_times_n, values_gpu_min_batch), fitted on the _vals timings of 194 points given the split above. The CPU computes eigenvalues alone by LAPACK’s two-stage reduction from N = 128, so the boundary need not be eigh’s.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| eigh’s boundary, in effect | 1.0610 | 3.90x | 27 | 1.078 | 0 |
| fitted (‘256’, ‘2048’, ‘32’) | 1.0075 | 1.28x | 7 | 1.000 | 0 |
5 combinations are within 0.5% of the best geomean: values_gpu_max_n 192 .. 384, values_gpu_min_batch_times_n 1024 .. 2048, values_gpu_min_batch 16 .. 32.
Held out: fitted on 105 points (‘256’, ‘2048’, ‘32’), scored on the other 89: geomean 1.0068x, worst 1.26x, against 1.0683x, worst 3.90x for eigh’s boundary on the same points.
Pass-to-pass ratio (max/min of the same measurement across passes), 1150 measurements: median 1.013, p90 1.135, max 3.09. The held-out verdicts use a bootstrap rather than this figure, since a mean over many points is far less noisy than one measurement.
| runtime | n | median | p90 | max |
|---|---|---|---|---|
| <1 ms | 462 | 1.039 | 1.412 | 3.09 |
| 1-3 ms | 118 | 1.007 | 1.025 | 1.19 |
| 3-10 ms | 177 | 1.009 | 1.029 | 1.10 |
| 10-30 ms | 104 | 1.011 | 1.033 | 1.07 |
| 30-100 ms | 115 | 1.008 | 1.035 | 1.08 |
| >100 ms | 174 | 1.011 | 1.055 | 1.16 |