metal-linalg

Eigensolver routing on Apple M5 Pro (20 GPU cores)

What every section and number below means: reading-reports.md.

Cost model scaled to this device from one probe point: block x0.14, cpu x0.08, whole-matrix x1.00 (1.00 is an M1).

Machine state: load 2.0/18 at the start, load 2.2/18 at the end; power mains.

Probe point after the sweep relative to before it: block x0.99, cpu x1.01, tg x1.03 (stable).

Generated by tuning/tune_eigh.py from 207 (N, batch) points, N in [2, 4, 8, 12, 16, 24, 32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048, 2560, 3072, 3584, 4096], batch in [1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096], four backends, two or more passes, min-of-repeats.

Answer

Row for kTuned[] in src/eigh.mm:

// device, GPU cores,   simd_max_n, block_min_n, block_min_n_batched, block_min_batch,   gpu_max_n, gpu_min_batch_times_n, gpu_min_batch,   values_gpu_max_n, values_gpu_min_batch_times_n, values_gpu_min_batch,   tridiag_min_n, values_tridiag_min_n, tridiag_max_batch, values_tridiag_max_batch,   ql_min_n, ql_max_n,   share_min_batch,   gpu_big_batch_max_n, gpu_big_batch_min,   values_band_min_n, values_band_width
{"Apple M5 Pro", 20,   0, 96, 0, 0,   32, 16384, 1,   48, 16384, 1,   1024, 1024, 4, 2,   2, 64,   2048,   64, 2048,   2048, 16},

To try it without rebuilding:

EIGH_SIMD_MAX_N=0 EIGH_BLOCK_MIN_N=96 EIGH_BLOCK_MIN_N_BATCHED=0 EIGH_BLOCK_MIN_BATCH=0 EIGH_GPU_MAX_N=32 EIGH_GPU_MIN_BATCH_TIMES_N=16384 EIGH_GPU_MIN_BATCH=1 EIGH_TRIDIAG_MIN_N=1024 EIGH_VALUES_TRIDIAG_MIN_N=1024 EIGH_TRIDIAG_MAX_BATCH=4 EIGH_VALUES_TRIDIAG_MAX_BATCH=2 EIGH_QL_MIN_N=2 EIGH_QL_MAX_N=64 EIGH_SHARE_MIN_BATCH=2048 EIGH_GPU_BIG_BATCH_MAX_N=64 EIGH_GPU_BIG_BATCH_MIN=2048 EIGH_VALUES_BAND_MIN_N=2048 EIGH_VALUES_BAND_WIDTH=16 EIGH_VALUES_GPU_MAX_N=48 EIGH_VALUES_GPU_MIN_BATCH_TIMES_N=16384 EIGH_VALUES_GPU_MIN_BATCH=1

The policy in effect on this device came from tuned-stale:Apple M5 Pro. It differs from the fitted one; see the warnings.

Against the best measured backend at every point the whole rule scores 1.0160 geometric-mean regret, worst 1.48x, 12 of 207 points losing more than 10%, and 1.001x the oracle’s total time. The decision is fitted in two stages, below, because the CPU routing would otherwise hide the GPU backend crossover.

Warnings

Stage 1: which GPU backend

Scored against the best GPU backend at each of the 207 points, as if there were no CPU: this is the rule a forced-GPU call (EIGH_DEVICE=gpu) and the detail entry points follow, and it is what a GPU with more cores will lean on.

rule geomean regret worst >10% total time / oracle est. picks
policy in effect (‘0’, ‘96’, ‘none’, ‘none’) 1.0240 1.59x 14 1.003 0
fitted (‘0’, ‘96’, ‘none’, ‘none’) 1.0240 1.59x 14 1.003 0

4 of 144 (simd_max_n, block_min_n) pairs are within 0.5% of the best geomean: simd_max_n 0 .. 8, block_min_n 96 .. 96.

xychart-beta
    title "Regret by block_min_n"
    x-axis "block_min_n" [32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048, 2560, 3072, 3584, 4096, none]
    y-axis "geometric-mean regret" 1.0 --> 2.67
    line [1.1999, 1.1051, 1.0416, 1.0240, 1.0480, 1.1198, 1.2458, 1.3874, 1.5577, 1.7313, 1.9439, 2.1286, 2.2673, 2.4141, 2.4726, 2.5322, 2.5933, 2.6559]
block_min_n 32 48 64 96 128 192 256 384 512 768 1024 1536 2048 2560 3072 3584 4096 none
geomean 1.1999 1.1051 1.0416 1.0240 1.0480 1.1198 1.2458 1.3874 1.5577 1.7313 1.9439 2.1286 2.2673 2.4141 2.4726 2.5322 2.5933 2.6559
worst 9.31x 4.72x 2.74x 1.59x 1.93x 2.57x 4.60x 6.13x 10.16x 15.82x 15.82x 15.82x 15.82x 15.82x 15.82x 15.82x 15.82x 15.82x
xychart-beta
    title "Regret by simd_max_n"
    x-axis "simd_max_n" [0, 2, 4, 8, 12, 16, 24, 32]
    y-axis "geometric-mean regret" 1.0 --> 1.21
    line [1.0240, 1.0243, 1.0290, 1.0265, 1.0570, 1.0809, 1.1237, 1.1909]
simd_max_n 0 2 4 8 12 16 24 32
geomean 1.0240 1.0243 1.0290 1.0265 1.0570 1.0809 1.1237 1.1909
worst 1.59x 1.59x 2.27x 2.27x 3.21x 3.53x 3.53x 4.68x

Held-out check of a batch-dependent crossover (block from a lower N once the batch is large enough). Fitted on 114 points, scored on the other 93; the verdict is a bootstrap over the test points.

rule fitted on train train geomean test geomean test worst verdict
two constants [0, 96, 1000000000, 1000000000] 1.0209 1.0277 1.59x baseline
batch-dependent block crossover {“block_lo”: 64, “batch_hi”: 32} 1.0104 1.0100 1.86x rejected (better in 94% of resamples, median gain 1.8%)

Best GPU backend per point (s simd, t threadgroup, B block, q ql, Q ql shared with the CPU), then what the split picks, the ql window of stage 1b included:

  N \ batch     1     2     4     8    16    32    64   128   256   512  1024  2048  4096
          2     q     q     q     s     q     s     q     s     s     t     s     q     q
          4     q     t     q     q     s     q     q     q     q     q     s     q     q
          8     q     q     q     s     s     q     s     s     s     q     s     t     s
         12     t     t     t     t     t     t     t     t     q     q     q     q     q
         16     t     t     t     t     t     t     t     t     q     q     q     q     q
         24     t     t     t     t     t     t     t     q     q     q     q     q     q
         32     t     t     t     t     t     t     q     q     q     q     q     q     q
         48     t     t     t     t     t     q     q     q     q     q     q     q     q
         64     t     t     t     t     t     q     q     q     q     q     q     q     q
         96     t     t     t     t     t     B     B     B     B     B     B     B     B
        128     B     B     B     B     B     B     B     B     B     B     B     B     B
        192     B     B     B     B     B     B     B     B     B     B     B     B     B
        256     B     B     B     B     B     B     B     B     B     B     B     .     .
        384     B     B     B     B     B     B     B     B     B     B     .     .     .
        512     B     B     B     B     B     B     B     B     .     .     .     .     .
        768     B     B     B     B     B     B     B     .     .     .     .     .     .
       1024     B     B     B     B     B     .     .     .     .     .     .     .     .
       1536     B     B     B     .     .     .     .     .     .     .     .     .     .
       2048     B     B     B     .     .     .     .     .     .     .     .     .     .
       2560     B     .     .     .     .     .     .     .     .     .     .     .     .
       3072     B     .     .     .     .     .     .     .     .     .     .     .     .
       3584     B     .     .     .     .     .     .     .     .     .     .     .     .
       4096     B     .     .     .     .     .     .     .     .     .     .     .     .
  N \ batch     1     2     4     8    16    32    64   128   256   512  1024  2048  4096
          2     q     q     q     q     q     q     q     q     q     q     q     Q     Q
          4     q     q     q     q     q     q     q     q     q     q     q     Q     Q
          8     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         12     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         16     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         24     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         32     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         48     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         64     q     q     q     q     q     q     q     q     q     q     q     Q     Q
         96     B     B     B     B     B     B     B     B     B     B     B     B     B
        128     B     B     B     B     B     B     B     B     B     B     B     B     B
        192     B     B     B     B     B     B     B     B     B     B     B     B     B
        256     B     B     B     B     B     B     B     B     B     B     B     .     .
        384     B     B     B     B     B     B     B     B     B     B     .     .     .
        512     B     B     B     B     B     B     B     B     .     .     .     .     .
        768     B     B     B     B     B     B     B     .     .     .     .     .     .
       1024     B     B     B     B     B     .     .     .     .     .     .     .     .
       1536     B     B     B     .     .     .     .     .     .     .     .     .     .
       2048     B     B     B     .     .     .     .     .     .     .     .     .     .
       2560     B     .     .     .     .     .     .     .     .     .     .     .     .
       3072     B     .     .     .     .     .     .     .     .     .     .     .     .
       3584     B     .     .     .     .     .     .     .     .     .     .     .     .
       4096     B     .     .     .     .     .     .     .     .     .     .     .     .

Stage 1b: the ql backend

Inside a window of N, ql (tridiagonalization and implicit QL, one threadgroup per matrix, N <= 87 on this device) instead of the Jacobi backend the split picks, fitted over the 207 points against the best GPU backend, ql included. Chosen: N = 2 .. 64.

rule geomean regret worst >10% total time / oracle est. picks
without ql (the split alone) 1.1759 3.48x 56 1.009 0
with ql for N in (2, 64) 1.0454 1.63x 35 1.000 0

2 windows are within 0.5% of the best geomean: ql_min_n 2 .. 4, ql_max_n 64 .. 64.

Held out: fitted on 114 points (window (2, 64)), scored on the other 93: geomean 1.0542x, worst 1.47x, against 1.1710x, worst 3.48x without ql.

ql over the best Jacobi backend, N x batch: 2x1 1.02x, 2x2 1.09x, 2x4 1.28x, 2x8 0.95x, 2x16 1.15x, 2x32 0.98x, 2x64 1.01x, 2x128 0.94x, 2x256 0.88x, 2x512 0.90x, 2x1024 0.97x, 2x2048 1.01x, 2x4096 1.01x, 4x1 1.10x, 4x2 0.97x, 4x4 1.12x, 4x8 1.43x, 4x16 0.90x, 4x32 1.31x, 4x64 1.07x, 4x128 1.51x, 4x256 1.00x, 4x512 1.55x, 4x1024 0.98x, 4x2048 1.29x, 4x4096 1.04x, 8x1 1.05x, 8x2 1.30x, 8x4 1.42x, 8x8 0.94x, 8x16 0.97x, 8x32 1.69x, 8x64 0.97x, 8x128 0.94x, 8x256 0.98x, 8x512 1.00x, 8x1024 0.97x, 8x2048 0.98x, 8x4096 0.95x, 12x1 0.83x, 12x2 0.83x, 12x4 0.91x, 12x8 0.95x, 12x16 0.98x, 12x32 0.95x, 12x64 0.99x, 12x128 0.97x, 12x256 1.17x, 12x512 1.20x, 12x1024 1.26x, 12x2048 1.41x, 12x4096 1.37x, 16x1 0.86x, 16x2 0.91x, 16x4 0.84x, 16x8 0.93x, 16x16 0.89x, 16x32 0.90x, 16x64 0.90x, 16x128 0.97x, 16x256 1.24x, 16x512 1.27x, 16x1024 1.49x, 16x2048 1.37x, 16x4096 1.46x, 24x1 0.74x, 24x2 0.75x, 24x4 0.73x, 24x8 0.77x, 24x16 0.73x, 24x32 0.76x, 24x64 0.92x, 24x128 1.22x, 24x256 1.79x, 24x512 1.94x, 24x1024 1.99x, 24x2048 2.18x, 24x4096 2.33x, 32x1 0.61x, 32x2 0.68x, 32x4 0.69x, 32x8 0.71x, 32x16 0.71x, 32x32 0.81x, 32x64 1.07x, 32x128 1.72x, 32x256 2.57x, 32x512 2.22x, 32x1024 2.69x, 32x2048 2.92x, 32x4096 3.07x, 48x1 0.83x, 48x2 0.85x, 48x4 0.83x, 48x8 0.75x, 48x16 0.87x, 48x32 1.08x, 48x64 1.74x, 48x128 2.59x, 48x256 2.56x, 48x512 3.13x, 48x1024 3.13x, 48x2048 3.35x, 48x4096 3.48x, 64x1 0.91x, 64x2 0.92x, 64x4 0.91x, 64x8 0.93x, 64x16 0.91x, 64x32 1.30x, 64x64 2.25x, 64x128 2.09x, 64x256 2.07x, 64x512 2.05x, 64x1024 2.13x, 64x2048 2.17x, 64x4096 2.14x

Stage 1c: sharing a batch with the CPU

From a batch on, ql_share: ql and the CPU path at once on one batch, the GPU taking chunks from the front and the CPU from the back. Fitted against the best GPU backend, the shared one included, on the 63 points where it was timed and ql is the GPU’s choice. Chosen: from batch 2048.

rule geomean regret worst >10% total time / oracle est. picks
ql alone 1.1150 1.91x 22 1.479 0
shared from batch 2048 1.1002 1.71x 19 1.102 0

ql_share over ql alone, N x batch: 2x64 1.52x, 2x128 1.10x, 2x256 0.63x, 2x512 0.60x, 2x1024 0.59x, 2x2048 0.61x, 2x4096 0.61x, 4x64 1.21x, 4x128 0.97x, 4x256 0.90x, 4x512 0.60x, 4x1024 0.64x, 4x2048 0.71x, 4x4096 0.74x, 8x64 1.20x, 8x128 0.81x, 8x256 0.64x, 8x512 0.66x, 8x1024 0.77x, 8x2048 0.72x, 8x4096 0.94x, 12x64 0.69x, 12x128 0.71x, 12x256 0.65x, 12x512 0.74x, 12x1024 0.69x, 12x2048 0.94x, 12x4096 1.07x, 16x64 0.75x, 16x128 0.77x, 16x256 0.70x, 16x512 0.82x, 16x1024 0.78x, 16x2048 1.14x, 16x4096 1.13x, 24x64 0.80x, 24x128 0.81x, 24x256 0.81x, 24x512 0.68x, 24x1024 1.09x, 24x2048 1.12x, 24x4096 1.22x, 32x64 0.84x, 32x128 0.85x, 32x256 0.85x, 32x512 0.85x, 32x1024 1.12x, 32x2048 1.10x, 32x4096 1.19x, 48x64 0.88x, 48x128 0.89x, 48x256 1.58x, 48x512 1.25x, 48x1024 1.22x, 48x2048 1.46x, 48x4096 1.46x, 64x64 0.92x, 64x128 0.96x, 64x256 1.36x, 64x512 1.50x, 64x1024 1.71x, 64x2048 1.87x, 64x4096 1.91x

Stage 2: GPU or CPU

Given the split above, GPU iff N <= gpu_max_n, batch * N >= gpu_min_batch_times_n and batch >= gpu_min_batch, scored against the best of all four backends. worst is over the points where the chosen backend was timed; a pick the cost model had to guess is listed in the warnings instead.

rule geomean regret worst >10% total time / oracle est. picks
oracle (best per point) 1.0000 1.00x 0 1.000 0
policy in effect (‘48’, ‘16384’, ‘1’) 1.0154 1.71x 12 1.001 0
fitted (‘32’, ‘16384’, ‘1’) 1.0160 1.48x 12 1.001 0

Large batches: the GPU also for N above gpu_max_n up to 64 in a batch of at least 2048, fitted with the product rule (per cap, the rule, the clause over it and the rule again given the clause, the best kept) (product rule alone 1.0267, worst 1.91x; chosen 1.0160, worst 1.48x).

18 of 1320 combinations are within 0.5% of the best geomean: gpu_max_n 48 .. 64, gpu_min_batch_times_n 8192 .. 16384, gpu_min_batch 1 .. 32.

xychart-beta
    title "Regret by gpu_min_batch_times_n"
    x-axis "gpu_min_batch_times_n" [0, 64, 128, 256, 512, 1024, 2048, 4096, 8192, 16384, none]
    y-axis "geometric-mean regret" 1.0 --> 1.78
    line [1.7675, 1.2981, 1.2023, 1.1350, 1.0922, 1.0699, 1.0430, 1.0274, 1.0181, 1.0160, 1.0408]
gpu_min_batch_times_n 0 64 128 256 512 1024 2048 4096 8192 16384 none
geomean 1.7675 1.2981 1.2023 1.1350 1.0922 1.0699 1.0430 1.0274 1.0181 1.0160 1.0408
worst 126.76x 12.59x 10.09x 5.57x 3.37x 3.18x 2.22x 2.13x 1.63x 1.48x 2.19x
xychart-beta
    title "Regret by gpu_min_batch"
    x-axis "gpu_min_batch" [1, 2, 4, 8, 16, 32]
    y-axis "geometric-mean regret" 1.0 --> 1.03
    line [1.0160, 1.0160, 1.0160, 1.0160, 1.0160, 1.0160]
gpu_min_batch 1 2 4 8 16 32
geomean 1.0160 1.0160 1.0160 1.0160 1.0160 1.0160
worst 1.48x 1.48x 1.48x 1.48x 1.48x 1.48x
xychart-beta
    title "Regret by gpu_max_n"
    x-axis "gpu_max_n" [16, 24, 32, 48, 64, 96, 128, 192, 256, 384, 512, 768, 1024, 1536, 2048, 2560, 3072, 3584, 4096, none]
    y-axis "geometric-mean regret" 1.0 --> 1.37
    line [1.0197, 1.0182, 1.0160, 1.0147, 1.0184, 1.0493, 1.0944, 1.1533, 1.2081, 1.2632, 1.3064, 1.3398, 1.3544, 1.3544, 1.3544, 1.3544, 1.3544, 1.3544, 1.3544, 1.3544]
gpu_max_n 16 24 32 48 64 96 128 192 256 384 512 768 1024 1536 2048 2560 3072 3584 4096 none
geomean 1.0197 1.0182 1.0160 1.0147 1.0184 1.0493 1.0944 1.1533 1.2081 1.2632 1.3064 1.3398 1.3544 1.3544 1.3544 1.3544 1.3544 1.3544 1.3544 1.3544
worst 1.60x 1.60x 1.48x 1.48x 1.71x 3.70x 4.65x 6.67x 7.75x 10.64x 11.01x 14.20x 14.20x 14.20x 14.20x 14.20x 14.20x 14.20x 14.20x 14.20x

Held-out check of a per-N boundary (a lookup table of the smallest batch at which the GPU wins, per N) against the product rule. Fitted on 114 points, scored on the other 93.

rule fitted on train train geomean test geomean test worst verdict
product rule [48, 16384, 1] 1.0180 1.0106 1.43x baseline
per-N table {“min_batch_by_n”: {“2”: null, “4”: null, “8”: null, “12”: 64, “16”: 512, “24”: 256, “32”: 256, “48”: 512, “64”: null, “96”: null, “128”: null, “192”: null, “256”: null, “384”: null, “512”: null, “768”: null, “1024”: null, “1536”: null, “2048”: null, “2560”: null, “3072”: null, “3584”: null, “4096”: null}} 1.0134 1.0278 1.76x rejected (better in 3% of resamples, median gain -1.6%)

Best backend per point (c CPU, s simd, t threadgroup, B block, q ql, Q ql shared with the CPU, . not measured), what the whole rule picks, and the speedup of the best GPU backend over the CPU:

  N \ batch     1     2     4     8    16    32    64   128   256   512  1024  2048  4096
          2     c     c     c     c     c     c     c     c     c     c     c     c     q
          4     c     c     c     c     c     c     c     c     c     c     c     c     q
          8     c     c     c     c     c     c     c     c     c     c     s     t     s
         12     c     c     c     c     c     c     t     c     c     c     q     q     Q
         16     c     c     c     c     c     c     c     c     c     q     q     Q     Q
         24     c     c     c     c     c     c     c     c     q     q     Q     Q     Q
         32     c     c     c     c     c     c     c     c     q     q     Q     Q     Q
         48     c     c     c     c     c     c     c     c     Q     Q     Q     Q     Q
         64     c     c     c     c     c     c     c     c     c     Q     Q     Q     Q
         96     c     c     c     c     c     c     c     c     c     c     c     c     c
        128     c     c     c     c     c     c     c     c     c     c     c     c     c
        192     c     c     c     c     c     c     c     c     c     c     c     c     c
        256     c     c     c     c     c     c     c     c     c     c     c     .     .
        384     c     c     c     c     c     c     c     c     c     c     .     .     .
        512     c     c     c     c     c     c     c     c     .     .     .     .     .
        768     c     c     c     c     c     c     c     .     .     .     .     .     .
       1024     c     c     c     c     c     .     .     .     .     .     .     .     .
       1536     c     c     c     .     .     .     .     .     .     .     .     .     .
       2048     c     c     c     .     .     .     .     .     .     .     .     .     .
       2560     c     .     .     .     .     .     .     .     .     .     .     .     .
       3072     c     .     .     .     .     .     .     .     .     .     .     .     .
       3584     c     .     .     .     .     .     .     .     .     .     .     .     .
       4096     c     .     .     .     .     .     .     .     .     .     .     .     .
  N \ batch     1     2     4     8    16    32    64   128   256   512  1024  2048  4096
          2     c     c     c     c     c     c     c     c     c     c     c     c     c
          4     c     c     c     c     c     c     c     c     c     c     c     c     Q
          8     c     c     c     c     c     c     c     c     c     c     c     Q     Q
         12     c     c     c     c     c     c     c     c     c     c     c     Q     Q
         16     c     c     c     c     c     c     c     c     c     c     q     Q     Q
         24     c     c     c     c     c     c     c     c     c     c     q     Q     Q
         32     c     c     c     c     c     c     c     c     c     q     q     Q     Q
         48     c     c     c     c     c     c     c     c     c     c     c     Q     Q
         64     c     c     c     c     c     c     c     c     c     c     c     Q     Q
         96     c     c     c     c     c     c     c     c     c     c     c     c     c
        128     c     c     c     c     c     c     c     c     c     c     c     c     c
        192     c     c     c     c     c     c     c     c     c     c     c     c     c
        256     c     c     c     c     c     c     c     c     c     c     c     .     .
        384     c     c     c     c     c     c     c     c     c     c     .     .     .
        512     c     c     c     c     c     c     c     c     .     .     .     .     .
        768     c     c     c     c     c     c     c     .     .     .     .     .     .
       1024     c     c     c     c     c     .     .     .     .     .     .     .     .
       1536     c     c     c     .     .     .     .     .     .     .     .     .     .
       2048     c     c     c     .     .     .     .     .     .     .     .     .     .
       2560     c     .     .     .     .     .     .     .     .     .     .     .     .
       3072     c     .     .     .     .     .     .     .     .     .     .     .     .
       3584     c     .     .     .     .     .     .     .     .     .     .     .     .
       4096     c     .     .     .     .     .     .     .     .     .     .     .     .
  N \ batch     1     2     4     8    16    32    64   128   256   512  1024  2048  4096
          2  0.01  0.02  0.04  0.04  0.06  0.10  0.16  0.37  0.64  0.62  0.76  0.77  1.02
          4  0.01  0.03  0.05  0.07  0.11  0.17  0.40  0.74  0.46  0.66  0.76  0.94  1.24
          8  0.02  0.04  0.09  0.13  0.21  0.46  0.83  0.53  0.66  0.77  1.00  1.12  1.36
         12  0.03  0.08  0.11  0.15  0.36  0.36  1.11  0.58  0.75  0.94  1.14  1.20  1.47
         16  0.04  0.08  0.11  0.16  0.33  0.39  0.53  0.61  0.83  1.05  1.22  1.27  1.46
         24  0.09  0.11  0.13  0.28  0.40  0.48  0.52  0.69  1.04  1.28  1.34  1.67  1.80
         32  0.11  0.12  0.14  0.25  0.42  0.39  0.45  0.69  1.07  1.10  1.43  1.68  1.74
         48  0.11  0.13  0.14  0.29  0.28  0.34  0.50  0.82  0.88  1.13  1.17  1.31  1.29
         64  0.12  0.13  0.13  0.22  0.23  0.33  0.49  0.51  0.67  0.80  0.87  0.83  0.82
         96  0.08  0.09  0.09  0.11  0.14  0.16  0.21  0.26  0.31  0.30  0.29  0.28  0.27
        128  0.08  0.09  0.10  0.11  0.12  0.18  0.22  0.26  0.25  0.24  0.23  0.22  0.21
        192  0.13  0.13  0.13  0.15  0.16  0.19  0.20  0.19  0.17  0.17  0.16  0.15  0.15
        256  0.19  0.18  0.18  0.18  0.17  0.19  0.18  0.15  0.14  0.14  0.13     .     .
        384  0.22  0.20  0.20  0.19  0.15  0.13  0.11  0.10  0.10  0.09     .     .     .
        512  0.29  0.26  0.25  0.19  0.13  0.11  0.10  0.09     .     .     .     .     .
        768  0.30  0.28  0.21  0.13  0.08  0.08  0.07     .     .     .     .     .     .
       1024  0.40  0.33  0.20  0.12  0.10     .     .     .     .     .     .     .     .
       1536  0.40  0.26  0.12     .     .     .     .     .     .     .     .     .     .
       2048  0.42  0.25  0.17     .     .     .     .     .     .     .     .     .     .
       2560  0.31     .     .     .     .     .     .     .     .     .     .     .     .
       3072  0.32     .     .     .     .     .     .     .     .     .     .     .     .
       3584  0.34     .     .     .     .     .     .     .     .     .     .     .     .
       4096  0.47     .     .     .     .     .     .     .     .     .     .     .     .

Stage 3: GPU or CPU, eigenvalues alone

The same rule for eigvalsh, with its own thresholds (values_gpu_max_n, values_gpu_min_batch_times_n, values_gpu_min_batch), fitted on the _vals timings of 207 points given the split above. The CPU computes eigenvalues alone by LAPACK’s two-stage reduction from N = 128, so the boundary need not be eigh’s.

rule geomean regret worst >10% total time / oracle est. picks
policy in effect 1.0483 3.77x 22 1.414 0
eigh’s fitted boundary 1.0447 3.77x 20 1.411 0
fitted (‘48’, ‘16384’, ‘1’) 1.0483 3.77x 22 1.414 0

12 combinations are within 0.5% of the best geomean: values_gpu_max_n 32 .. 48, values_gpu_min_batch_times_n 16384 .. 16384, values_gpu_min_batch 1 .. 32.

Held out: fitted on 114 points (‘48’, ‘16384’, ‘1’), scored on the other 93: geomean 1.0280x, worst 1.64x, against 1.0309x, worst 1.64x for eigh’s boundary on the same points.

Stage 4: the tridiag backend instead of the CPU

Where the rule above chooses the CPU, the tridiag backend from a threshold N on (0: never), for batches up to a cap (0: any; it solves a batch one matrix after another, the CPU path spreads one over every core), fitted over the measured N and batches against the best of all backends, tridiag included, on the points where tridiag was timed (N >= 128, within the cost cap): the region the threshold decides.

  threshold batch cap geomean regret worst without tridiag: geomean worst held out (fitted on half)
with eigenvectors 1024 4 1.0120 1.51x 1.2534 9.01x from 768, batch <= 4: 1.0191 vs 1.0962
eigenvalues alone 1024 2 1.0173 1.55x 1.1177 3.77x from 1536, batch <= 2: 1.0171 vs 1.0482

tridiag over the CPU (with eigenvectors), N x batch: 128x1 0.30x, 128x2 0.20x, 128x4 0.10x, 128x8 0.06x, 128x16 0.04x, 128x32 0.03x, 128x64 0.03x, 128x128 0.03x, 128x256 0.03x, 128x512 0.03x, 192x1 0.48x, 192x2 0.29x, 192x4 0.15x, 192x8 0.09x, 192x16 0.06x, 192x32 0.05x, 192x64 0.05x, 192x128 0.05x, 192x256 0.04x, 256x1 0.69x, 256x2 0.39x, 256x4 0.22x, 256x8 0.13x, 256x16 0.08x, 256x32 0.08x, 256x64 0.07x, 256x128 0.07x, 256x256 0.07x, 384x1 0.87x, 384x2 0.53x, 384x4 0.31x, 384x8 0.19x, 384x16 0.12x, 384x32 0.11x, 384x64 0.10x, 384x128 0.10x, 512x1 1.16x, 512x2 0.80x, 512x4 0.49x, 512x8 0.27x, 512x16 0.18x, 512x32 0.17x, 512x64 0.16x, 512x128 0.16x, 768x1 1.51x, 768x2 1.13x, 768x4 0.59x, 768x8 0.38x, 768x16 0.28x, 768x32 0.27x, 768x64 0.26x, 1024x1 2.27x, 1024x2 1.65x, 1024x4 0.90x, 1024x8 0.64x, 1024x16 0.57x, 1536x1 3.16x, 1536x2 2.14x, 1536x4 1.17x, 2048x1 4.37x, 2048x2 3.34x, 2048x4 2.32x, 2560x1 4.53x, 3072x1 5.67x, 3584x1 6.52x, 4096x1 9.01x

tridiag over the CPU (eigenvalues alone), N x batch: 128x1 0.24x, 128x2 0.14x, 128x4 0.07x, 128x8 0.05x, 128x16 0.03x, 128x32 0.02x, 128x64 0.02x, 128x128 0.02x, 128x256 0.02x, 128x512 0.02x, 192x1 0.34x, 192x2 0.19x, 192x4 0.10x, 192x8 0.06x, 192x16 0.05x, 192x32 0.03x, 192x64 0.04x, 192x128 0.03x, 192x256 0.03x, 256x1 0.44x, 256x2 0.24x, 256x4 0.14x, 256x8 0.08x, 256x16 0.05x, 256x32 0.05x, 256x64 0.04x, 256x128 0.04x, 256x256 0.04x, 384x1 0.60x, 384x2 0.34x, 384x4 0.20x, 384x8 0.12x, 384x16 0.08x, 384x32 0.07x, 384x64 0.07x, 384x128 0.06x, 512x1 0.84x, 512x2 0.48x, 512x4 0.28x, 512x8 0.17x, 512x16 0.11x, 512x32 0.10x, 512x64 0.09x, 512x128 0.09x, 768x1 1.25x, 768x2 0.70x, 768x4 0.40x, 768x8 0.25x, 768x16 0.17x, 768x32 0.15x, 768x64 0.13x, 1024x1 1.61x, 1024x2 0.92x, 1024x4 0.55x, 1024x8 0.32x, 1024x16 0.21x, 1536x1 2.04x, 1536x2 1.37x, 1536x4 0.72x, 2048x1 2.33x, 2048x2 1.58x, 2048x4 0.82x, 2560x1 2.38x, 3072x1 2.42x, 3584x1 2.37x, 4096x1 2.44x

Stage 4b: the band backend for eigenvalues alone

For eigenvalues alone, where the rules above choose the CPU or tridiag, the band backend (the two-stage reduction: A to a band on the GPU in blocks whose work is matrix products, the band to tridiagonal on the CPU’s cores, then bisection on the GPU) from a threshold N on (0: never), within tridiag’s batch cap, fitted on the 30 points where it was timed (N >= 512) against the CPU, tridiag and band.

Chosen: 2048 (in effect: 3072): 1.0172 geometric-mean regret, worst 1.25x; without band 1.0587, worst 1.55x.

Its band’s width: 16 (in effect: 16); geometric mean of each width’s time over the best width’s at each point: 8 1.148, 16 1.001, 32 1.317. 16, the default, unless another is better by more than 1%.

Thresholds within the fit’s tolerance of the best, and within 3% of it on the points where the two choose differently: 1536, 2048.

band over tridiag, N x batch: 512x1 0.93x, 512x2 1.03x, 512x4 1.13x, 512x8 1.20x, 512x16 1.22x, 512x32 1.24x, 512x64 1.25x, 512x128 1.26x, 768x1 0.90x, 768x2 0.99x, 768x4 1.10x, 768x8 1.16x, 768x16 1.18x, 768x32 1.20x, 768x64 1.22x, 1024x1 0.87x, 1024x2 0.99x, 1024x4 1.12x, 1024x8 1.18x, 1024x16 1.22x, 1536x1 0.99x, 1536x2 1.05x, 1536x4 1.13x, 2048x1 0.95x, 2048x2 1.17x, 2048x4 1.32x, 2560x1 1.11x, 3072x1 1.29x, 3584x1 1.38x, 4096x1 1.51x

band over the CPU, N x batch: 512x1 0.78x, 512x2 0.49x, 512x4 0.31x, 512x8 0.20x, 512x16 0.14x, 512x32 0.13x, 512x64 0.11x, 512x128 0.11x, 768x1 1.12x, 768x2 0.70x, 768x4 0.44x, 768x8 0.29x, 768x16 0.21x, 768x32 0.18x, 768x64 0.16x, 1024x1 1.39x, 1024x2 0.91x, 1024x4 0.62x, 1024x8 0.38x, 1024x16 0.25x, 1536x1 2.03x, 1536x2 1.44x, 1536x4 0.81x, 2048x1 2.22x, 2048x2 1.84x, 2048x4 1.08x, 2560x1 2.65x, 3072x1 3.11x, 3584x1 3.27x, 4096x1 3.69x

Noise floor

Pass-to-pass ratio (max/min of the same measurement across passes), 1766 measurements: median 1.008, p90 1.099, max 3.03. The held-out verdicts use a bootstrap rather than this figure, since a mean over many points is far less noisy than one measurement.

runtime n median p90 max
<1 ms 811 1.026 1.212 3.03
1-3 ms 180 1.005 1.041 1.61
3-10 ms 215 1.005 1.030 1.33
10-30 ms 168 1.004 1.020 1.08
30-100 ms 176 1.003 1.020 1.11
>100 ms 216 1.002 1.011 1.12