What every section and number below means: reading-reports.md.
Cost model scaled to this device from one probe point: bidiag x1.00, block x0.25, cpu x0.08, gk x0.40, jacobi x0.10, qr x0.18, qrblock x0.23 (1.00 is an M1).
Machine state: load 4.8/18 at the start, load 4.8/18 at the end; power mains.
Probe point after the sweep relative to before it: block x0.99, cpu x1.00, gk x1.00, jacobi x0.98, qr x0.99, qrblock x1.01 (stable).
Generated by tuning/tune_svd.py from 291 (shape, batch) points, 45 shapes with M >= N, batch in [1, 4, 16, 64, 256, 1024, 4096], seven backends (and the CPU and bidiag again for singular values alone), two or more passes, min-of-repeats.
Row for kTuned[] in src/svd.mm:
// device, GPU cores, qr_min_rows, qr_min_k, block_min_k, block_min_k_batched, block_min_batch, gpu_max_k, gpu_min_batch_times_k, gpu_min_batch, gpu_max_l, values_gpu_max_k, values_gpu_min_batch_times_k, values_gpu_min_batch, values_gpu_max_l, bidiag_min_k, values_bidiag_min_k, bidiag_max_batch, values_bidiag_max_batch, gk_min_k, gk_max_k, share_min_batch, gpu_big_batch_max_k, gpu_big_batch_min, values_band_min_k
{"Apple M5 Pro", 20, 512, 32, 192, 64, 64, 56, 16384, 1, 256, 56, 16384, 1, 56, 1024, 1024, 2, 2, 8, 80, 1024, 80, 1024, 1536},
To try it without rebuilding:
SVD_QR_MIN_ROWS=512 SVD_QR_MIN_K=32 SVD_BLOCK_MIN_K=192 SVD_BLOCK_MIN_K_BATCHED=64 SVD_BLOCK_MIN_BATCH=64 SVD_GPU_MAX_K=56 SVD_GPU_MIN_BATCH_TIMES_K=16384 SVD_GPU_MIN_BATCH=1 SVD_GPU_MAX_L=256 SVD_BIDIAG_MIN_K=1024 SVD_VALUES_BIDIAG_MIN_K=1024 SVD_BIDIAG_MAX_BATCH=2 SVD_VALUES_BIDIAG_MAX_BATCH=2 SVD_GK_MIN_K=8 SVD_GK_MAX_K=80 SVD_SHARE_MIN_BATCH=1024 SVD_GPU_BIG_BATCH_MAX_K=80 SVD_GPU_BIG_BATCH_MIN=1024 SVD_VALUES_BAND_MIN_K=1536 SVD_VALUES_GPU_MAX_K=56 SVD_VALUES_GPU_MIN_BATCH_TIMES_K=16384 SVD_VALUES_GPU_MIN_BATCH=1 SVD_VALUES_GPU_MAX_L=56
The policy in effect on this device came from tuned:Apple M5 Pro. Against the best measured backend at every point the fitted rule scores 1.0198 geometric-mean regret, worst 1.58x, 19 of 291 points losing more than 10%, and 1.106x the oracle’s total time.
Scored against the best GPU backend at each of the 291 points, as if there were no CPU: this is the rule a forced-GPU call (SVD_DEVICE=gpu) follows, and what a GPU with more cores will lean on. Two independent choices. Precondition with QR iff the long side is at least qr_min_rows, the short side at least qr_min_k, and the long side at least twice the short one. Block kernel iff the short side is at least block_min_k, or at least block_min_k_batched in a batch of block_min_batch or more.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| policy in effect (‘512’, ‘32’, ‘192’, ‘64’, ‘64’) | 1.0193 | 1.92x | 19 | 1.009 | 0 |
| fitted (‘512’, ‘32’, ‘192’, ‘64’, ‘64’) | 1.0193 | 1.92x | 19 | 1.009 | 0 |
Without a batch term, 3 of 735 combinations are within 0.5% of the best geomean: qr_min_rows 256 .. 512, qr_min_k 32 .. 32, block_min_k 96 .. 128.
With the batch term adopted (below), 4 combinations are within 0.5% of the best: block_min_k 128 .. 256, block_min_k_batched 56 .. 64, block_min_batch 64 .. 256. The curves vary one constant around the chosen combination.
xychart-beta
title "Regret by block_min_k"
x-axis "block_min_k" [32, 40, 48, 56, 64, 80, 96, 128, 192, 256, 384, 512, 768, 1024, none]
y-axis "geometric-mean regret" 1.0 --> 1.27
line [1.2571, 1.1453, 1.1177, 1.0990, 1.0852, 1.0408, 1.0324, 1.0266, 1.0193, 1.0238, 1.0595, 1.0755, 1.0967, 1.1238, 1.1542]
| block_min_k | 32 | 40 | 48 | 56 | 64 | 80 | 96 | 128 | 192 | 256 | 384 | 512 | 768 | 1024 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.2571 | 1.1453 | 1.1177 | 1.0990 | 1.0852 | 1.0408 | 1.0324 | 1.0266 | 1.0193 | 1.0238 | 1.0595 | 1.0755 | 1.0967 | 1.1238 | 1.1542 |
| worst | 5.86x | 4.49x | 3.29x | 2.98x | 2.73x | 2.30x | 1.92x | 1.92x | 1.92x | 1.92x | 2.66x | 5.04x | 8.22x | 15.87x | 24.46x |
xychart-beta
title "Regret by block_min_k_batched"
x-axis "block_min_k_batched" [32, 40, 48, 56, 64, 80, 96, 128]
y-axis "geometric-mean regret" 1.0 --> 1.07
line [1.0484, 1.0357, 1.0262, 1.0212, 1.0193, 1.0405, 1.0427, 1.0512]
| block_min_k_batched | 32 | 40 | 48 | 56 | 64 | 80 | 96 | 128 |
|---|---|---|---|---|---|---|---|---|
| geomean | 1.0484 | 1.0357 | 1.0262 | 1.0212 | 1.0193 | 1.0405 | 1.0427 | 1.0512 |
| worst | 2.63x | 2.53x | 1.99x | 1.92x | 1.92x | 2.48x | 2.48x | 2.48x |
xychart-beta
title "Regret by block_min_batch"
x-axis "block_min_batch" [4, 16, 64, 256, 1024, 4096]
y-axis "geometric-mean regret" 1.0 --> 1.09
line [1.0616, 1.0400, 1.0193, 1.0279, 1.0487, 1.0747]
| block_min_batch | 4 | 16 | 64 | 256 | 1024 | 4096 |
|---|---|---|---|---|---|---|
| geomean | 1.0616 | 1.0400 | 1.0193 | 1.0279 | 1.0487 | 1.0747 |
| worst | 2.73x | 2.51x | 1.92x | 2.35x | 2.47x | 2.55x |
xychart-beta
title "Regret by qr_min_rows"
x-axis "qr_min_rows" [16, 32, 64, 128, 256, 512, 1024, 2048, none]
y-axis "geometric-mean regret" 1.0 --> 1.20
line [1.0283, 1.0283, 1.0283, 1.0252, 1.0176, 1.0193, 1.0490, 1.1108, 1.1848]
| qr_min_rows | 16 | 32 | 64 | 128 | 256 | 512 | 1024 | 2048 | none |
|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.0283 | 1.0283 | 1.0283 | 1.0252 | 1.0176 | 1.0193 | 1.0490 | 1.1108 | 1.1848 |
| worst | 1.53x | 1.53x | 1.53x | 1.92x | 1.92x | 1.92x | 2.56x | 4.25x | 6.54x |
xychart-beta
title "Regret by qr_min_k"
x-axis "qr_min_k" [8, 16, 32, 64, 128, 256]
y-axis "geometric-mean regret" 1.0 --> 1.16
line [1.0213, 1.0213, 1.0193, 1.0397, 1.0747, 1.1481]
| qr_min_k | 8 | 16 | 32 | 64 | 128 | 256 |
|---|---|---|---|---|---|---|
| geomean | 1.0213 | 1.0213 | 1.0193 | 1.0397 | 1.0747 | 1.1481 |
| worst | 1.92x | 1.92x | 1.92x | 3.67x | 4.51x | 6.54x |
Held-out check of a batch-dependent block crossover (block from a smaller k once the batch is large enough), fitted on 163 points and scored on the other 128; the verdict is a bootstrap over the test points.
| rule | fitted on train | train geomean | test geomean | test worst | verdict |
|---|---|---|---|---|---|
| one block crossover | [“256”, “32”, “96”, “none”, “none”] | 1.0407 | 1.0642 | 2.20x | baseline |
| batch-dependent block crossover | {“block_min”: “192”, “block_lo”: “64”, “batch_hi”: “64”} | 1.0166 | 1.0188 | 1.92x | justified (better in 100% of resamples, median gain 4.3%) |
Best GPU backend per point (J whole-matrix kernel, B block kernel, j and b the same after QR, g gk, G gk shared with the CPU), then what the rule picks, the gk window included:
M x N \ batch 1 4 16 64 256 1024 4096
4 x 4 g J g g J g g
8 x 8 J g J g g g g
16 x 8 g g J g g g g
32 x 8 J g g g g g g
64 x 8 J g g g g g g
128 x 8 g J J g g g g
256 x 8 g J g g g J J
16 x 16 g J J J g g g
32 x 16 J g J g g g g
64 x 16 J J g J g g g
128 x 16 g g J g g g g
256 x 16 J g J g g g g
512 x 16 J J J J g g g
24 x 24 g J g g g g g
32 x 32 J J J g g g g
64 x 32 J J j g g g g
128 x 32 J J J g g g g
256 x 32 j J J J g g g
512 x 32 J j J j g g g
1024 x 32 j j J j g g .
40 x 40 J J J g g g g
48 x 48 J J J g g g g
56 x 56 J J J g g g g
64 x 64 J J J g g g g
128 x 64 J J J g B g g
256 x 64 j J J g g g g
512 x 64 j j j g g g .
1024 x 64 j j j g g g .
2048 x 64 j j j j g . b
80 x 80 J J J g g g g
96 x 96 J J J B B B B
128 x 128 J J J B B B B
256 x 128 j j j B b b b
512 x 128 j j j b b b b
1024 x 128 j j j b b b b
2048 x 128 j j j b b b .
192 x 192 B B B B B B .
256 x 256 B B B B B B .
512 x 256 B B b b b b .
1024 x 256 b b b b b b .
2048 x 256 b b b b b . .
384 x 384 B B B B B . .
512 x 512 B B B B . . .
768 x 768 B B B . . . .
1024 x 1024 B B B . . . .
M x N \ batch 1 4 16 64 256 1024 4096
4 x 4 J J J J J J J
8 x 8 g g g g g G G
16 x 8 g g g g g G G
32 x 8 g g g g g G G
64 x 8 g g g g g G G
128 x 8 g g g g g G G
256 x 8 g g g g g G G
16 x 16 g g g g g G G
32 x 16 g g g g g G G
64 x 16 g g g g g G G
128 x 16 g g g g g G G
256 x 16 g g g g g G G
512 x 16 g g g g g G G
24 x 24 g g g g g G G
32 x 32 g g g g g G G
64 x 32 g g g g g G G
128 x 32 g g g g g G G
256 x 32 g g g g g G G
512 x 32 g g g g g G G
1024 x 32 g g g g g G .
40 x 40 g g g g g G G
48 x 48 g g g g g G G
56 x 56 g g g g g G G
64 x 64 g g g g g G G
128 x 64 g g g g g G G
256 x 64 g g g g g G G
512 x 64 g g g g g G .
1024 x 64 g g g g g G .
2048 x 64 g g g g g . G
80 x 80 g g g g g G G
96 x 96 J J J B B B B
128 x 128 J J J B B B B
256 x 128 J J J B B B B
512 x 128 j j j b b b b
1024 x 128 j j j b b b b
2048 x 128 j j j b b b .
192 x 192 B B B B B B .
256 x 256 B B B B B B .
512 x 256 b b b b b b .
1024 x 256 b b b b b b .
2048 x 256 b b b b b . .
384 x 384 B B B B B . .
512 x 512 B B B B . . .
768 x 768 B B B . . . .
1024 x 1024 B B B . . . .
Inside a window of k = min(M, N), gk (Householder bidiagonalization and implicit QR, one threadgroup per matrix, k <= 83 on this device; on the matrix itself where it fits, else after a QR) instead of the Jacobi backend the split picks, fitted over the 291 points against the best GPU backend, gk included. Chosen: k = 4 .. 56.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| without gk (the split alone) | 1.1860 | 2.96x | 105 | 1.040 | 0 |
| with gk for k in (4, 56) | 1.0947 | 2.39x | 62 | 1.022 | 0 |
Refitted once stage 1c chose to share batches with the CPU (gk shared leads the Jacobi backends where gk alone did not): k = 4 .. 56 became k = 8 .. 80, the window in the row.
2 windows are within 0.5% of the best geomean: gk_min_k 4 .. 8, gk_max_k 56 .. 56.
Held out: fitted on 163 points (window (4, 80)), scored on the other 128: geomean 1.1451x, worst 2.39x, against 1.1457x, worst 2.93x without gk.
gk over the best Jacobi backend, M x N x batch: 4x4x1 1.19x, 4x4x4 0.94x, 4x4x16 1.08x, 4x4x64 1.27x, 4x4x256 0.92x, 4x4x1024 1.10x, 4x4x4096 1.13x, 8x8x1 0.75x, 8x8x4 1.36x, 8x8x16 0.91x, 8x8x64 1.01x, 8x8x256 1.15x, 8x8x1024 1.40x, 8x8x4096 1.62x, 16x8x1 1.13x, 16x8x4 1.48x, 16x8x16 1.00x, 16x8x64 1.61x, 16x8x256 1.41x, 16x8x1024 1.38x, 16x8x4096 1.39x, 32x8x1 0.95x, 32x8x4 1.56x, 32x8x16 1.52x, 32x8x64 1.60x, 32x8x256 1.46x, 32x8x1024 1.18x, 32x8x4096 1.19x, 64x8x1 0.88x, 64x8x4 1.54x, 64x8x16 1.59x, 64x8x64 1.68x, 64x8x256 2.36x, 64x8x1024 1.09x, 64x8x4096 1.10x, 128x8x1 1.55x, 128x8x4 0.91x, 128x8x16 1.00x, 128x8x64 1.77x, 128x8x256 1.03x, 128x8x1024 1.04x, 128x8x4096 1.01x, 256x8x1 1.48x, 256x8x4 0.96x, 256x8x16 1.88x, 256x8x64 1.85x, 256x8x256 1.00x, 256x8x1024 0.96x, 256x8x4096 0.94x, 16x16x1 1.53x, 16x16x4 0.71x, 16x16x16 0.90x, 16x16x64 0.93x, 16x16x256 1.68x, 16x16x1024 2.02x, 16x16x4096 2.36x, 32x16x1 0.90x, 32x16x4 1.35x, 32x16x16 0.72x, 32x16x64 1.64x, 32x16x256 1.48x, 32x16x1024 1.61x, 32x16x4096 1.73x, 64x16x1 0.64x, 64x16x4 0.90x, 64x16x16 1.17x, 64x16x64 0.98x, 64x16x256 1.60x, 64x16x1024 1.46x, 64x16x4096 1.51x, 128x16x1 1.15x, 128x16x4 1.13x, 128x16x16 0.77x, 128x16x64 1.08x, 128x16x256 1.35x, 128x16x1024 1.32x, 128x16x4096 1.28x, 256x16x1 0.80x, 256x16x4 1.26x, 256x16x16 0.83x, 256x16x64 1.05x, 256x16x256 1.14x, 256x16x1024 1.19x, 256x16x4096 1.22x, 512x16x1 0.65x, 512x16x4 0.63x, 512x16x16 0.51x, 512x16x64 0.76x, 512x16x256 1.08x, 512x16x1024 1.05x, 512x16x4096 1.07x, 24x24x1 1.02x, 24x24x4 0.93x, 24x24x16 1.10x, 24x24x64 1.02x, 24x24x256 2.15x, 24x24x1024 2.23x, 24x24x4096 2.47x, 32x32x1 0.50x, 32x32x4 0.54x, 32x32x16 0.57x, 32x32x64 1.08x, 32x32x256 2.39x, 32x32x1024 2.62x, 32x32x4096 2.41x, 64x32x1 0.52x, 64x32x4 0.51x, 64x32x16 0.79x, 64x32x64 1.02x, 64x32x256 1.58x, 64x32x1024 1.71x, 64x32x4096 1.65x, 128x32x1 0.51x, 128x32x4 0.66x, 128x32x16 0.56x, 128x32x64 1.07x, 128x32x256 1.23x, 128x32x1024 1.41x, 128x32x4096 1.22x, 256x32x1 0.57x, 256x32x4 0.45x, 256x32x16 0.42x, 256x32x64 0.78x, 256x32x256 1.21x, 256x32x1024 1.31x, 256x32x4096 1.21x, 512x32x1 0.61x, 512x32x4 0.62x, 512x32x16 0.54x, 512x32x64 0.97x, 512x32x256 1.15x, 512x32x1024 1.13x, 512x32x4096 1.06x, 1024x32x1 0.60x, 1024x32x4 0.65x, 1024x32x16 0.65x, 1024x32x64 0.97x, 1024x32x256 1.04x, 1024x32x1024 1.02x, 40x40x1 0.67x, 40x40x4 0.70x, 40x40x16 0.70x, 40x40x64 1.50x, 40x40x256 2.43x, 40x40x1024 2.82x, 40x40x4096 2.71x, 48x48x1 0.74x, 48x48x4 0.73x, 48x48x16 0.78x, 48x48x64 1.55x, 48x48x256 2.38x, 48x48x1024 2.96x, 48x48x4096 2.93x, 56x56x1 0.72x, 56x56x4 0.70x, 56x56x16 0.68x, 56x56x64 1.59x, 56x56x256 2.30x, 56x56x1024 2.42x, 56x56x4096 2.41x, 64x64x1 0.62x, 64x64x4 0.59x, 64x64x16 0.66x, 64x64x64 1.65x, 64x64x256 1.79x, 64x64x1024 1.75x, 64x64x4096 1.77x, 128x64x1 0.55x, 128x64x4 0.56x, 128x64x16 0.50x, 128x64x64 1.07x, 128x64x256 0.95x, 128x64x1024 1.24x, 128x64x4096 1.17x, 256x64x1 0.62x, 256x64x4 0.62x, 256x64x16 0.58x, 256x64x64 1.16x, 256x64x256 1.03x, 256x64x1024 1.18x, 256x64x4096 1.10x, 512x64x1 0.58x, 512x64x4 0.63x, 512x64x16 0.66x, 512x64x64 1.09x, 512x64x256 1.10x, 512x64x1024 1.09x, 1024x64x1 0.62x, 1024x64x4 0.66x, 1024x64x16 0.82x, 1024x64x64 1.08x, 1024x64x256 1.02x, 1024x64x1024 1.04x, 2048x64x1 0.67x, 2048x64x4 0.79x, 2048x64x16 0.83x, 2048x64x64 0.98x, 2048x64x256 1.04x, 80x80x1 0.68x, 80x80x4 0.75x, 80x80x16 0.72x, 80x80x64 1.41x, 80x80x256 1.91x, 80x80x1024 2.10x, 80x80x4096 2.05x
gk over the CPU, M x N x batch: 4x4x1 0.02x, 4x4x4 0.15x, 4x4x16 0.23x, 4x4x64 0.42x, 4x4x256 0.45x, 4x4x1024 1.04x, 4x4x4096 1.81x, 8x8x1 0.04x, 8x8x4 0.10x, 8x8x16 0.38x, 8x8x64 0.58x, 8x8x256 0.81x, 8x8x1024 1.32x, 8x8x4096 1.67x, 16x8x1 0.04x, 16x8x4 0.11x, 16x8x16 0.35x, 16x8x64 0.59x, 16x8x256 0.94x, 16x8x1024 1.55x, 16x8x4096 1.88x, 32x8x1 0.06x, 32x8x4 0.12x, 32x8x16 0.41x, 32x8x64 0.64x, 32x8x256 1.18x, 32x8x1024 1.30x, 32x8x4096 1.62x, 64x8x1 0.05x, 64x8x4 0.14x, 64x8x16 0.35x, 64x8x64 0.64x, 64x8x256 0.97x, 64x8x1024 1.15x, 64x8x4096 1.57x, 128x8x1 0.06x, 128x8x4 0.18x, 128x8x16 0.51x, 128x8x64 0.76x, 128x8x256 0.93x, 128x8x1024 1.04x, 128x8x4096 1.22x, 256x8x1 0.07x, 256x8x4 0.21x, 256x8x16 0.48x, 256x8x64 0.77x, 256x8x256 0.87x, 256x8x1024 1.14x, 256x8x4096 0.96x, 16x16x1 0.06x, 16x16x4 0.13x, 16x16x16 0.27x, 16x16x64 0.52x, 16x16x256 1.05x, 16x16x1024 1.36x, 16x16x4096 1.82x, 32x16x1 0.10x, 32x16x4 0.15x, 32x16x16 0.35x, 32x16x64 0.66x, 32x16x256 1.32x, 32x16x1024 1.68x, 32x16x4096 2.18x, 64x16x1 0.10x, 64x16x4 0.17x, 64x16x16 0.41x, 64x16x64 0.72x, 64x16x256 1.19x, 64x16x1024 1.46x, 64x16x4096 1.71x, 128x16x1 0.12x, 128x16x4 0.19x, 128x16x16 0.45x, 128x16x64 0.72x, 128x16x256 0.94x, 128x16x1024 1.37x, 128x16x4096 1.24x, 256x16x1 0.15x, 256x16x4 0.21x, 256x16x16 0.49x, 256x16x64 0.72x, 256x16x256 0.75x, 256x16x1024 1.03x, 256x16x4096 0.98x, 512x16x1 0.13x, 512x16x4 0.18x, 512x16x16 0.27x, 512x16x64 0.45x, 512x16x256 0.84x, 512x16x1024 0.64x, 512x16x4096 0.56x, 24x24x1 0.07x, 24x24x4 0.13x, 24x24x16 0.31x, 24x24x64 0.54x, 24x24x256 1.22x, 24x24x1024 1.42x, 24x24x4096 1.67x, 32x32x1 0.09x, 32x32x4 0.13x, 32x32x16 0.24x, 32x32x64 0.49x, 32x32x256 1.22x, 32x32x1024 1.56x, 32x32x4096 1.66x, 64x32x1 0.10x, 64x32x4 0.14x, 64x32x16 0.28x, 64x32x64 0.56x, 64x32x256 0.97x, 64x32x1024 1.39x, 64x32x4096 1.46x, 128x32x1 0.11x, 128x32x4 0.16x, 128x32x16 0.31x, 128x32x64 0.54x, 128x32x256 0.80x, 128x32x1024 0.98x, 128x32x4096 1.00x, 256x32x1 0.11x, 256x32x4 0.15x, 256x32x16 0.24x, 256x32x64 0.42x, 256x32x256 0.71x, 256x32x1024 0.77x, 256x32x4096 0.87x, 512x32x1 0.16x, 512x32x4 0.21x, 512x32x16 0.27x, 512x32x64 0.49x, 512x32x256 0.65x, 512x32x1024 0.66x, 512x32x4096 0.74x, 1024x32x1 0.20x, 1024x32x4 0.29x, 1024x32x16 0.37x, 1024x32x64 0.57x, 1024x32x256 0.61x, 1024x32x1024 0.60x, 40x40x1 0.11x, 40x40x4 0.15x, 40x40x16 0.29x, 40x40x64 0.61x, 40x40x256 1.02x, 40x40x1024 1.37x, 40x40x4096 1.38x, 48x48x1 0.11x, 48x48x4 0.14x, 48x48x16 0.27x, 48x48x64 0.57x, 48x48x256 1.04x, 48x48x1024 1.25x, 48x48x4096 1.29x, 56x56x1 0.10x, 56x56x4 0.14x, 56x56x16 0.25x, 56x56x64 0.54x, 56x56x256 0.85x, 56x56x1024 0.91x, 56x56x4096 0.96x, 64x64x1 0.12x, 64x64x4 0.13x, 64x64x16 0.22x, 64x64x64 0.53x, 64x64x256 0.73x, 64x64x1024 0.83x, 64x64x4096 0.83x, 128x64x1 0.13x, 128x64x4 0.16x, 128x64x16 0.20x, 128x64x64 0.49x, 128x64x256 0.57x, 128x64x1024 0.79x, 128x64x4096 0.71x, 256x64x1 0.14x, 256x64x4 0.19x, 256x64x16 0.24x, 256x64x64 0.52x, 256x64x256 0.61x, 256x64x1024 0.75x, 256x64x4096 0.69x, 512x64x1 0.18x, 512x64x4 0.26x, 512x64x16 0.32x, 512x64x64 0.54x, 512x64x256 0.64x, 512x64x1024 0.77x, 1024x64x1 0.24x, 1024x64x4 0.35x, 1024x64x16 0.42x, 1024x64x64 0.62x, 1024x64x256 0.63x, 1024x64x1024 0.81x, 2048x64x1 0.32x, 2048x64x4 0.45x, 2048x64x16 0.52x, 2048x64x64 0.63x, 2048x64x256 0.66x, 80x80x1 0.15x, 80x80x4 0.17x, 80x80x16 0.23x, 80x80x64 0.47x, 80x80x256 0.68x, 80x80x1024 0.72x, 80x80x4096 0.65x
From a batch on, gk_share: gk and the CPU path at once on one batch, the GPU taking chunks from the front and the CPU from the back. Fitted against the best GPU backend, the shared one included, on the 91 points where it was timed and gk is the GPU’s choice. Chosen: from batch 1024.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| gk alone | 1.1307 | 1.74x | 36 | 1.360 | 0 |
| shared from batch 1024 | 1.0754 | 1.75x | 19 | 1.021 | 0 |
gk_share over gk alone, M x N x batch: 4x4x64 1.18x, 4x4x256 0.95x, 4x4x1024 0.60x, 4x4x4096 0.63x, 8x8x64 0.76x, 8x8x256 0.62x, 8x8x1024 0.57x, 8x8x4096 0.91x, 16x8x64 0.72x, 16x8x256 0.63x, 16x8x1024 0.58x, 16x8x4096 0.78x, 32x8x64 0.70x, 32x8x256 0.54x, 32x8x1024 0.70x, 32x8x4096 1.01x, 64x8x64 0.70x, 64x8x256 0.72x, 64x8x1024 0.80x, 64x8x4096 0.98x, 128x8x64 0.72x, 128x8x256 0.79x, 128x8x1024 1.05x, 128x8x4096 1.03x, 256x8x64 0.77x, 256x8x256 0.89x, 256x8x1024 1.01x, 256x8x4096 1.21x, 16x16x64 0.79x, 16x16x256 0.73x, 16x16x1024 0.92x, 16x16x4096 1.08x, 32x16x64 0.78x, 32x16x256 0.77x, 32x16x1024 0.82x, 32x16x4096 0.84x, 64x16x64 0.76x, 64x16x256 0.87x, 64x16x1024 0.96x, 64x16x4096 1.16x, 128x16x64 0.81x, 128x16x256 1.15x, 128x16x1024 1.11x, 128x16x4096 1.07x, 256x16x64 0.83x, 256x16x256 1.25x, 256x16x1024 1.44x, 256x16x4096 1.48x, 512x16x64 0.94x, 512x16x256 1.00x, 512x16x1024 1.26x, 512x16x4096 1.73x, 24x24x64 0.84x, 24x24x256 0.89x, 24x24x1024 1.17x, 24x24x4096 1.22x, 32x32x64 0.94x, 32x32x256 1.04x, 32x32x1024 1.17x, 32x32x4096 1.27x, 64x32x64 0.87x, 64x32x256 1.51x, 64x32x1024 1.23x, 64x32x4096 1.31x, 128x32x64 0.94x, 128x32x256 1.26x, 128x32x1024 1.57x, 128x32x4096 1.64x, 256x32x64 1.01x, 256x32x256 1.04x, 256x32x1024 1.11x, 256x32x4096 1.27x, 512x32x64 0.99x, 512x32x256 1.09x, 512x32x1024 1.20x, 512x32x4096 1.44x, 1024x32x64 1.03x, 1024x32x256 1.07x, 1024x32x1024 1.60x, 40x40x64 0.91x, 40x40x256 1.44x, 40x40x1024 1.27x, 40x40x4096 1.36x, 48x48x64 0.91x, 48x48x256 1.52x, 48x48x1024 1.39x, 48x48x4096 1.46x, 56x56x64 0.93x, 56x56x256 1.11x, 56x56x1024 1.68x, 56x56x4096 1.74x
GPU iff k <= gpu_max_k, l <= gpu_max_l (l = max(M, N)), batch * k >= gpu_min_batch_times_k and batch >= gpu_min_batch, with k = min(M, N), scored against the best of the CPU and the GPU backends (gk included). worst is over the points where the chosen backend was timed; a pick the cost model had to guess is listed in the warnings instead.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| oracle (best per point) | 1.0000 | 1.00x | 0 | 1.000 | 0 |
| policy in effect (‘56’, ‘16384’, ‘1’, ‘256’) | 1.0198 | 1.58x | 19 | 1.106 | 1 |
| fitted (‘56’, ‘16384’, ‘1’, ‘256’) | 1.0198 | 1.58x | 19 | 1.106 | 1 |
84 of 6171 combinations are within 0.5% of the best geomean: gpu_max_k 56 .. 80, gpu_min_batch_times_k 4096 .. 16384, gpu_min_batch 1 .. 16, gpu_max_l 128 .. none.
Large batches: the GPU also for k above gpu_max_k up to 80 (and l <= gpu_max_l) in a batch of at least 1024, fitted with the product rule (per cap, the rule, the clause over it and the rule again given the clause, the best kept) (product rule alone 1.0284, worst 1.58x; chosen 1.0198, worst 1.58x).
xychart-beta
title "Regret by gpu_min_batch_times_k"
x-axis "gpu_min_batch_times_k" [0, 64, 128, 256, 512, 1024, 2048, 4096, 8192, 16384, none]
y-axis "geometric-mean regret" 1.0 --> 1.58
line [1.5645, 1.2490, 1.1963, 1.1209, 1.0942, 1.0510, 1.0387, 1.0216, 1.0227, 1.0198, 1.0730]
| gpu_min_batch_times_k | 0 | 64 | 128 | 256 | 512 | 1024 | 2048 | 4096 | 8192 | 16384 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.5645 | 1.2490 | 1.1963 | 1.1209 | 1.0942 | 1.0510 | 1.0387 | 1.0216 | 1.0227 | 1.0198 | 1.0730 |
| worst | 61.97x | 7.95x | 7.82x | 4.17x | 4.17x | 2.38x | 2.38x | 1.75x | 1.75x | 1.58x | 2.18x |
xychart-beta
title "Regret by gpu_max_l"
x-axis "gpu_max_l" [16, 24, 32, 40, 48, 56, 64, 80, 96, 128, 192, 256, 384, 512, 768, 1024, 2048, none]
y-axis "geometric-mean regret" 1.0 --> 1.09
line [1.0741, 1.0696, 1.0595, 1.0552, 1.0509, 1.0476, 1.0350, 1.0321, 1.0321, 1.0241, 1.0241, 1.0198, 1.0198, 1.0211, 1.0211, 1.0212, 1.0260, 1.0260]
| gpu_max_l | 16 | 24 | 32 | 40 | 48 | 56 | 64 | 80 | 96 | 128 | 192 | 256 | 384 | 512 | 768 | 1024 | 2048 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.0741 | 1.0696 | 1.0595 | 1.0552 | 1.0509 | 1.0476 | 1.0350 | 1.0321 | 1.0321 | 1.0241 | 1.0241 | 1.0198 | 1.0198 | 1.0211 | 1.0211 | 1.0212 | 1.0260 | 1.0260 |
| worst | 2.18x | 2.18x | 1.98x | 1.98x | 1.98x | 1.98x | 1.65x | 1.65x | 1.65x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x |
xychart-beta
title "Regret by gpu_min_batch"
x-axis "gpu_min_batch" [1, 4, 16]
y-axis "geometric-mean regret" 1.0 --> 1.03
line [1.0198, 1.0198, 1.0198]
| gpu_min_batch | 1 | 4 | 16 |
|---|---|---|---|
| geomean | 1.0198 | 1.0198 | 1.0198 |
| worst | 1.58x | 1.58x | 1.58x |
xychart-beta
title "Regret by gpu_max_k"
x-axis "gpu_max_k" [8, 16, 24, 32, 40, 48, 56, 64, 80, 96, 128, 192, 256, 384, 512, 768, 1024, none]
y-axis "geometric-mean regret" 1.0 --> 1.09
line [1.0198, 1.0198, 1.0198, 1.0198, 1.0198, 1.0198, 1.0198, 1.0246, 1.0259, 1.0342, 1.0501, 1.0604, 1.0768, 1.0768, 1.0768, 1.0768, 1.0768, 1.0768]
| gpu_max_k | 8 | 16 | 24 | 32 | 40 | 48 | 56 | 64 | 80 | 96 | 128 | 192 | 256 | 384 | 512 | 768 | 1024 | none |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| geomean | 1.0198 | 1.0198 | 1.0198 | 1.0198 | 1.0198 | 1.0198 | 1.0198 | 1.0246 | 1.0259 | 1.0342 | 1.0501 | 1.0604 | 1.0768 | 1.0768 | 1.0768 | 1.0768 | 1.0768 | 1.0768 |
| worst | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.58x | 1.74x | 1.74x | 2.34x | 2.47x | 4.32x | 4.71x | 4.71x | 4.71x | 4.71x | 4.71x | 4.71x |
Held-out check of a per-k boundary against the product rule, fitted on 163 points and scored on the other 128; the verdict is a bootstrap.
| rule | fitted on train | train geomean | test geomean | test worst | verdict |
|---|---|---|---|---|---|
| product rule | [56, 16384, 1, 256] | 1.0218 | 1.0172 | 1.55x | baseline |
| per-k table | {“min_batch_by_k”: {“4”: null, “8”: 256, “16”: 256, “24”: 4096, “32”: 1024, “40”: 256, “48”: 256, “56”: 1024, “64”: 1024, “80”: null, “96”: null, “128”: null, “192”: null, “256”: null, “384”: null, “512”: null, “768”: null, “1024”: null}} | 1.0213 | 1.0360 | 1.81x | rejected (better in 0% of resamples, median gain -1.8%) |
Best backend per point (c CPU, J whole-matrix kernel, B block kernel, j and b the same after QR, g gk, G gk shared with the CPU, . not measured), what the rule picks, and the speedup of the best GPU backend over the CPU:
M x N \ batch 1 4 16 64 256 1024 4096
4 x 4 c c c c c g g
8 x 8 c c c c c g g
16 x 8 c c c c c g g
32 x 8 c c c c g g G
64 x 8 c c c c c g g
128 x 8 c c c c c G G
256 x 8 c c c c c J G
16 x 16 c c c c g g G
32 x 16 c c c c g g g
64 x 16 c c c c g g G
128 x 16 c c c c G G G
256 x 16 c c c c c G G
512 x 16 c c c c c c c
24 x 24 c c c c g G G
32 x 32 c c c c G G G
64 x 32 c c c c G G G
128 x 32 c c c c G G G
256 x 32 c c c c c c G
512 x 32 c c c c c c G
1024 x 32 c c c c c c .
40 x 40 c c c c G G G
48 x 48 c c c c G G G
56 x 56 c c c c c G G
64 x 64 c c c c G G G
128 x 64 c c c c c G G
256 x 64 c c c c c G G
512 x 64 c c c c c G .
1024 x 64 c c c c c G .
2048 x 64 c c c c c . c
80 x 80 c c c c G G G
96 x 96 c c c c c c c
128 x 128 c c c c c c c
256 x 128 c c c c c c c
512 x 128 c c c c c c c
1024 x 128 c c c c c c c
2048 x 128 c c c c c c .
192 x 192 c c c c c c .
256 x 256 c c c c c c .
512 x 256 c c c c c c .
1024 x 256 c c c c c b .
2048 x 256 c c c c c . .
384 x 384 c c c c c . .
512 x 512 c c c c . . .
768 x 768 c c c . . . .
1024 x 1024 c c c . . . .
M x N \ batch 1 4 16 64 256 1024 4096
4 x 4 c c c c c c J
8 x 8 c c c c c c G
16 x 8 c c c c c c G
32 x 8 c c c c c c G
64 x 8 c c c c c c G
128 x 8 c c c c c c G
256 x 8 c c c c c c G
16 x 16 c c c c c G G
32 x 16 c c c c c G G
64 x 16 c c c c c G G
128 x 16 c c c c c G G
256 x 16 c c c c c G G
512 x 16 c c c c c c c
24 x 24 c c c c c G G
32 x 32 c c c c c G G
64 x 32 c c c c c G G
128 x 32 c c c c c G G
256 x 32 c c c c c G G
512 x 32 c c c c c c c
1024 x 32 c c c c c c .
40 x 40 c c c c c G G
48 x 48 c c c c c G G
56 x 56 c c c c c G G
64 x 64 c c c c c G G
128 x 64 c c c c c G G
256 x 64 c c c c c G G
512 x 64 c c c c c c .
1024 x 64 c c c c c c .
2048 x 64 c c c c c . c
80 x 80 c c c c c G G
96 x 96 c c c c c c c
128 x 128 c c c c c c c
256 x 128 c c c c c c c
512 x 128 c c c c c c c
1024 x 128 c c c c c c c
2048 x 128 c c c c c c .
192 x 192 c c c c c c .
256 x 256 c c c c c c .
512 x 256 c c c c c c .
1024 x 256 c c c c c c .
2048 x 256 c c c c c . .
384 x 384 c c c c c . .
512 x 512 c c c c . . .
768 x 768 c c c . . . .
1024 x 1024 c c c . . . .
M x N \ batch 1 4 16 64 256 1024 4096
4 x 4 0.02 0.16 0.23 0.42 0.49 1.04 1.81
8 x 8 0.05 0.10 0.42 0.58 0.81 1.32 1.67
16 x 8 0.04 0.11 0.35 0.59 0.94 1.55 1.88
32 x 8 0.06 0.12 0.41 0.64 1.18 1.30 1.62
64 x 8 0.06 0.14 0.35 0.64 0.97 1.15 1.57
128 x 8 0.06 0.19 0.51 0.76 0.93 1.04 1.22
256 x 8 0.07 0.22 0.48 0.77 0.87 1.19 1.02
16 x 16 0.06 0.19 0.30 0.56 1.05 1.36 1.82
32 x 16 0.11 0.15 0.49 0.66 1.32 1.68 2.18
64 x 16 0.15 0.19 0.41 0.73 1.19 1.46 1.71
128 x 16 0.12 0.19 0.59 0.72 0.94 1.37 1.24
256 x 16 0.19 0.21 0.59 0.72 0.75 1.03 0.98
512 x 16 0.19 0.28 0.54 0.59 0.84 0.64 0.56
24 x 24 0.07 0.13 0.31 0.54 1.22 1.42 1.67
32 x 32 0.17 0.24 0.42 0.49 1.22 1.56 1.66
64 x 32 0.19 0.28 0.35 0.56 0.97 1.39 1.46
128 x 32 0.23 0.24 0.56 0.54 0.80 0.98 1.00
256 x 32 0.20 0.34 0.57 0.54 0.71 0.77 0.87
512 x 32 0.26 0.34 0.50 0.51 0.65 0.66 0.74
1024 x 32 0.33 0.45 0.57 0.59 0.61 0.60 .
40 x 40 0.16 0.21 0.42 0.61 1.02 1.37 1.38
48 x 48 0.15 0.20 0.34 0.57 1.04 1.25 1.29
56 x 56 0.15 0.19 0.37 0.54 0.85 0.91 0.96
64 x 64 0.19 0.22 0.33 0.53 0.73 0.83 0.83
128 x 64 0.24 0.29 0.39 0.49 0.60 0.79 0.71
256 x 64 0.23 0.31 0.41 0.52 0.61 0.75 0.69
512 x 64 0.31 0.40 0.49 0.54 0.64 0.77 .
1024 x 64 0.38 0.52 0.51 0.62 0.63 0.81 .
2048 x 64 0.48 0.57 0.62 0.64 0.66 . 0.86
80 x 80 0.22 0.23 0.31 0.47 0.68 0.72 0.65
96 x 96 0.23 0.29 0.34 0.36 0.49 0.46 0.43
128 x 128 0.21 0.27 0.33 0.42 0.45 0.40 0.41
256 x 128 0.28 0.37 0.41 0.59 0.56 0.59 0.57
512 x 128 0.34 0.34 0.46 0.57 0.62 0.66 0.63
1024 x 128 0.39 0.40 0.56 0.62 0.62 0.78 0.72
2048 x 128 0.45 0.50 0.62 0.69 0.73 0.90 .
192 x 192 0.21 0.21 0.26 0.28 0.25 0.23 .
256 x 256 0.31 0.29 0.25 0.24 0.23 0.21 .
512 x 256 0.44 0.43 0.35 0.36 0.33 0.32 .
1024 x 256 0.47 0.48 0.55 0.43 0.40 gpu .
2048 x 256 0.57 0.56 0.54 0.53 0.51 . .
384 x 384 0.34 0.30 0.19 0.17 0.15 . .
512 x 512 0.53 0.39 0.18 0.16 . . .
768 x 768 0.52 0.35 0.15 . . . .
1024 x 1024 0.68 0.30 0.25 . . . .
The rule of stage 2 again, with constants of its own, for svdvals: both sides skip the vectors, by different amounts. Fitted on the 145 points where gk and the CPU were timed for singular values alone and gk is the GPU’s choice. Chosen: GPU iff k <= 56, l <= 56, batch * k >= 16384 and batch >= 1.
| rule | geomean regret | worst | >10% | total time / oracle | est. picks |
|---|---|---|---|---|---|
| CPU always | 1.0042 | 1.22x | 3 | 1.009 | 0 |
| as with vectors (stage 2’s rule) | 1.0042 | 1.22x | 3 | 1.009 | 0 |
| fitted (‘56’, ‘16384’, ‘1’, ‘56’) | 1.0042 | 1.22x | 3 | 1.009 | 0 |
Held out: fitted on 87 points ((‘80’, ‘16384’, ‘1’, ‘80’)), scored on the other 58: geomean 1.0196x, worst 1.79x, against 1.0071x, worst 1.20x as with vectors.
gk over the CPU for singular values alone, M x N x batch: 8x8x1 0.03x, 8x8x4 0.10x, 8x8x16 0.25x, 8x8x64 0.69x, 8x8x256 0.74x, 16x8x1 0.03x, 16x8x4 0.10x, 16x8x16 0.33x, 16x8x64 0.69x, 16x8x256 0.86x, 32x8x1 0.03x, 32x8x4 0.12x, 32x8x16 0.37x, 32x8x64 0.60x, 32x8x256 1.09x, 64x8x1 0.03x, 64x8x4 0.12x, 64x8x16 0.35x, 64x8x64 1.15x, 64x8x256 0.91x, 128x8x1 0.04x, 128x8x4 0.12x, 128x8x16 0.45x, 128x8x64 1.22x, 128x8x256 0.89x, 256x8x1 0.05x, 256x8x4 0.15x, 256x8x16 0.56x, 256x8x64 0.83x, 256x8x256 0.92x, 16x16x1 0.04x, 16x16x4 0.10x, 16x16x16 0.31x, 16x16x64 0.54x, 16x16x256 0.84x, 32x16x1 0.04x, 32x16x4 0.11x, 32x16x16 0.37x, 32x16x64 0.52x, 32x16x256 0.88x, 64x16x1 0.04x, 64x16x4 0.14x, 64x16x16 0.39x, 64x16x64 0.62x, 64x16x256 0.90x, 128x16x1 0.05x, 128x16x4 0.17x, 128x16x16 0.33x, 128x16x64 0.63x, 128x16x256 0.76x, 256x16x1 0.07x, 256x16x4 0.19x, 256x16x16 0.47x, 256x16x64 0.68x, 256x16x256 0.89x, 512x16x1 0.12x, 512x16x4 0.20x, 512x16x16 0.32x, 512x16x64 0.46x, 512x16x256 0.44x, 24x24x1 0.04x, 24x24x4 0.09x, 24x24x16 0.31x, 24x24x64 0.44x, 24x24x256 0.89x, 32x32x1 0.04x, 32x32x4 0.11x, 32x32x16 0.24x, 32x32x64 0.41x, 32x32x256 0.84x, 64x32x1 0.06x, 64x32x4 0.12x, 64x32x16 0.30x, 64x32x64 0.45x, 64x32x256 0.61x, 128x32x1 0.06x, 128x32x4 0.13x, 128x32x16 0.32x, 128x32x64 0.44x, 128x32x256 0.56x, 256x32x1 0.09x, 256x32x4 0.14x, 256x32x16 0.22x, 256x32x64 0.40x, 256x32x256 0.75x, 512x32x1 0.12x, 512x32x4 0.20x, 512x32x16 0.29x, 512x32x64 0.38x, 512x32x256 0.49x, 1024x32x1 0.15x, 1024x32x4 0.25x, 1024x32x16 0.31x, 1024x32x64 0.62x, 1024x32x256 0.46x, 40x40x1 0.04x, 40x40x4 0.08x, 40x40x16 0.21x, 40x40x64 0.36x, 40x40x256 0.67x, 48x48x1 0.05x, 48x48x4 0.08x, 48x48x16 0.19x, 48x48x64 0.32x, 48x48x256 0.67x, 56x56x1 0.05x, 56x56x4 0.08x, 56x56x16 0.18x, 56x56x64 0.32x, 56x56x256 0.68x, 64x64x1 0.06x, 64x64x4 0.07x, 64x64x16 0.16x, 64x64x64 0.32x, 64x64x256 0.56x, 128x64x1 0.08x, 128x64x4 0.11x, 128x64x16 0.17x, 128x64x64 0.30x, 128x64x256 0.40x, 256x64x1 0.08x, 256x64x4 0.13x, 256x64x16 0.21x, 256x64x64 0.35x, 256x64x256 0.44x, 512x64x1 0.11x, 512x64x4 0.17x, 512x64x16 0.22x, 512x64x64 0.37x, 512x64x256 0.46x, 1024x64x1 0.15x, 1024x64x4 0.22x, 1024x64x16 0.29x, 1024x64x64 0.41x, 1024x64x256 0.47x, 2048x64x1 0.20x, 2048x64x4 0.27x, 2048x64x16 0.36x, 2048x64x64 0.44x, 2048x64x256 0.48x, 80x80x1 0.11x, 80x80x4 0.13x, 80x80x16 0.18x, 80x80x64 0.39x, 80x80x256 0.88x
Where the rule above chooses the CPU, the bidiag backend (GPU bidiagonalization, then LAPACK’s bidiagonal solve) from a threshold k = min(M, N) on (0: never), for batches up to a cap (0: any; it solves a batch one matrix after another, the CPU path spreads one over every core), fitted on the points where bidiag was timed (k >= 128, within the cost cap): the region the threshold decides. With vectors it is scored against the best of all backends, bidiag included; for singular values alone (svdvals), bidiag against the CPU at the points where the rule chooses the CPU.
| threshold | batch cap | geomean regret | worst | without bidiag: geomean | worst | held out (fitted on half) | |
|---|---|---|---|---|---|---|---|
| with vectors | 1024 | 2 | 1.0052 | 1.22x | 1.0573 | 2.32x | from 1024, batch <= 2: 1.0077 vs 1.0247 |
| singular values alone | 1024 | 2 | 1.0041 | 1.32x | 1.0685 | 2.66x | from 1024, batch <= 2: 1.0101 vs 1.0284 |
bidiag over the CPU (with vectors), M x N x batch: 128x128x1 0.39x, 128x128x4 0.15x, 128x128x16 0.05x, 128x128x64 0.04x, 128x128x256 0.04x, 256x128x1 0.46x, 256x128x4 0.19x, 256x128x16 0.06x, 256x128x64 0.05x, 256x128x256 0.05x, 512x128x1 0.51x, 512x128x4 0.17x, 512x128x16 0.07x, 512x128x64 0.06x, 512x128x256 0.06x, 1024x128x1 0.56x, 1024x128x4 0.19x, 1024x128x16 0.09x, 1024x128x64 0.08x, 1024x128x256 0.07x, 2048x128x1 0.59x, 2048x128x4 0.21x, 2048x128x16 0.10x, 2048x128x64 0.10x, 2048x128x256 0.09x, 192x192x1 0.46x, 192x192x4 0.15x, 192x192x16 0.06x, 192x192x64 0.05x, 192x192x256 0.04x, 256x256x1 0.65x, 256x256x4 0.22x, 256x256x16 0.08x, 256x256x64 0.07x, 512x256x1 0.66x, 512x256x4 0.24x, 512x256x16 0.08x, 512x256x64 0.08x, 1024x256x1 0.71x, 1024x256x4 0.24x, 1024x256x16 0.13x, 1024x256x64 0.09x, 2048x256x1 0.77x, 2048x256x4 0.25x, 2048x256x16 0.12x, 2048x256x64 0.11x, 384x384x1 0.73x, 384x384x4 0.29x, 384x384x16 0.11x, 384x384x64 0.10x, 512x512x1 1.02x, 512x512x4 0.43x, 512x512x16 0.17x, 512x512x64 0.16x, 768x768x1 1.14x, 768x768x4 0.53x, 768x768x16 0.29x, 1024x1024x1 1.42x, 1024x1024x4 0.66x, 1024x1024x16 0.64x, 1536x1536x1 1.49x, 1536x1536x2 1.13x, 1536x1536x4 0.76x, 2048x2048x1 1.81x, 2048x2048x2 1.46x, 2048x2048x4 1.22x, 3072x3072x1 2.15x, 4096x4096x1 2.32x
bidiag over the CPU (singular values alone), M x N x batch: 128x128x1 0.36x, 128x128x4 0.14x, 128x128x16 0.05x, 128x128x64 0.04x, 128x128x256 0.04x, 256x128x1 0.38x, 256x128x4 0.14x, 256x128x16 0.06x, 256x128x64 0.04x, 256x128x256 0.04x, 512x128x1 0.38x, 512x128x4 0.13x, 512x128x16 0.06x, 512x128x64 0.04x, 512x128x256 0.04x, 1024x128x1 0.39x, 1024x128x4 0.12x, 1024x128x16 0.06x, 1024x128x64 0.04x, 1024x128x256 0.04x, 2048x128x1 0.39x, 2048x128x4 0.15x, 2048x128x16 0.06x, 2048x128x64 0.05x, 2048x128x256 0.04x, 192x192x1 0.31x, 192x192x4 0.11x, 192x192x16 0.04x, 192x192x64 0.03x, 192x192x256 0.03x, 256x256x1 0.44x, 256x256x4 0.14x, 256x256x16 0.06x, 256x256x64 0.05x, 512x256x1 0.44x, 512x256x4 0.14x, 512x256x16 0.07x, 512x256x64 0.05x, 1024x256x1 0.42x, 1024x256x4 0.15x, 1024x256x16 0.07x, 1024x256x64 0.05x, 2048x256x1 0.44x, 2048x256x4 0.15x, 2048x256x16 0.07x, 2048x256x64 0.06x, 384x384x1 0.57x, 384x384x4 0.19x, 384x384x16 0.08x, 384x384x64 0.07x, 512x512x1 0.84x, 512x512x4 0.28x, 512x512x16 0.12x, 512x512x64 0.11x, 768x768x1 0.98x, 768x768x4 0.34x, 768x768x16 0.24x, 1024x1024x1 1.55x, 1024x1024x4 0.48x, 1024x1024x16 0.67x, 1536x1536x1 1.71x, 1536x1536x2 1.07x, 1536x1536x4 0.65x, 2048x2048x1 2.15x, 2048x2048x2 1.64x, 2048x2048x4 1.32x, 3072x3072x1 2.57x, 4096x4096x1 2.66x
For singular values alone, where the rules above choose the CPU or bidiag, the band backend (the two-stage reduction: A to a band on the GPU in blocks whose work is matrix products, then LAPACK’s band-to-bidiagonal on the CPU) from a threshold k on (0: never), within bidiag’s batch cap, fitted on the 18 points where it was timed (k >= 512) against the CPU, bidiag and band.
Chosen: 1536 (in effect: never): 1.0434 geometric-mean regret, worst 2.15x; without band 1.2311, worst 3.24x.
band over bidiag, M x N x batch: 512x512x1 0.74x, 512x512x4 0.79x, 512x512x16 0.80x, 512x512x64 0.80x, 768x768x1 0.80x, 768x768x4 0.84x, 768x768x16 0.86x, 1024x1024x1 0.83x, 1024x1024x4 0.95x, 1024x1024x16 0.98x, 1536x1536x1 1.13x, 1536x1536x2 1.20x, 1536x1536x4 1.28x, 2048x2048x1 1.33x, 2048x2048x2 1.51x, 2048x2048x4 1.62x, 3072x3072x1 2.21x, 4096x4096x1 3.24x
band over the CPU, M x N x batch: 512x512x1 0.62x, 512x512x4 0.22x, 512x512x16 0.10x, 512x512x64 0.09x, 768x768x1 0.78x, 768x768x4 0.28x, 768x768x16 0.20x, 1024x1024x1 1.28x, 1024x1024x4 0.46x, 1024x1024x16 0.66x, 1536x1536x1 1.93x, 1536x1536x2 1.29x, 1536x1536x4 0.83x, 2048x2048x1 2.87x, 2048x2048x2 2.48x, 2048x2048x4 2.15x, 3072x3072x1 5.68x, 4096x4096x1 8.62x
Pass-to-pass ratio, 2122 measurements: median 1.017, p90 1.095, max 2.89.
| runtime | n | median | p90 | max |
|---|---|---|---|---|
| <1 ms | 738 | 1.030 | 1.249 | 2.89 |
| 1-3 ms | 378 | 1.013 | 1.118 | 1.73 |
| 3-10 ms | 345 | 1.013 | 1.057 | 1.45 |
| 10-30 ms | 226 | 1.012 | 1.049 | 1.22 |
| 30-100 ms | 193 | 1.011 | 1.045 | 1.09 |
| >100 ms | 242 | 1.012 | 1.044 | 1.16 |