Raw output of the experiments in this folder, Apple M5 Pro (Mac17,8), 20 GPU cores, 18 CPU cores, 48 GB, macOS 27.0.1, on mains, 2026-10-03. Times in ms; min of repeats. See ../performance-headroom-apple-m5-pro.md for what each measures. == exp_batch 18 (with the 250 ms warm-up; the tables in the report) == cpu workers 18 (hardware 18) op shape batch | serial par gpu hybrid | par/s gpu/p hyb/b | split eigh 8x8 4096 | 9.01 0.64 0.50 0.47 | 14.0x 1.29x 1.06x | gpu 1612/4096 eigh 16x16 4096 | 32.97 2.16 2.17 1.28 | 15.2x 1.00x 1.70x | gpu 2045/4096 eigh 32x32 256 | 8.19 0.57 1.50 0.52 | 14.3x 0.38x 1.10x | gpu 50/256 eigh 32x32 4096 | 130.92 9.42 16.07 6.43 | 13.9x 0.59x 1.46x | gpu 1287/4096 eigh 64x64 256 | 35.39 2.46 9.83 2.20 | 14.4x 0.25x 1.12x | gpu 36/256 eigh 64x64 2048 | 283.93 18.72 75.90 16.34 | 15.2x 0.25x 1.15x | gpu 405/2048 eigh 128x128 256 | 122.59 9.82 42.05 9.79 | 12.5x 0.23x 1.00x | gpu 34/256 eigh 256x256 64 | 123.44 12.08 72.61 15.71 | 10.2x 0.17x 0.77x | gpu 6/64 eigh 512x512 16 | 132.80 15.31 129.46 29.63 | 8.7x 0.12x 0.52x | gpu 1/16 svd 8x8 4096 | 18.42 1.26 1.23 0.94 | 14.6x 1.03x 1.31x | gpu 1765/4096 svd 32x32 4096 | 210.33 14.98 24.82 10.64 | 14.0x 0.60x 1.41x | gpu 1542/4096 svd 64x64 256 | 45.67 3.52 8.45 2.87 | 13.0x 0.42x 1.23x | gpu 53/256 svd 128x128 64 | 63.73 5.73 13.87 4.99 | 11.1x 0.41x 1.15x | gpu 13/64 svd 256x256 16 | 53.79 5.60 24.25 13.15 | 9.6x 0.23x 0.43x | gpu 2/16 svd 512x512 4 | 62.58 18.14 48.55 33.42 | 3.4x 0.37x 0.54x | gpu 1/4 svd 1024x64 64 | 33.36 4.91 10.30 4.58 | 6.8x 0.48x 1.07x | gpu 14/64 svd 2048x256 16 | 129.85 18.91 38.72 22.52 | 6.9x 0.49x 0.84x | gpu 4/16 qr 16x16 10000 | 21.33 2.04 3.80 1.83 | 10.4x 0.54x 1.12x | gpu 2449/10000 qr 64x64 1000 | 26.98 2.26 4.46 2.19 | 12.0x 0.51x 1.03x | gpu 285/1000 qr 256x128 1000 | 294.89 37.40 35.30 22.90 | 7.9x 1.06x 1.54x | gpu 514/1000 qr 512x512 32 | 93.79 12.43 16.86 10.03 | 7.5x 0.74x 1.24x | gpu 16/32 qr 1024x512 16 | 91.58 15.40 18.08 12.09 | 5.9x 0.85x 1.27x | gpu 6/16 == exp_batch 18 (first run, warm-up of one call; for comparison) == cpu workers 18 (hardware 18) op shape batch | serial par gpu hybrid | par/s gpu/p hyb/b | split eigh 8x8 4096 | 9.10 0.64 1.76 0.52 | 14.2x 0.37x 1.25x | gpu 1261/4096 eigh 16x16 4096 | 33.55 3.67 2.29 1.90 | 9.1x 1.61x 1.21x | gpu 2902/4096 eigh 32x32 256 | 8.04 0.59 1.50 0.53 | 13.7x 0.39x 1.10x | gpu 50/256 eigh 32x32 4096 | 129.64 10.03 16.15 6.03 | 12.9x 0.62x 1.66x | gpu 1334/4096 eigh 64x64 256 | 35.58 2.49 9.77 2.10 | 14.3x 0.25x 1.19x | gpu 36/256 eigh 64x64 2048 | 283.77 18.74 75.21 16.01 | 15.1x 0.25x 1.17x | gpu 347/2048 eigh 128x128 256 | 123.24 9.96 42.18 11.20 | 12.4x 0.24x 0.89x | gpu 42/256 eigh 256x256 64 | 122.73 12.23 73.73 21.65 | 10.0x 0.17x 0.56x | gpu 8/64 eigh 512x512 16 | 132.19 16.08 129.57 33.92 | 8.2x 0.12x 0.47x | gpu 1/16 svd 8x8 4096 | 18.42 1.54 1.26 0.72 | 12.0x 1.22x 1.74x | gpu 1913/4096 svd 32x32 4096 | 224.84 15.37 25.06 10.70 | 14.6x 0.61x 1.44x | gpu 1557/4096 svd 64x64 256 | 48.16 3.58 8.40 3.07 | 13.4x 0.43x 1.17x | gpu 54/256 svd 128x128 64 | 64.50 5.79 13.83 5.01 | 11.1x 0.42x 1.15x | gpu 16/64 svd 256x256 16 | 53.85 5.81 24.32 13.29 | 9.3x 0.24x 0.44x | gpu 2/16 svd 512x512 4 | 62.24 18.93 49.16 33.44 | 3.3x 0.39x 0.57x | gpu 1/4 svd 1024x64 64 | 34.55 5.04 10.30 4.02 | 6.9x 0.49x 1.26x | gpu 18/64 svd 2048x256 16 | 130.59 19.21 38.51 22.64 | 6.8x 0.50x 0.85x | gpu 4/16 qr 16x16 10000 | 20.19 2.16 4.34 1.70 | 9.3x 0.50x 1.27x | gpu 2825/10000 qr 64x64 1000 | 26.40 2.47 4.21 2.22 | 10.7x 0.59x 1.11x | gpu 370/1000 qr 256x128 1000 | 295.28 37.56 35.25 22.36 | 7.9x 1.07x 1.58x | gpu 516/1000 qr 512x512 32 | 95.36 12.58 16.61 10.78 | 7.6x 0.76x 1.17x | gpu 10/32 qr 1024x512 16 | 93.37 15.93 18.94 12.80 | 5.9x 0.84x 1.24x | gpu 6/16 == exp_gemm == device Apple M5 Pro GPU bandwidth: read 282 GB/s, copy 274 GB/s (read+write) CPU bandwidth (18 threads, vDSP_sve): read 230 GB/s === 2048 x 2048 x 2048 === MPS fp32 2.27 ms 7.56 TFLOP/s err 6.6e-08 mm_f32_64x32 2.22 ms 7.74 TFLOP/s err 7.6e-08 mm_f32_64x64 2.26 ms 7.62 TFLOP/s err 9.7e-08 mm_f32_128x64 2.44 ms 7.05 TFLOP/s err 5.3e-08 mm_f32r_64x64 1.71 ms 10.03 TFLOP/s err 4.6e-05 mm_f32r_128x64 1.32 ms 13.05 TFLOP/s err 6.5e-05 mm_f16_64x64 0.85 ms 20.31 TFLOP/s err 1.9e-05 mm_f16_128x64 1.42 ms 12.13 TFLOP/s err 2.0e-05 mm_bf16_128x64 0.87 ms 19.74 TFLOP/s err 1.7e-04 Accelerate sgemm (CPU) 6.36 ms 2.70 TFLOP/s err 6.1e-08 concurrent: GPU MPS 6.10 TFLOP/s and CPU 2.63 TFLOP/s at once -> 8.74 TFLOP/s combined === 4096 x 4096 x 4096 === MPS fp32 19.41 ms 7.08 TFLOP/s err 7.0e-08 mm_f32_64x32 18.31 ms 7.50 TFLOP/s err 7.3e-08 mm_f32_64x64 17.84 ms 7.70 TFLOP/s err 1.0e-07 mm_f32_128x64 18.21 ms 7.55 TFLOP/s err 8.0e-08 mm_f32r_64x64 15.74 ms 8.73 TFLOP/s err 4.1e-05 mm_f32r_128x64 11.85 ms 11.60 TFLOP/s err 3.3e-05 mm_f16_64x64 6.24 ms 22.02 TFLOP/s err 1.3e-05 mm_f16_128x64 6.01 ms 22.87 TFLOP/s err 1.4e-05 mm_bf16_128x64 6.03 ms 22.80 TFLOP/s err 1.2e-04 Accelerate sgemm (CPU) 54.33 ms 2.53 TFLOP/s err 8.8e-08 concurrent: GPU MPS 7.37 TFLOP/s and CPU 2.36 TFLOP/s at once -> 9.74 TFLOP/s combined == exp_large (kd = 64) == N | trid vec trid val | cpu vec cpu val | ssytrd sy2sb sb2st ssterf sstedc floor 1024 | 34.5 26.4 | 40.3 18.0 | 15.8 3.9 20.6 5.4 10.8 2.5 2048 | 115.3 90.0 | 234.3 79.3 | 110.2 20.5 85.2 20.8 42.0 20.3 4096 | 520.7 387.5 | 2456.5 464.8 | 1727.8 171.8 327.0 79.4 174.8 162.5 8192 | 3056.1 2342.4 | nan 2554.5 | 14954.2 1083.7 1277.9 308.7 812.0 1299.7 == KD=16|32|128 exp_large 4096 (sy2sb, sb2st columns are the ones that change) == kd=16 4096 | 528.2 391.2 | 2452.6 469.3 | 1749.5 358.4 117.6 80.2 178.8 162.5 kd=32 4096 | 529.7 394.3 | 2462.7 470.9 | 1740.5 249.4 130.3 80.8 178.9 162.5 kd=128 4096 | 528.4 395.3 | 2498.8 471.6 | 1786.8 152.6 388.4 79.9 178.6 162.5 == exp_svdlarge (sbdsdc on the matrix's own bidiagonal) == k | bid vec bid val | cpu vec cpu val | sbdsdc I sbdsdc N floor (cpu sgebrd 2048: 184 ms) 2048 | 315.1 171.8 | 456.4 213.8 | 134.2 23.7 81.2 (cpu sgebrd 4096: 1888 ms) 4096 | 1783.4 1029.9 | 3612.4 2003.4 | 741.9 92.6 649.8 == exp_tdql, first version (one 17.7 KB layout for every N) == N batch | proto lib gpu cpu par cpu ser | vs gpu vs par | resid orth dw bad 8 4096 | 1.60 0.49 0.89 9.56 | 0.3x 0.56x | 5.7e-07 6.6e-07 4.4e-07 0 16 4096 | 3.52 2.27 2.40 33.04 | 0.6x 0.68x | 7.4e-07 9.5e-07 5.8e-07 0 32 256 | 1.18 1.60 0.60 8.19 | 1.4x 0.51x | 9.5e-07 1.0e-06 5.9e-07 0 32 4096 | 11.71 16.21 9.18 129.10 | 1.4x 0.78x | 9.1e-07 1.0e-06 5.1e-07 0 48 2048 | 13.29 29.23 11.33 155.42 | 2.2x 0.85x | 1.1e-06 1.4e-06 4.4e-07 0 64 256 | 3.72 10.03 2.66 35.69 | 2.7x 0.72x | 1.3e-06 1.5e-06 5.1e-07 0 64 2048 | 21.94 75.83 19.85 284.80 | 3.5x 0.90x | 1.4e-06 1.6e-06 5.1e-07 0 64 8192 | 85.49 305.19 80.33 1134.13 | 3.6x 0.94x | 1.3e-06 1.5e-06 7.1e-07 0 == exp_tdql (4.7 KB layout for N <= 32, 17.7 KB above; 250 ms warm-up) == N batch | proto lib gpu cpu par cpu ser | vs gpu vs par | resid orth dw bad 8 4096 | 0.46 0.50 1.13 9.50 | 1.1x 2.46x | 5.7e-07 6.6e-07 4.4e-07 0 16 4096 | 1.22 2.15 2.48 34.77 | 1.8x 2.03x | 7.4e-07 9.5e-07 5.8e-07 0 24 4096 | 2.40 7.10 5.83 79.89 | 3.0x 2.43x | 8.0e-07 8.8e-07 6.5e-07 0 32 256 | 0.53 1.49 0.92 8.46 | 2.8x 1.74x | 1.0e-06 1.0e-06 4.8e-07 0 32 2048 | 2.13 8.26 4.41 67.39 | 3.9x 2.08x | 9.3e-07 1.0e-06 5.3e-07 0 32 4096 | 4.04 16.26 10.27 134.01 | 4.0x 2.54x | 9.8e-07 1.1e-06 5.7e-07 0 48 2048 | 13.11 29.26 11.71 160.03 | 2.2x 0.89x | 1.2e-06 1.2e-06 7.5e-07 0 64 256 | 3.57 9.97 2.34 34.28 | 2.8x 0.66x | 1.4e-06 1.5e-06 4.8e-07 0 64 2048 | 22.26 75.54 18.54 273.61 | 3.4x 0.83x | 1.3e-06 1.5e-06 4.5e-07 0 == exp_stage tdql64.metallib 64 / tdql32.metallib 32 == stage<=1: tridiagonalization; <=2: + forming Q; <=5: + QL without updating Z; <=4: the whole kernel N=64 batch 2048: stage<=1 4.17 ms (TGmem 17680) stage<=2 7.52 ms stage<=5 18.10 ms stage<=4 22.30 ms N=32 batch 2048: stage<=1 0.33 ms (TGmem 4752) stage<=2 0.60 ms stage<=5 1.59 ms stage<=4 1.99 ms N=32 batch 2048 with the 17.7 KB layout: stage<=1 1.67 ms stage<=2 2.23 ms stage<=3 5.85 ms stage<=4 5.93 ms (with sqrt and two divides in the QL chase instead of rsqrt: N=64 stage<=3 23.14 ms, stage<=4 22.25 ms) == exp_mrrr == N | sstedc sstemr | orth dc orth mr | max|dw| 2048 | 42.7 292.3 | 2.0e-06 8.8e-05 | 3.5e-06 4096 | 175.7 1506.0 | 2.7e-06 1.0e-04 | 2.3e-06 8192 | 807.1 7676.5 | 3.8e-06 3.4e-04 | 2.3e-06 == exp_bisect == N | ssterf stebz x1 stebz x18 | max|dw| 2048 | 20.6 327.6 21.1 | 1.1e-05 4096 | 79.3 1180.0 75.6 | 1.3e-05 8192 | 308.2 4456.7 277.4 | 4.7e-06 == exp_stedc_threads == N=2048 sstedc: multi-threaded 42.4 ms, single-threaded 43.4 ms (1.0x) N=4096 sstedc: multi-threaded 175.3 ms, single-threaded 186.5 ms (1.1x)