Status: done in 2.15.0 (2026-10-07); see Done.
Six shader files are built with -fno-fast-math (CMakeLists.txt):
Eigh_Jacobi, Eigh_BlockJacobi, Eigh_Tridiag, Svd_Jacobi,
Svd_BlockJacobi, Svd_Bidiag. That makes / and sqrt() IEEE sequences,
and a kernel with any of them in it compiles all of its arithmetic in IEEE
mode, even if that division or square root is never executed (2.15.0 found
an untaken sqrt() costing the TSQR top a third of its time). In 2.15.0 the
band panels, the one-stage reductions’ reflectors and bisection were moved to
fast division and square root with a Newton step (div1, sqrt1). Do the
same, kernel by kernel, where it pays, in the rest.
The whole library built once with every -fno-fast-math removed, against the
normal build, same session, M5 Pro (before the 2.15.0 changes to the
one-stage reductions):
| point | IEEE | fast math |
|---|---|---|
| SVD Jacobi, 1024 x 32x32 | 7.38 ms | 6.10 ms (1.21x) |
| SVD Jacobi, 64 x 128x128 | 33.6 | 33.5 |
| SVD block Jacobi, 16 x 256x256 | 25.2 | 24.6 (1.02x) |
| eigh threadgroup mode, 1024 x 32x32 | 4.67 | 4.61 |
| eigh threadgroup mode, 256 x 64x64 | 10.40 | 10.44 |
| eigh simd mode, 1024 x 16x16 | 0.906 | 0.904 |
| eigh block Jacobi, 16 x 256x256 | 20.7 | 20.5 |
eigvalsh tridiag, 2048 |
39.8 | 35.2 (1.13x) |
svdvals bidiag, 2048 |
93.8 | 85.4 (1.10x) |
After 2.15.0’s div1/sqrt1 in the one-stage reductions, eigvalsh tridiag
at 2048 is 36.0 ms (most of the gain), svdvals bidiag 91.8: about 7% is
left there, which division and square root alone do not account for (IEEE
max, comparisons and the like, which only fast math relaxes).
Svd_Jacobi.metal: the rotation’s angle (a square root and divisions per
pair) to div1/sqrt1; check its accuracy on the suite’s hard cases
(graded, nearly rank-deficient, tiny values) before timing.About a day, mostly checking accuracy.
1.2x for the SVD’s Jacobi kernel on batches of small matrices; up to 7% for
svdvals on bidiag. On the M5 Pro the Jacobi kernel is rarely chosen
(golub_kahan and the CPU take most of its region), but estimated policies
on other Macs use it more.
div1 and sqrt1 in shaders/Svd_Bidiag.metal (and their copy in
shaders/Eigh_Tridiag.metal); the rotation in shaders/Svd_Jacobi.metal and
shaders/eigh_jacobi_common.h.
The SVD’s Jacobi kernel. Its rotation (a square root, three divisions
and an rsqrt a pair) and its output (a square root a column, a division an
entry of U) now use the fast approximations with Newton steps, j_div,
j_sqrt and j_rsqrt in eigh_jacobi_common.h. With one Newton step for
rsqrt, c came out a little low every time (c^2 + s^2 < 1) and V’s
orthogonality was 10x worse (8.9e-5 at 512 x 512, against 7.1e-6): four
cases of the suite failed. With two steps it is better than the IEEE
sequence’s (4.5e-6 at 512, 2.3e-6 at 200 against 3.2e-6), and the suite
passes. Against the build before, alternating runs, M5 Pro:
| point | before | after | |
|---|---|---|---|
| 1024 x 32x32 | 7.58 ms | 6.46 | 1.17x |
| 4096 x 16x16 | 5.87 | 5.09 | 1.15x |
| 256 x 48x48 | 5.52 | 4.86 | 1.14x |
| 64 x 128x128 | 33.7 | 32.5 | 1.03x |
| 16 x 256x256 | 66.1 | 66.2 | 1.00x |
| 1 x 512x512 | 262 | 249 | 1.06x |
(Fast math for the whole library had given 1.21x at 1024 x 32x32.)
The one-stage reductions’ 7%: gone already. Svd_Bidiag.metal built
entirely with fast math, against the current build: svdvals on bidiag
85.4 ms at 2048 either way (the 91.8 above was measured before the one-stage
reflectors’ div1), and no difference beyond 2% for bidiag with vectors,
band with or without; Eigh_Tridiag.metal likewise (eigvalsh on
tridiag 35.1 ms at 2048 against 35.4, eigh 54.3 against 53.3). Nothing
left there to find.
The eigensolver’s Jacobi kernels: skipped, as planned (within 1-2% with fast math).