metal-linalg routes every call by measurements made on each kind of Mac. This page lists every Apple Silicon chip and how current its measurements are, for each decomposition. It is regenerated from the submitted runs on every merge.
| state | what it means | |
|---|---|---|
| 🟢 | current | measured at the current kernels, every backend timed |
| 🟡 | incomplete | still valid, but a newer backend was never timed on this chip, so it stays off there |
| 🟠 | stale | measured on older kernels; still used, as the best available, until remeasured (runs from before 2.9.0, when the CPU path ran on one core, are out of date and not used: such a chip is estimated) |
| 🔴 | not measured | estimated from a measured Mac and published benchmarks (how): conservative, so it misses some GPU wins |
Have one of these Macs? One command measures it, in about an hour and a half, and a pull request submits it: how to contribute. Every run improves the library for everyone with that chip; the library prints a notice on chips that need one.
3 current, 0 incomplete, 0 stale and 111 not measured, over 38 chip configurations and 3 decompositions.
| chip | GPU cores | QR | eigh | SVD |
|---|---|---|---|---|
| M1 | ||||
| Apple M1 | 7 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 | 8 | 🔴 not measured estimated its runs predate 2.9.0: out of date | 🔴 not measured estimated its runs predate 2.9.0: out of date | 🔴 not measured estimated |
| Apple M1 Pro | 14 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 Pro | 16 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 Max | 24 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 Max | 32 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 Ultra | 48 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M1 Ultra | 64 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| M2 | ||||
| Apple M2 | 8 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 | 10 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Pro | 16 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Pro | 19 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Max | 30 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Max | 38 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Ultra | 60 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M2 Ultra | 76 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| M3 | ||||
| Apple M3 | 8 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 | 10 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Pro | 14 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Pro | 18 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Max | 30 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Max | 40 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Ultra | 60 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M3 Ultra | 80 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| M4 | ||||
| Apple M4 | 8 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M4 | 10 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M4 Pro | 16 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M4 Pro | 20 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M4 Max | 32 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M4 Max | 40 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| M5 | ||||
| Apple M5 | 8 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 | 10 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 Pro | 16 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 Pro | 20 | 🟢 current 1 run, latest 2026-10-07 measured with 2.16.0 | 🟢 current 1 run, latest 2026-10-07 measured with 2.16.0 | 🟢 current 1 run, latest 2026-10-07 measured with 2.16.0 |
| Apple M5 Max | 32 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 Max | 40 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 Ultra | 64 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
| Apple M5 Ultra | 80 | 🔴 not measured estimated | 🔴 not measured estimated | 🔴 not measured estimated |
Every run under docs/results/, and what each decomposition’s measurements in it are used for.
| run | chip | GPU cores | date | machine | macOS | library | QR | eigh | SVD |
|---|---|---|---|---|---|---|---|---|---|
| legacy | Apple M1 | 8 | not recorded | MacBook Pro (13-inch, M1, 2020) | ? | before 2.6 | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | — |
| 20261007-9f2589 | Apple M5 Pro | 20 | 2026-10-07 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.16.0 | used | used | used |
| 20261007-8633ac | Apple M5 Pro | 20 | 2026-10-07 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.16.0 | measured at kernel version 6; superseded | — | — |
| 20261007-82345e | Apple M5 Pro | 20 | 2026-10-07 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.15.0 | measured at kernel version 5; superseded | — | — |
| 20261007-246324 | Apple M5 Pro | 20 | 2026-10-07 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.15.0 | measured at kernel version 4; superseded | measured at kernel version 6; superseded | measured at kernel version 8; superseded |
| 20261004-fd9bd8 | Apple M5 Pro | 20 | 2026-10-04 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.12.0 | — | measured at kernel version 4; superseded | measured at kernel version 6; superseded |
| 20261004-4d6208 | Apple M5 Pro | 20 | 2026-10-04 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.12.0 | measured at kernel version 3; superseded | — | — |
| 20261004-06bc11 | Apple M5 Pro | 20 | 2026-10-04 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.13.0 | — | measured at kernel version 5; superseded | measured at kernel version 7; superseded |
| 20261003-d26059 | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | before 2.6 | — | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | — |
| 20261003-c0878c | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.11.0 | measured at kernel version 3; superseded | measured at kernel version 3; superseded | measured at kernel version 5; superseded |
| 20261003-af087d | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.10.0 | measured at kernel version 2; superseded | — | — |
| 20261003-847f0e | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.9.0 | measured at kernel version 2; superseded | — | — |
| 20261003-2d2c19 | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.9.0 | measured at kernel version 2; superseded | measured at kernel version 2; superseded | measured at kernel version 3; superseded |
| 20261003-106b6c | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.10.0 | — | — | measured at kernel version 4; superseded |
| 20261003-064803 | Apple M5 Pro | 20 | 2026-10-03 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | 2.7.0 | — | — | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated |
| 20261002-9d19ba | Apple M5 Pro | 20 | 2026-10-02 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | before 2.6 | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | — | — |
| 20261002-153352 | Apple M5 Pro | 20 | 2026-10-02 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | before 2.6 | — | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | — |
| 20261001-c76e82 | Apple M5 Pro | 20 | 2026-10-01 | MacBook Pro (16-inch, M5 Pro) | 27.0.1 | before 2.6 | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated |
| 20260930-27b6c2 | Apple M5 Pro | 20 | 2026-09-30 | MacBook Pro (16-inch, M5 Pro) | 26.6 | before 2.6 | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated | measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated |
A decomposition’s measurements go stale only when something they time changes: its kernels, their launch parameters, its CPU path, or what it calls. Releases that change nothing of the kind keep every measurement current. Each change that does bumps that decomposition’s kernel version (tuning/kernels.py):
| decomposition | version | since | date | why |
|---|---|---|---|---|
| QR | 1 | 2.0.0 | 2026-09-30 | the measurements as 2.0 introduced them |
| eigh | 1 | 2.0.0 | 2026-09-30 | the measurements as 2.0 introduced them |
| SVD | 1 | 2.0.0 | 2026-09-30 | the measurements as 2.0 introduced them |
| SVD | 2 | 2.2.3 | 2026-10-02 | QR, which the QR-preconditioned SVD backends call first, routes small problems to the CPU since its CPU boundary was measured; those backends’ timings have changed |
| QR | 2 | 2.9.0 | 2026-10-03 | the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved |
| eigh | 2 | 2.9.0 | 2026-10-03 | the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved |
| SVD | 3 | 2.9.0 | 2026-10-03 | the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved |
| SVD | 4 | 2.10.0 | 2026-10-03 | QR keeps the smallest matrices on the CPU at any batch (gpu_min_k), which moves the timings of the QR-preconditioned backends |
| QR | 3 | 2.11.0 | 2026-10-03 | the CPU path factors a wide matrix by its leading square block and one matrix product, 10-40x faster than sgeqrf on the whole matrix |
| eigh | 3 | 2.11.0 | 2026-10-03 | the tridiag backend pipelines a batch over two slots (CPU solve of one matrix while the GPU reduces the next), 1.4-1.5x per matrix for batches of 2048 x 2048 |
| SVD | 5 | 2.11.0 | 2026-10-03 | golub_kahan splits its column sums over lanes (1.2-1.4x on tall matrices) and runs the QR iteration as a second dispatch from k = 40 (singular values alone) or 60; the bidiag backend pipelines a batch over two slots (1.5-1.7x per matrix for batches of 2048 x 2048) |
| eigh | 4 | 2.12.0 | 2026-10-04 | the tridiag backend’s reduction takes three dispatches per column instead of seven, and copies the matrix in on every core: 1.3-1.7x for eigenvalues alone, 1.2-1.5x with vectors |
| SVD | 6 | 2.12.0 | 2026-10-04 | the bidiag backend’s reduction takes four dispatches per column instead of twelve, and copies the matrix in on every core: 1.2-1.8x for singular values alone, 1.1-1.3x with vectors |
| eigh | 5 | 2.13.0 | 2026-10-04 | the tridiag backend’s eigenvalues alone come from bisection on the GPU rather than ssterf from N = 512: 6 ms against 79 at 4096 |
| SVD | 7 | 2.13.0 | 2026-10-04 | the bidiag backend’s singular values alone come from bisection on the GPU from k = 1024, and below that from sbdsqr (dqds) rather than sbdsdc: 12 ms against 93 at 4096 |
| eigh | 6 | 2.15.0 | 2026-10-07 | the tridiag backend’s eigenvectors come from a divide and conquer on every core (sstedc’s 176 ms to 50 at 4096); the band backend’s panels are faster (the TSQR top a tree, no IEEE division), its small products are kernels of their own and its trailing update is on the lower triangle: eigh with vectors 1.3-1.5x and eigvalsh 1.1-1.2x at 2048-8192 |
| SVD | 8 | 2.15.0 | 2026-10-07 | the bidiag backend’s singular vectors come from a divide and conquer on every core (sbdsdc’s 753 ms to 100 at 4096): the SVD with vectors 1.7-1.9x at 2048-4096; the band backend’s panels and small products are faster: svdvals 1.1-1.3x; and the band backend takes singular vectors too (band_min_k): 2.3x bidiag at 4096 |
| QR | 4 | 2.15.0 | 2026-10-07 | the reduced backend hands one matrix, or a few large ones, to the blocked QR (the band reduction’s panels, aggregates of 128 columns, MPS products): 2.1x at 1024, 2.6x at 2048, 3.7x at 4096, 5x on tall 4096 x 1024 and 8192 x 512 |
| QR | 5 | 2.15.0 | 2026-10-07 | the blocked QR takes a batch at once and any height (padded to whole panels, batched MPS products): every shape the reduced backend’s streaming kernels took, 1.8-3.5x faster; batches of 512-2048 now beat the CPU (16 x 1024^2: 25 ms against 48); the large clause counts rows and k, sqrt(M k), and the grid has tall large shapes |
| QR | 6 | 2.16.0 | 2026-10-08 | the unblocked backend is new Householder kernels (in a simdgroup’s registers up to 32 x 128, else blocked in a threadgroup up to 4096 rows, its updates 8 x 8 simdgroup matrix products; its own kernel retired): 1024 of 128 x 128 in 6.4 ms against 17.8 for the blocked QR and 25 on the CPU, 4096 of 32 x 32 in 1.4 against 2.9 on the CPU; MLX’s own buffers no longer wrapped again; the grid’s kernel crossover goes down to 64 rows and its mid-size batches up to 384 |
| QR | 7 (current) | 2.16.0 | 2026-10-08 | the sweeps keep MLX’s buffer cache on, as an MLX program has it (off, every call’s outputs were fresh pages the GPU maps at about 12 us a MB: 4096 of 64 x 64 6.4 ms on the GPU against 4.0 with it on); the blocked kernel gives a small batch more simdgroups a matrix (one 384 x 384 2.0 ms against 3.0) and reads an aligned input directly; the GPU-or-CPU rule is on sqrt(M k) |
| eigh | 7 (current) | 2.16.0 | 2026-10-08 | the sweeps keep MLX’s buffer cache on, as an MLX program has it, and MLX’s own buffers are no longer wrapped again: the GPU backends’ calls on large batches 10-40% cheaper |
| SVD | 9 (current) | 2.16.0 | 2026-10-08 | the sweeps keep MLX’s buffer cache on, as an MLX program has it, and MLX’s own buffers are no longer wrapped again: the GPU backends’ calls on large batches 10-40% cheaper; the QR-preconditioned backends run on 2.16.0’s QR |