metal-linalg

Measurements

metal-linalg routes every call by measurements made on each kind of Mac. This page lists every Apple Silicon chip and how current its measurements are, for each decomposition. It is regenerated from the submitted runs on every merge.

  state what it means
🟢 current measured at the current kernels, every backend timed
🟡 incomplete still valid, but a newer backend was never timed on this chip, so it stays off there
🟠 stale measured on older kernels; still used, as the best available, until remeasured (runs from before 2.9.0, when the CPU path ran on one core, are out of date and not used: such a chip is estimated)
🔴 not measured estimated from a measured Mac and published benchmarks (how): conservative, so it misses some GPU wins

Have one of these Macs? One command measures it, in about an hour and a half, and a pull request submits it: how to contribute. Every run improves the library for everyone with that chip; the library prints a notice on chips that need one.

3 current, 0 incomplete, 0 stale and 111 not measured, over 38 chip configurations and 3 decompositions.

By chip

chipGPU coresQReighSVD
M1
Apple M17🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M18🔴 not measured
estimated
its runs predate 2.9.0: out of date
🔴 not measured
estimated
its runs predate 2.9.0: out of date
🔴 not measured
estimated
Apple M1 Pro14🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M1 Pro16🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M1 Max24🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M1 Max32🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M1 Ultra48🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M1 Ultra64🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
M2
Apple M28🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M210🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Pro16🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Pro19🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Max30🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Max38🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Ultra60🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M2 Ultra76🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
M3
Apple M38🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M310🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Pro14🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Pro18🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Max30🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Max40🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Ultra60🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M3 Ultra80🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
M4
Apple M48🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M410🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M4 Pro16🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M4 Pro20🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M4 Max32🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M4 Max40🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
M5
Apple M58🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M510🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M5 Pro16🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M5 Pro20🟢 current
1 run, latest 2026-10-07
measured with 2.16.0
🟢 current
1 run, latest 2026-10-07
measured with 2.16.0
🟢 current
1 run, latest 2026-10-07
measured with 2.16.0
Apple M5 Max32🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M5 Max40🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M5 Ultra64🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated
Apple M5 Ultra80🔴 not measured
estimated
🔴 not measured
estimated
🔴 not measured
estimated

Submitted runs

Every run under docs/results/, and what each decomposition’s measurements in it are used for.

run chip GPU cores date machine macOS library QR eigh SVD
legacy Apple M1 8 not recorded MacBook Pro (13-inch, M1, 2020) ? before 2.6 measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated —
20261007-9f2589 Apple M5 Pro 20 2026-10-07 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.16.0 used used used
20261007-8633ac Apple M5 Pro 20 2026-10-07 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.16.0 measured at kernel version 6; superseded — —
20261007-82345e Apple M5 Pro 20 2026-10-07 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.15.0 measured at kernel version 5; superseded — —
20261007-246324 Apple M5 Pro 20 2026-10-07 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.15.0 measured at kernel version 4; superseded measured at kernel version 6; superseded measured at kernel version 8; superseded
20261004-fd9bd8 Apple M5 Pro 20 2026-10-04 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.12.0 — measured at kernel version 4; superseded measured at kernel version 6; superseded
20261004-4d6208 Apple M5 Pro 20 2026-10-04 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.12.0 measured at kernel version 3; superseded — —
20261004-06bc11 Apple M5 Pro 20 2026-10-04 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.13.0 — measured at kernel version 5; superseded measured at kernel version 7; superseded
20261003-d26059 Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 before 2.6 — measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated —
20261003-c0878c Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.11.0 measured at kernel version 3; superseded measured at kernel version 3; superseded measured at kernel version 5; superseded
20261003-af087d Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.10.0 measured at kernel version 2; superseded — —
20261003-847f0e Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.9.0 measured at kernel version 2; superseded — —
20261003-2d2c19 Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.9.0 measured at kernel version 2; superseded measured at kernel version 2; superseded measured at kernel version 3; superseded
20261003-106b6c Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.10.0 — — measured at kernel version 4; superseded
20261003-064803 Apple M5 Pro 20 2026-10-03 MacBook Pro (16-inch, M5 Pro) 27.0.1 2.7.0 — — measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated
20261002-9d19ba Apple M5 Pro 20 2026-10-02 MacBook Pro (16-inch, M5 Pro) 27.0.1 before 2.6 measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated — —
20261002-153352 Apple M5 Pro 20 2026-10-02 MacBook Pro (16-inch, M5 Pro) 27.0.1 before 2.6 — measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated —
20261001-c76e82 Apple M5 Pro 20 2026-10-01 MacBook Pro (16-inch, M5 Pro) 27.0.1 before 2.6 measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated
20260930-27b6c2 Apple M5 Pro 20 2026-09-30 MacBook Pro (16-inch, M5 Pro) 26.6 before 2.6 measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated measured before 2.9.0, when the CPU path ran on one core: out of date, so this Mac is estimated

Kernel versions

A decomposition’s measurements go stale only when something they time changes: its kernels, their launch parameters, its CPU path, or what it calls. Releases that change nothing of the kind keep every measurement current. Each change that does bumps that decomposition’s kernel version (tuning/kernels.py):

decomposition version since date why
QR 1 2.0.0 2026-09-30 the measurements as 2.0 introduced them
eigh 1 2.0.0 2026-09-30 the measurements as 2.0 introduced them
SVD 1 2.0.0 2026-09-30 the measurements as 2.0 introduced them
SVD 2 2.2.3 2026-10-02 QR, which the QR-preconditioned SVD backends call first, routes small problems to the CPU since its CPU boundary was measured; those backends’ timings have changed
QR 2 2.9.0 2026-10-03 the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved
eigh 2 2.9.0 2026-10-03 the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved
SVD 3 2.9.0 2026-10-03 the CPU path spreads a batch over every core (7.5-15x faster for batches of small matrices on an M5 Pro), so every GPU-or-CPU boundary has moved
SVD 4 2.10.0 2026-10-03 QR keeps the smallest matrices on the CPU at any batch (gpu_min_k), which moves the timings of the QR-preconditioned backends
QR 3 2.11.0 2026-10-03 the CPU path factors a wide matrix by its leading square block and one matrix product, 10-40x faster than sgeqrf on the whole matrix
eigh 3 2.11.0 2026-10-03 the tridiag backend pipelines a batch over two slots (CPU solve of one matrix while the GPU reduces the next), 1.4-1.5x per matrix for batches of 2048 x 2048
SVD 5 2.11.0 2026-10-03 golub_kahan splits its column sums over lanes (1.2-1.4x on tall matrices) and runs the QR iteration as a second dispatch from k = 40 (singular values alone) or 60; the bidiag backend pipelines a batch over two slots (1.5-1.7x per matrix for batches of 2048 x 2048)
eigh 4 2.12.0 2026-10-04 the tridiag backend’s reduction takes three dispatches per column instead of seven, and copies the matrix in on every core: 1.3-1.7x for eigenvalues alone, 1.2-1.5x with vectors
SVD 6 2.12.0 2026-10-04 the bidiag backend’s reduction takes four dispatches per column instead of twelve, and copies the matrix in on every core: 1.2-1.8x for singular values alone, 1.1-1.3x with vectors
eigh 5 2.13.0 2026-10-04 the tridiag backend’s eigenvalues alone come from bisection on the GPU rather than ssterf from N = 512: 6 ms against 79 at 4096
SVD 7 2.13.0 2026-10-04 the bidiag backend’s singular values alone come from bisection on the GPU from k = 1024, and below that from sbdsqr (dqds) rather than sbdsdc: 12 ms against 93 at 4096
eigh 6 2.15.0 2026-10-07 the tridiag backend’s eigenvectors come from a divide and conquer on every core (sstedc’s 176 ms to 50 at 4096); the band backend’s panels are faster (the TSQR top a tree, no IEEE division), its small products are kernels of their own and its trailing update is on the lower triangle: eigh with vectors 1.3-1.5x and eigvalsh 1.1-1.2x at 2048-8192
SVD 8 2.15.0 2026-10-07 the bidiag backend’s singular vectors come from a divide and conquer on every core (sbdsdc’s 753 ms to 100 at 4096): the SVD with vectors 1.7-1.9x at 2048-4096; the band backend’s panels and small products are faster: svdvals 1.1-1.3x; and the band backend takes singular vectors too (band_min_k): 2.3x bidiag at 4096
QR 4 2.15.0 2026-10-07 the reduced backend hands one matrix, or a few large ones, to the blocked QR (the band reduction’s panels, aggregates of 128 columns, MPS products): 2.1x at 1024, 2.6x at 2048, 3.7x at 4096, 5x on tall 4096 x 1024 and 8192 x 512
QR 5 2.15.0 2026-10-07 the blocked QR takes a batch at once and any height (padded to whole panels, batched MPS products): every shape the reduced backend’s streaming kernels took, 1.8-3.5x faster; batches of 512-2048 now beat the CPU (16 x 1024^2: 25 ms against 48); the large clause counts rows and k, sqrt(M k), and the grid has tall large shapes
QR 6 2.16.0 2026-10-08 the unblocked backend is new Householder kernels (in a simdgroup’s registers up to 32 x 128, else blocked in a threadgroup up to 4096 rows, its updates 8 x 8 simdgroup matrix products; its own kernel retired): 1024 of 128 x 128 in 6.4 ms against 17.8 for the blocked QR and 25 on the CPU, 4096 of 32 x 32 in 1.4 against 2.9 on the CPU; MLX’s own buffers no longer wrapped again; the grid’s kernel crossover goes down to 64 rows and its mid-size batches up to 384
QR 7 (current) 2.16.0 2026-10-08 the sweeps keep MLX’s buffer cache on, as an MLX program has it (off, every call’s outputs were fresh pages the GPU maps at about 12 us a MB: 4096 of 64 x 64 6.4 ms on the GPU against 4.0 with it on); the blocked kernel gives a small batch more simdgroups a matrix (one 384 x 384 2.0 ms against 3.0) and reads an aligned input directly; the GPU-or-CPU rule is on sqrt(M k)
eigh 7 (current) 2.16.0 2026-10-08 the sweeps keep MLX’s buffer cache on, as an MLX program has it, and MLX’s own buffers are no longer wrapped again: the GPU backends’ calls on large batches 10-40% cheaper
SVD 9 (current) 2.16.0 2026-10-08 the sweeps keep MLX’s buffer cache on, as an MLX program has it, and MLX’s own buffers are no longer wrapped again: the GPU backends’ calls on large batches 10-40% cheaper; the QR-preconditioned backends run on 2.16.0’s QR