metal-linalg

Measuring your Mac

metal-linalg chooses the fastest kernel for each problem, or the CPU, from measurements made on each kind of Mac. If yours isn’t in the table below, or you’d like to add to its measurements, it takes one command and about an hour and a half of the Mac’s time. Any number of people with the same Mac can contribute: each run is saved under its own ID, and the runs are combined.

| Mac | GPU cores | QR | eigh | SVD | runs | |—|—|—|—|—|—| | Apple M1 | 8 | out of date | out of date | — | 1 | | Apple M5 Pro | 20 | measured | measured | measured | 18 |

How to measure and send the results is in CONTRIBUTING.md: one command, python3 tuning/run.py, then a pull request. This page is the reference behind it: what a run records, what to do if it stops with a problem, and how the runs become the library’s settings.

Which measurements are current

The measurements page lists every Apple Silicon chip, colour-coded per decomposition:

A library release does not by itself make measurements stale: most change nothing they time. What does is a change to a decomposition’s kernels, their launch parameters, its CPU path or what it calls, and each such change bumps that decomposition’s kernel version in tuning/kernels.py, with a line saying why. Every run records the versions it measured, and the tables use the runs at the current version, or else the newest older ones rather than none.

The library says so too. On a Mac without measurements for a decomposition, or with stale or incomplete ones, it prints one line, once per decomposition per process, with a link to these instructions (Python raises it as a metal_linalg.CalibrationWarning at import instead). The policy sources say the same: estimated:<chip> (from <measured chip>, ...), tuned-stale:<chip> or tuned-incomplete:<chip>. METAL_LINALG_NO_CALIBRATION_NOTICE=1 turns the notice off.

Macs nobody has measured

A Mac without measurements is routed by an estimate. The library takes a Mac that has been measured at the current kernels (the M5 Pro so far) and uses its timings refitted, by the same analysis that turns a run into settings, as if its GPU were slower against its CPU by as much as the two Macs differ:

On Macs simulated from the M5 Pro’s timings with a GPU 1 to 4 times weaker, the estimates come within 0-3% of the best routing on average, where the untuned default they replace was 1.3-3x slower; the study has the method, the tables, and every Mac’s estimate. The estimated rows are refitted from the measurements whenever those change (tuning/generate_tables.py, by the same Action), so each new measurement also improves the estimates for its neighbours.

METAL_LINALG_ESTIMATE_AS="<chip>[:<GPU cores>[:<CPU cores>]] estimates as if the library ran on that Mac, even on a measured one: METAL_LINALG_ESTIMATE_AS="Apple M4 Max:40" with any program shows, in its policy sources, what an M4 Max gets.

What is recorded

The exact model of Mac (for example “MacBook Pro (16-inch, M5 Pro)”, since the same chip runs at different speeds in machines that cool it differently), its chip, CPU and GPU core counts, memory and built-in display, the macOS and MLX versions, and, at the start and after each part of the run, the load average, the power source and charger, the power mode and any thermal warnings. And the timings. Nothing that identifies you or your particular Mac: no names, hostnames or serial numbers.

If something goes wrong

what you see what to do
“This Mac is not ready to measure” plug in, turn Low Power Mode off, quit other apps, and run it again
“CMake is not installed” / “MLX is not installed” brew install cmake mlx
a correctness test failed please open an issue with the log file it names; don’t send timings
some results “NOT trustworthy” something interfered (another app, a thermal change); run it again with the Mac idle. Untrustworthy results are never used
you stopped it part-way delete the incomplete docs/results/<your Mac>/<ID>/ folder and run it again

python3 tuning/run.py --quick is a 15-minute smoke test of the whole pipeline. Its results are written to build-tuning/quick/ and are not for sending.

python3 tuning/run.py --only qr (or eigh, svd, or a comma-separated list) measures only those decompositions: QR in about 3 minutes. It is for remeasuring one after its routing or kernels change; the settings of the others keep coming from earlier runs of the same Mac. The submission is sent like any other.


Maintainers: a GitHub Action (.github/workflows/tuned-policies.yml) checks each results pull request, shows the effect on the tables, and after the merge regenerates src/tuned/ from every run and commits it. Nothing needs doing by hand. How the measurements, the combining and the Action work is in tuning-details.md; what every number in a report means is in reading-reports.md.