metal-linalg chooses the fastest kernel for each problem, or the CPU, from measurements made on each kind of Mac. If yours isn’t in the table below, or you’d like to add to its measurements, it takes one command and about an hour and a half of the Mac’s time. Any number of people with the same Mac can contribute: each run is saved under its own ID, and the runs are combined.
| Mac | GPU cores | QR | eigh | SVD | runs | |—|—|—|—|—|—| | Apple M1 | 8 | out of date | out of date | — | 1 | | Apple M5 Pro | 20 | measured | measured | measured | 18 |
How to measure and send the results is in
CONTRIBUTING.md: one command,
python3 tuning/run.py, then a pull request. This page is the reference
behind it: what a run records, what to do if it stops with a problem, and
how the runs become the library’s settings.
The measurements page lists every Apple Silicon chip, colour-coded per decomposition:
A library release does not by itself make measurements stale: most change
nothing they time. What does is a change to a decomposition’s kernels, their
launch parameters, its CPU path or what it calls, and each such change bumps
that decomposition’s kernel version in
tuning/kernels.py, with a line saying why. Every
run records the versions it measured, and the tables use the runs at the
current version, or else the newest older ones rather than none.
The library says so too. On a Mac without measurements for a decomposition,
or with stale or incomplete ones, it prints one line, once per decomposition
per process, with a link to these instructions (Python raises it as a
metal_linalg.CalibrationWarning at import instead). The policy sources say
the same: estimated:<chip> (from <measured chip>, ...), tuned-stale:<chip>
or tuned-incomplete:<chip>. METAL_LINALG_NO_CALIBRATION_NOTICE=1 turns the
notice off.
A Mac without measurements is routed by an estimate. The library takes a Mac that has been measured at the current kernels (the M5 Pro so far) and uses its timings refitted, by the same analysis that turns a run into settings, as if its GPU were slower against its CPU by as much as the two Macs differ:
tuning/chip_specs.py
lists every Mac’s);On Macs simulated from the M5 Pro’s timings with a GPU 1 to 4 times weaker,
the estimates come within 0-3% of the best routing on average, where the
untuned default they replace was 1.3-3x slower; the
study has the method, the tables, and every
Mac’s estimate. The estimated rows are refitted from the measurements whenever
those change (tuning/generate_tables.py, by the same Action), so each new
measurement also improves the estimates for its neighbours.
METAL_LINALG_ESTIMATE_AS="<chip>[:<GPU cores>[:<CPU cores>]] estimates as if
the library ran on that Mac, even on a measured one: METAL_LINALG_ESTIMATE_AS="Apple
M4 Max:40" with any program shows, in its policy sources, what an M4 Max gets.
The exact model of Mac (for example “MacBook Pro (16-inch, M5 Pro)”, since the same chip runs at different speeds in machines that cool it differently), its chip, CPU and GPU core counts, memory and built-in display, the macOS and MLX versions, and, at the start and after each part of the run, the load average, the power source and charger, the power mode and any thermal warnings. And the timings. Nothing that identifies you or your particular Mac: no names, hostnames or serial numbers.
| what you see | what to do |
|---|---|
| “This Mac is not ready to measure” | plug in, turn Low Power Mode off, quit other apps, and run it again |
| “CMake is not installed” / “MLX is not installed” | brew install cmake mlx |
| a correctness test failed | please open an issue with the log file it names; don’t send timings |
| some results “NOT trustworthy” | something interfered (another app, a thermal change); run it again with the Mac idle. Untrustworthy results are never used |
| you stopped it part-way | delete the incomplete docs/results/<your Mac>/<ID>/ folder and run it again |
python3 tuning/run.py --quick is a 15-minute smoke test of the whole
pipeline. Its results are written to build-tuning/quick/ and are not for
sending.
python3 tuning/run.py --only qr (or eigh, svd, or a comma-separated
list) measures only those decompositions: QR in about 3 minutes. It is for
remeasuring one after its routing or kernels change; the settings of the
others keep coming from earlier runs of the same Mac. The submission is sent
like any other.
Maintainers: a GitHub Action (.github/workflows/tuned-policies.yml)
checks each results pull request, shows the effect on the tables, and after
the merge regenerates src/tuned/ from every run and commits it. Nothing
needs doing by hand. How the measurements, the combining and the Action work
is in tuning-details.md; what every number in a report
means is in reading-reports.md.