How the measurements behind the per-device routing work, for maintainers and
the curious. Contributors only need tuning.md.
tuning/run.py doescmake and MLX installed, on mains power,
load average under half the CPU count, Low Power Mode off. It stops if any
check fails; --anyway measures regardless and the results are marked
untrustworthy.build-tuning/.test_qr, test_eigh and test_svd, and stops if any fails: timings
of a wrong answer mean nothing.Creates docs/results/<device>/<id>/ and runs the three harnesses into it,
one after another:
| step | harness | measures | time on an M5 Pro |
|---|---|---|---|
| QR | tuning/tune_qr.py |
both GPU backends and the CPU on 185 shapes, square, tall, wide and near-square, batch 1 to 16384 | 4 min |
| eigh | tuning/tune_eigh.py --max-n 4096 |
CPU, the whole-matrix kernel in both modes, block Jacobi, ql, tridiag, each again for eigenvalues alone; N from 2 to 4096, batch 1 to 4096 (small batches above 1024) | 25 min |
| SVD | tuning/tune_svd.py --max-k 4096 |
CPU, both Jacobi kernels, each with and without QR, square k up to 1024 and tall shapes, batch 1 to 4096; golub_kahan up to its limit; bidiag and the CPU, with and without vectors, from k = 128 up to 4096 (small batches above 1024) | 50 min |
Each runs two passes in random order, so thermal drift is not mistaken for a size effect and a noise floor can be measured.
submission.json and summary.md, and prints the rows and how to
send them.--quick runs coarser grids, one pass each, into build-tuning/quick/: a
smoke test of the pipeline, never a submission.
Memory. Every sweep skips a shape whose estimated peak memory, six copies
of its arrays in float32 (input, outputs, workspaces, the correctness check),
exceeds 35% of the Mac’s RAM: about 2.8 GB on an 8 GB Mac, 17 GB on a 48 GB
one (tuning/submissions.py, MEMORY_FRACTION). A shape that swaps would
time the disk, or stop the run. A smaller Mac therefore measures a smaller
grid; the fitted thresholds are chosen among the shapes measured, as always.
METAL_LINALG_TUNING_MEMORY_GB=<n> sets the budget instead.
docs/results/<device>/<id>/
submission.json the spec, versions, conditions and fitted rows
summary.md the same, readable
qr/ eigh/ svd/ per decomposition: raw.csv, results.json, report.md, policy.json
qr.log eigh.log svd.log
<device> is the chip and GPU core count, e.g. apple-m5-pro-20gpu: the
same pair kTuned[] is keyed on, so a binned part with fewer cores is its own
device. <id> is the UTC date and six random hex digits, e.g.
20260930-27b6c2, unique without any coordination between contributors.
submission.json records:
| field | contents |
|---|---|
device |
the chip and GPU core count, as Metal reports them; the key for kTuned[] |
machine |
the exact model: product name (“MacBook Pro (16-inch, M5 Pro)”), model identifier (Mac17,8) and built-in display resolution. One chip ships in machines that cool it differently, a 14-inch and a 16-inch MacBook Pro for instance, and only this tells them apart |
cpu, memory_gb |
CPU cores per performance level, memory |
macos, macos_build, mlx, metal_linalg, epoch |
the software it ran |
conditions |
at the start and after each decomposition: load average, CPU count, power source, charger and its wattage, battery charge, power mode, Low Power Mode, and macOS’s thermal and performance warnings |
results |
per decomposition: trustworthy or not and why, the fitted row, and the probe point’s drift over the sweep |
minutes |
how long each decomposition took |
It records nothing that identifies the person or the particular Mac: no serial numbers, hardware UUIDs, hostnames or user names.
The M1’s runs predate this layout and are kept as apple-m1-8gpu/legacy/
(QR and eigh only).
The rows the library uses are generated, never edited by hand:
docs/results/<device>/<id>/ every run submitted (any number per device)
| tuning/combine.py one device's runs -> its rows
v
docs/results/<device>/combined/ the combined report per decomposition, summary.md
| tuning/generate_tables.py every device
v
src/tuned/{qr,eigh,svd}.inc included by kTuned[] in src/qr.mm, src/eigh.mm, src/svd.mm
What a timing is. Each sweep tool times one shape in a process of its own, a call through the MLX API as a program makes it, the median of its reps after two warm-ups, with MLX’s buffer cache on as MLX has it by default (since 2.16.0; off before, which made every call’s outputs fresh pages for the GPU to map, and timed the allocation as much as the decomposition).
How runs combine (tuning/submissions.py): for each decomposition,
every run whose measurements of it are trustworthy is used, and smoke tests,
interrupted runs, untrustworthy results and runs from an older measurement
epoch are skipped (the combined summary.md says which and why). Within a
run, each backend’s time at a point is the fastest of its passes, since
interference only ever slows a run down; across runs, the median of those, so
one unusual machine cannot move the result. The analysis is then the same as
for a single run, on the combined times. The noise floor stays per run, so it
measures run-to-run noise rather than the spread between machines.
The epoch (EPOCH in tuning/submissions.py, recorded in every
submission.json). When a kernel or its launch parameters change enough that
earlier timings no longer describe the library, bump it: older runs then stop
counting, and a device is estimated from the measured ones (and its own
measurements stop being an anchor for the others) until it is measured again.
Diagnosing disagreement. The combined summary.md for each device ends
with a table of its runs: each run’s machine, memory, macOS and conditions,
and for each decomposition whether the settings that run fitted on its own
match the combined ones, naming every setting that differs. The Action’s
summary shows the same table on every results pull request. Runs from one
model that consistently disagree with another model of the same chip point to
a real difference between the machines rather than noise: a 14-inch MacBook
Pro that throttles where the 16-inch does not, for instance. The tables are
keyed on the chip and GPU core count only, so such machines share a row
today; if the difference is real, the key would have to include the model
identifier, which the library can read at run time (sysctl hw.model).
The Action (.github/workflows/tuned-policies.yml) runs on every pull
request and every push to main that touches docs/results/ or tuning/,
on a Linux runner: the analysis is plain Python and needs no GPU.
tuning/validate_submissions.py (the folder and
file layout, the ID, that the device folder matches the spec, the CSV
columns and values, no smoke tests, no oversized files), then
tuning/generate_tables.py, and lists the resulting rows and the change to
src/tuned/ in the run’s summary. A malformed submission fails the check.main it does the same and commits src/tuned/ and the
combined reports if they changed, as github-actions[bot]. If main is
protected against direct pushes, allow the Action to push or change this
step to open a pull request.By hand, the same steps are:
python3 tuning/validate_submissions.py
python3 tuning/generate_tables.py # rewrites src/tuned/ and the combined reports
python3 tuning/generate_tables.py --check # changes nothing; fails if they are out of date
python3 tuning/combine.py docs/results/apple-m5-pro-20gpu # one device, to look at
After the tables change, cmake --build build && ctest --test-dir build and
./build/sweep_qr --policy (or sweep_eigh, sweep_svd) on a measured Mac
should report tuned:<device name>.
QR (src/qr.mm): device name, GPU cores, the row crossover for small
batches, the row crossover for large batches, and the batch that separates
them. The two crossovers are equal unless a batch-dependent split survived
held-out validation. Then the CPU boundary, gpu_max_k,
gpu_min_batch_times_k, gpu_min_batch and gpu_min_k (the GPU only from
this k, so that the smallest matrices stay on the CPU at any batch; 0 in a run
from before 2.10.0), the large-matrix clause,
gpu_large_min_k and gpu_large_max_batch: the GPU also from the first in
a batch of at most the second (0: any; 0, 0: never, which a run from
before 2.9.0 gives), the size being k before 2.15.0 and sqrt(M k) since,
so that a tall matrix counts by its rows too. Since the CPU path spreads a batch over every core
it wins batches of small and mid-size matrices, while one large matrix is
still faster on the GPU, and one product rule cannot say both. Last,
share_min_batch: from this batch a GPU batch is shared with the CPU path
(0: never, which a run from before 2.12.0 gives), fitted on the share
timings before the GPU-or-CPU boundary, which is then fitted with it in
effect.
Eigensolver (src/eigh.mm): device name, GPU cores, then simd_max_n,
block_min_n, block_min_n_batched, block_min_batch, gpu_max_n,
gpu_min_batch_times_n, gpu_min_batch, then values_gpu_max_n,
values_gpu_min_batch_times_n, values_gpu_min_batch, then tridiag_min_n and
values_tridiag_min_n, tridiag_max_batch and values_tridiag_max_batch,
then ql_min_n and ql_max_n, as documented on
EighPolicy in include/metal_linalg/core.h. The first four choose the GPU
backend, fitted against the best GPU backend alone; the next three are the CPU
boundary: GPU iff N is at most gpu_max_n, batch × N at least
gpu_min_batch_times_n and the batch at least gpu_min_batch. The last
three are the same boundary for eigenvalues alone (eigvalsh), fitted on the
_vals timings; 0, 0, 0 (a run from before those were measured) means “as
for eigenvectors”. The last two are where the tridiag backend takes over from
the CPU, with eigenvectors and without (0: never, which a run from before the
backend existed gives), for batches up to the two caps after them (0: any
batch; the backend solves a batch one matrix after another, the CPU path
spreads one over every core); stage 4 of tune_eigh.py fits each threshold
with its cap on the points where tridiag was timed (the backend pipelines a
batch over two slots since 2.11.0, so the caps grew). Then the window of N in
which the ql backend replaces the Jacobi backend the first four pick (0, 0:
never, which a run from before 2.9.0 gives); stage 1b fits it over the finished
split, against the best GPU backend, on the points where ql was timed (N up
to the device’s limit, 87 with 32 KB of threadgroup memory). The very last is
share_min_batch: from this batch, a batch that goes to ql is shared with
the CPU path, the GPU and the CPU solving it at once (0: never, which a run
from before 2.11.0 gives); stage 1c fits it against the best GPU backend, the
shared one (ql_share, timed from batch 64) included. After it,
gpu_big_batch_max_n and gpu_big_batch_min: the GPU also for N above
gpu_max_n up to the first in a batch of at least the second (0, 0: never,
which a run from before 2.12.0 gives), fitted in stage 2 together with the
product rule: for each gpu_max_n, the rule fitted alone, the clause over it,
and the rule fitted again given the clause; the best combination is kept.
Last, values_band_min_n: for eigenvalues alone, from this N the band
backend (the two-stage reduction) instead of tridiag or the CPU, within
values_tridiag_max_batch (0: never, which a run from before 2.13.0 gives);
stage 4b fits it, after tridiag’s thresholds, on the points where
band_vals was timed (N >= 512). Then values_band_width, the band’s width
(8, 16 or 32; 0: 16, which a run from before 2.15.0 gives): stage 4b chooses
it first, from band8_vals, band_vals and band32_vals at those points,
and fits the threshold on its times.
SVD (src/svd.mm): device name, GPU cores, then qr_min_rows,
qr_min_k, block_min_k, block_min_k_batched, block_min_batch,
gpu_max_k, gpu_min_batch_times_k, gpu_min_batch and gpu_max_l, then
the same four for singular values alone (values_gpu_max_k,
values_gpu_min_batch_times_k, values_gpu_min_batch, values_gpu_max_l;
values_gpu_min_batch = 0, which a run from before 2.11.0 gives, means “as
with vectors”), then bidiag_min_k, values_bidiag_min_k, bidiag_max_batch
and values_bidiag_max_batch, then gk_min_k and gk_max_k, then
share_min_batch, as documented on SvdPolicy in
include/metal_linalg/core.h: whether to precondition with QR, which kernel,
then the CPU boundary in the eigensolver’s form (with a cap on the long side,
l = max(M, N)), and again for svdvals (stage 2b of tune_svd.py, fitted on
gk_vals and cpu_vals where gk is timed), then where the bidiag
backend takes over from the CPU, with vectors and for singular values alone
(0: never, which a run from before the backend existed gives), and up to which
batch (0: any; as tridiag, it solves a batch one matrix after another);
stage 3 of tune_svd.py fits each threshold with its cap on the points where
bidiag was timed. Last, gk_min_k and gk_max_k: the window of k in which
the golub_kahan backend replaces the Jacobi backends on the GPU (0, 0: never,
which a run from before 2.10.0 gives). Stage 1b of tune_svd.py fits it over
the Jacobi split, against the best GPU backend, on the points where gk was
timed (k up to the device’s limit, 83 with 32 KB of threadgroup memory), as
tune_eigh.py fits the ql window; the CPU boundary is then fitted with it
in place. And share_min_batch, as for the eigensolver: from this batch a
golub_kahan batch is shared with the CPU path (stage 1c, on gk_share),
and gpu_big_batch_max_k and gpu_big_batch_min, the large-batch clause, as
for the eigensolver. Last, values_band_min_k: for singular values alone,
from this k the band backend (the two-stage reduction) instead of bidiag
or the CPU, within values_bidiag_max_batch (0: never, which a run from
before 2.13.0 gives); stage 3b fits it, after bidiag’s threshold, on the
points where band_vals was timed (k >= 512, where bidiag_vals is). Then
values_band_width, chosen by stage 3b as stage 4b does for the
eigensolver. And band_min_k (since 2.15.0): with singular vectors, from
this k the band backend instead of bidiag or the CPU, within
bidiag_max_batch (0: never, which a run from before 2.15.0 gives); stage 3c
fits it as stage 3b does, on the points where band was timed with vectors
(k >= 512).
What every section and number in a report means (regret, the flat region,
the curves, the held-out checks, the decision surfaces, the noise floor) is
explained in reading-reports.md. In short, whether a
run is usable:
| warning | meaning | what to do |
|---|---|---|
| the machine was not idle | load average above half the CPU count, or Low Power Mode on, at the start or end | stop the other jobs and rerun |
| the machine changed state during the sweep | a probe point timed before and after moved by more than 25%: thermal throttling or a job that started midway | rerun, on mains, once the machine has cooled |
| single pass | --quick was used, so there is no noise floor |
rerun without --quick |
| (QR) median pass-to-pass ratio above 1.10 | the machine was not stable | rerun |
Any of these makes run.py mark that decomposition untrustworthy, and
combine.py leaves it out. The other warnings do not invalidate a run:
| warning | what to do |
|---|---|
| the GPU is still ahead of the CPU at the largest N (or k) measured | the cap reported is a lower bound. Harmless at 1024: a lone matrix stays on the CPU through the minimum batch, and batches above 1024 take seconds either way |
this device has no entry in kTuned[] |
expected on a new device |
| the fitted policy differs from the one in effect | update the device’s row |
| chosen rule loses more than 25% at … | informational: where a simple rule is furthest from the best backend |
| chosen rule picks a backend that was not timed at … | see “est. picks” in reading-reports.md |
| the best feature is no longer M (QR) | a structural change, not a moved threshold; worth an issue rather than a new row |
The harnesses run on their own too, which is useful when changing one of them:
cmake --build build --target sweep_qr sweep_eigh sweep_svd
python3 tuning/tune_qr.py build/sweep_qr # writes qr-tune-results/
python3 tuning/tune_eigh.py build/sweep_eigh --max-n 4096 # writes eigh-tune-results/
python3 tuning/tune_svd.py build/sweep_svd --max-k 4096 # writes svd-tune-results/
| option | harness | use |
|---|---|---|
--out DIR |
all | where to write |
--passes N |
all | more than two passes, for a noisy machine |
--full |
QR | a denser grid, about three times longer |
--max-n, --max-k |
eigh, SVD | the largest size on the grid (default 512; larger sizes up to 4096 added up to this) |
--quick |
eigh, SVD | one pass on a coarse grid; a smoke test only |
--limit S |
all | per-point timeout in seconds |
To try a policy before rebuilding, copy the environment line from a report,
e.g. EIGH_GPU_MAX_N=128 ./build/benchmark_eigh or
QR_M_CROSSOVER=320 ./build/benchmark_qr.
Re-analysing without the GPU, for example after a harness change:
python3 tuning/tune_eigh.py --reanalyse docs/results/apple-m5-pro-20gpu/20260930-27b6c2/eigh/raw.csv --out /tmp/eigh
Several raw.csv files may be given: they combine as in section 3 (files in
one submission merge by min-of-passes, so a sweep can be topped up). Passing
the sweep binary as well makes the report describe the policy that binary
resolves; without it, the policy in effect when the sweep ran. --from
results.json re-renders a report from its JSON.
The routing decides which backend runs. How each eigensolver and SVD backend is launched (threads per matrix, inner sweeps, simdgroups per matrix) scales with the detected core count and is not part of the per-device table, but it can be inspected:
./build/benchmark_eigh --tune # five tables; --tune 1..5 for one
./build/benchmark_svd --tune # simdgroups per matrix
./build/probe_occupancy # threadgroup-memory limits for QR
The M1 tables are in
studies/eigh-launch-parameters-apple-m1.md
and studies/svd-design-notes.md.
| symptom | cause and fix |
|---|---|
Impacting Interactivity in an error message |
macOS stopped a GPU command buffer that ran for several seconds. The library splits large batches to avoid this; during a sweep the point is retried and then recorded as failed. Lower EIGH_CHUNK_MS or SVD_CHUNK_MS (default 750) if it recurs |
GPU Hang Error in an error message |
with the display busy, macOS stopped a threadgroup that ran for more than about a quarter of a second. Since 2.15.0 the whole-matrix Jacobi kernels split a long solve over dispatches (EIGH_DISPATCH_MS, SVD_DISPATCH_MS, default 40); if another kernel shows it, measure with the Mac idle |
rows with ok = 0 in raw.csv |
the backend failed its correctness gate or timed out at that point, and the point is excluded. A few are harmless; many at small sizes indicate a real fault |
sweep_eigh --policy failed (or another sweep) |
the binary is older than the harness; rebuild it |
missing Metal Toolchain |
only matters when changing a shader; the build otherwise uses shaders/prebuilt/. xcodebuild -downloadComponent MetalToolchain installs it |
| very different answers from two runs | the machine was not in the same state for both; compare the noise floor and machine-state lines of the two reports |
Studies, the reasoning behind each harness and the results in full:
studies/qr-routing-apple-m1.mdstudies/eigh-routing-apple-m1.mdstudies/routing-apple-m5-pro.mdstudies/svd-design-notes.mdEnvironment overrides. All take effect without a rebuild and are reported in the policy source.
| variable | effect |
|---|---|
QR_M_CROSSOVER |
QR: rows at which the grid-parallel backend takes over |
QR_GPU_MAX_K, QR_GPU_MIN_K, QR_GPU_MIN_BATCH_TIMES_K, QR_GPU_MIN_BATCH |
QR: the GPU/CPU boundary |
QR_SHARE_MIN_BATCH |
QR: a GPU batch shared with the CPU from this batch (0: never) |
QR_GPU_LARGE_MIN_K, QR_GPU_LARGE_MAX_BATCH |
QR: the GPU also from this k, for batches up to this (0: never / any batch) |
QR_DEVICE=gpu or cpu |
QR: bypass the GPU/CPU boundary |
EIGH_SIMD_MAX_N, EIGH_BLOCK_MIN_N |
eigensolver: the GPU backend split |
EIGH_BLOCK_MIN_N_BATCHED, EIGH_BLOCK_MIN_BATCH |
eigensolver: batch-dependent block crossover, 0 for off |
EIGH_GPU_MAX_N, EIGH_GPU_MIN_BATCH_TIMES_N, EIGH_GPU_MIN_BATCH |
eigensolver: the GPU/CPU boundary |
EIGH_VALUES_GPU_MAX_N, EIGH_VALUES_GPU_MIN_BATCH_TIMES_N, EIGH_VALUES_GPU_MIN_BATCH |
eigensolver, eigenvalues alone: the GPU/CPU boundary |
EIGH_TRIDIAG_MIN_N, EIGH_VALUES_TRIDIAG_MIN_N |
eigensolver: the tridiag backend instead of the CPU from this N (0: never) |
EIGH_TRIDIAG_MAX_BATCH, EIGH_VALUES_TRIDIAG_MAX_BATCH |
eigensolver: the tridiag backend only for batches up to this (0: any) |
EIGH_QL_MIN_N, EIGH_QL_MAX_N |
eigensolver: the ql backend on the GPU for N in this window (EIGH_QL_MAX_N=0: never) |
EIGH_SHARE_MIN_BATCH |
eigensolver: a ql batch shared with the CPU from this batch (0: never) |
EIGH_GPU_BIG_BATCH_MAX_N, EIGH_GPU_BIG_BATCH_MIN |
eigensolver: the GPU also for N above gpu_max_n up to this, in batches of at least this (0: never) |
EIGH_VALUES_BAND_MIN_N |
eigenvalues alone: the band backend (the two-stage reduction) from this N (0: never) |
EIGH_VALUES_BAND_WIDTH |
eigenvalues alone: the band backend’s band width, 8, 16 or 32 (0: 16) |
EIGH_BAND_WIDTH |
the eigensolver’s band backend’s band width where the policy’s is 0: 8, 16 (default) or 32 |
METAL_LINALG_CPU_THREADS |
every decomposition: CPU threads a batch is spread over (default: every core) |
EIGH_DEVICE=tridiag |
eigensolver: every call on the tridiag backend |
EIGH_DEVICE=band |
eigensolver: eigenvalues alone on the band backend (with eigenvectors, tridiag) |
EIGH_DEVICE=gpu or cpu |
eigensolver: bypass the GPU/CPU boundary |
SVD_QR_MIN_ROWS, SVD_QR_MIN_K |
SVD: when the QR-preconditioned backends are used |
SVD_BLOCK_MIN_K |
SVD: the short side from which the block kernel is used |
SVD_BLOCK_MIN_K_BATCHED, SVD_BLOCK_MIN_BATCH |
SVD: batch-dependent block crossover, 0 for off |
SVD_GPU_MAX_K, SVD_GPU_MIN_BATCH_TIMES_K, SVD_GPU_MIN_BATCH |
SVD: the GPU/CPU boundary |
SVD_BIDIAG_MIN_K, SVD_VALUES_BIDIAG_MIN_K |
SVD: the bidiag backend instead of the CPU from this k (0: never) |
SVD_BIDIAG_MAX_BATCH, SVD_VALUES_BIDIAG_MAX_BATCH |
SVD: the bidiag backend only for batches up to this (0: any) |
SVD_GK_MIN_K, SVD_GK_MAX_K |
SVD: the golub_kahan backend on the GPU for k in this window (SVD_GK_MAX_K=0: never) |
SVD_SHARE_MIN_BATCH |
SVD: a golub_kahan batch shared with the CPU from this batch (0: never) |
SVD_GPU_BIG_BATCH_MAX_K, SVD_GPU_BIG_BATCH_MIN |
SVD: the GPU also for k above gpu_max_k up to this, in batches of at least this (0: never) |
SVD_VALUES_BAND_MIN_K |
singular values alone: the band backend (the two-stage reduction) from this k (0: never) |
SVD_VALUES_BAND_WIDTH |
singular values alone: the band backend’s band width, 8, 16 or 32 (0: 16) |
SVD_BAND_WIDTH |
the band backend’s band width where the policy’s is 0: 8, 16 (default) or 32 |
SVD_VALUES_GPU_MAX_K, SVD_VALUES_GPU_MIN_BATCH_TIMES_K, SVD_VALUES_GPU_MIN_BATCH, SVD_VALUES_GPU_MAX_L |
SVD, singular values alone: the GPU/CPU boundary (SVD_VALUES_GPU_MIN_BATCH=0: as with vectors) |
SVD_DEVICE=bidiag |
SVD: every call on the bidiag backend |
SVD_DEVICE=band |
SVD: singular values alone on the band backend (with vectors, bidiag) |
SVD_DEVICE=gpu or cpu |
SVD: bypass the GPU/CPU boundary |
Programmatic overrides. set_qr_policy(), set_eigh_policy() and
set_svd_policy() take precedence over both the environment and the table;
qr_policy_source(), eigh_policy_source() and svd_policy_source() report
which is in effect.