QR decomposition, symmetric eigendecomposition (eigh) and singular value
decomposition (SVD) for batches of matrices on Apple Silicon GPUs, for
MLX, PyTorch and
plain float buffers. A C++ library, installed with Homebrew or built from
source inside your own project, with Python packages for mlx.core arrays
and for torch tensors, a C API, and a Swift package (on [Float], and on
mlx-swift’s MLXArray).
Each solver has several Metal kernels, one per regime (small matrices in large
batches, large matrices spread over the whole GPU, long thin matrices), and
every call is routed to the fastest of them, or to LAPACK on the CPU (a batch
spread over every core), by a policy measured on the device it runs on. MLX’s own linalg::eigh and
linalg::svd (mx.linalg.eigh and mx.linalg.svd in Python) run only on the CPU.
Contributions welcome: measure your Mac. The routing is only as good as the measurements behind it, and every new chip needs its own. See which Macs are measured: every Apple Silicon chip, colour-coded per decomposition (current, out of date, or not measured yet). So far an M5 Pro has current measurements (an M1’s predate 2.9.0 and are no longer used); every other Mac runs settings estimated from them and published benchmarks, which lean toward the CPU and miss some of what its GPU can do. If you have an Apple Silicon Mac, one command measures it (
python3 tuning/run.py, about an hour and a half of the Mac’s time) and produces a results folder to send as a pull request. Each run improves the library for everyone with that Mac, and runs from several people with the same Mac are combined. Contributions are what keep the library up to date as Apple ships new chips: how to contribute.
Formerly
qr-apple-silicon. The project was renamed in version 2.0, when it grew from QR to QR, symmetric eigendecomposition and SVD. Links togithub.com/c0rmac/qr-apple-siliconredirect here; to update an existing clone, rungit remote set-url origin https://github.com/c0rmac/metal-linalg.git. The changes from 1.x are listed in CHANGELOG.md.
| operation | functions | GPU kernels | CPU path | details |
|---|---|---|---|---|
| QR | qr_accelerated |
Householder in one simdgroup’s registers, or blocked in one threadgroup, per matrix; grid-parallel blocked Householder | LAPACK sgeqrf, sorgqr |
docs/qr.md |
| symmetric eigendecomposition | eigh_accelerated, eigvalsh_accelerated |
whole-matrix Jacobi; block Jacobi; tridiagonalization and implicit QL in one threadgroup per matrix (N <= 87); Householder tridiagonalization for large N (with LAPACK’s tridiagonal solver, or bisection on the GPU for eigenvalues alone); for eigenvalues alone of large N, a two-stage reduction (to a band on the GPU, then to tridiagonal on every CPU core) | LAPACK ssyevd; ssyevd_2stage for eigenvalues alone from N = 128 |
docs/eigh.md |
| thin SVD | svd_accelerated, svdvals_accelerated |
whole-matrix one-sided Jacobi; block one-sided Jacobi; either after QR for tall input; bidiagonalization and implicit QR in one threadgroup per matrix (k <= 83); Householder bidiagonalization for large k (with LAPACK’s bidiagonal solver, or bisection on the GPU for singular values alone); for singular values alone of large k, a two-stage reduction (to a band on the GPU, then to bidiagonal on every CPU core) | LAPACK sgesdd |
docs/svd.md |
On the CPU a batch is spread over every core, each solving whole matrices
(set_cpu_threads() or METAL_LINALG_CPU_THREADS caps it).
Input is any batch shape [..., M, N], any real dtype (computed in float32),
any magnitude from 1e-30 to 1e+37, rank-deficient or not. The eigensolver and
the SVD return NaN for a non-finite matrix rather than raising, leaving the
rest of its batch intact. The QR and SVD factors are the thin ones,
K = min(M, N).
The same solvers and routing, from five places, each on an Apple Silicon Mac:
| platform | works on | install | |
|---|---|---|---|
| C++ | MLX arrays (mlx::core::array) |
Homebrew, or CMake from source | C++ |
| Python, MLX | MLX arrays (mlx.core.array) |
pip install metal-linalg |
With MLX |
| Python, PyTorch | torch tensors, on the CPU or MPS |
pip install metal-linalg-torch |
With PyTorch |
| Swift | [Float], or mlx-swift’s MLXArray |
Swift Package Manager | Swift |
| C and Objective-C | plain float buffers, no MLX | as for C++ | C and Objective-C |
On an Apple M5 Pro (20 GPU cores), against a CPU path that spreads every call over all 18 CPU cores, the GPU wins in two places, and the router sends work there and nowhere else.
One large matrix: up to 10x, and 11.8x at 8192. eigh and the SVD keep LAPACK’s method and move its memory-bound reduction (to tridiagonal or bidiagonal form) and its back-transformation to the GPU; the divide and conquer between them runs on every CPU core (since 2.15.0). For the eigenvalues or singular values alone, large matrices are reduced in two stages (since 2.13.0): to a band on the GPU, in blocks whose work is matrix products, then to tridiagonal or bidiagonal on every CPU core, and the values come from bisection on the GPU. Since 2.15.0 the SVD with vectors can take the two stages too, both stages’ transformations applied on the GPU while the CPU chases the band and solves the bidiagonal problem. QR runs on the same panels and matrix products (since 2.15.0, a batch at once). One N×N matrix against the CPU:
| 1024 | 1536 | 2048 | 3072 | 4096 | |
|---|---|---|---|---|---|
| svdvals, singular values alone | 1.86x | 2.57x | 3.77x | 7.10x | 9.95x |
| eigh, with eigenvectors | 2.29x | 3.16x | 4.49x | 5.91x | 9.24x |
| eigvalsh, eigenvalues alone | 1.62x | 2.01x | 2.31x | 3.10x | 3.69x |
| SVD, with vectors | 2.77x | 3.65x | 5.57x | 8.10x | 10.4x |
| QR | 2.54x | 3.34x | 5.24x | 6.87x | 10.3x |
At 8192, the singular values alone take 1.07 s against the CPU’s 12.6 s
(11.8x), eigh 2.05 s against 18.2 s (8.9x), and the eigenvalues alone 0.70 s
against 2.64 s (3.8x, against LAPACK’s own two-stage driver). The M5 Pro uses
these backends from N = 1024, for one matrix or a few: a batch of them is
pipelined, the CPU solving one matrix’s small problem while the GPU reduces
the next, but the CPU path spreads a large batch over its cores. The SVD
with vectors’ row is the two-stage band backend, which the M5 Pro routes to
from k = 1024 (bidiag, the one-stage reduction, 4.2x at 4096). QR’s GPU
path takes batches too: 16 of 1024×1024 2.1x, 4 of 2048×2048 4.1x, and tall
matrices, one 8192×512 5.6x. At 4096, 2.14 had svdvals at 8.62x, eigh 5.62x,
eigvalsh 2.91x, the SVD with vectors 2.32x and QR 2.7x.
Large batches of small matrices: up to 5.6x. LAPACK’s own methods, a
matrix to a threadgroup or a simdgroup, carry the GPU’s lead: the
eigensolver’s ql kernel (tridiagonalization and QL), the SVD’s
golub_kahan (bidiagonalization and implicit QR, 1.6-3x faster than the
Jacobi kernels it replaced), and since 2.16.0 QR’s Householder kernels (up to
32 columns in a simdgroup’s registers; above that blocked in a threadgroup,
the updates as 8×8 simdgroup matrix products), which take QR’s batches from
the CPU from 64×64 at 256 matrices and lead by 4.7-5.6x at 128×128. eigh’s
and the SVD’s large batches are shared, the GPU and the CPU solving them at
once (from 4096 matrices for eigh, 256 for the SVD), up to 64×64 for eigh and
80×80 for the SVD. The best GPU route against the CPU alone:
| lone matrix | batch 16 | batch 256 | batch 4096 | |
|---|---|---|---|---|
| eigh 24×24 | 0.08x | 0.46x | 1.20x | 2.37x |
| eigh 32×32 | 0.11x | 0.30x | 1.20x | 2.36x |
| eigh 64×64 | 0.17x | 0.20x | 0.99x | 1.69x |
| eigvalsh 32×32 | 0.37x | 0.33x | 0.71x | 1.64x |
| SVD 16×16 | 0.06x | 0.27x | 1.40x | 2.38x |
| SVD 32×32 | 0.11x | 0.30x | 1.33x | 2.44x |
| SVD 48×48 | 0.16x | 0.30x | 1.66x | 2.11x |
| SVD 64×64 | 0.20x | 0.31x | 1.13x | 1.71x |
| QR 16×16 | 0.01x | 0.15x | 0.53x | 1.42x |
| QR 32×32 | 0.02x | 0.34x | 0.55x | 3.13x |
| QR 64×64 | 0.03x | 0.43x | 1.35x | 2.34x |
| QR 128×128 | 0.29x | 1.22x | 4.65x | 5.57x |
| QR 256×256 | 0.85x | 1.45x | 2.87x | 2.79x |
Lone small matrices and small batches stay on the CPU, which is 3-100x faster
there: since 2.9.0 it spreads a batch over every core (the
performance-headroom study
has why); so do batches of mid-size matrices (about 96 to 512) for eigh and
the SVD. The numbers for one large matrix are 2.16.0’s, measured side by
side with the CPU path (sweep_eigh, sweep_svd, sweep_qr, median, MLX’s
buffer cache on as an MLX program has it; at 8192 the CPU’s from 2.15.0, its
path unchanged since); the batches’ too, on 2.16.0 (each GPU backend against
the CPU path, the best of two passes, sharing from batch 64 as the routing
sweeps time it), and QR’s tall matrices from
20261007-9f2589. The full tables are in the per-solver docs; how the two-stage reduction
got there is in the two-stage study,
and 2.15.0’s changes in its study.
Which kernel is fastest, and where the GPU overtakes the CPU, depends on the chip, so each solver’s thresholds are a table of measured policies keyed on the Metal device name and GPU core count:
| device | QR | eigh | SVD | |—|—|—|—| | Apple M1, 8 GPU cores | estimated (out of date) | estimated (out of date) | estimated | | Apple M5 Pro, 20 GPU cores | measured | measured | measured | | anything else | estimated | estimated | estimated |
Every chip, and what is current: the measurements page.
qr_policy_source(), eigh_policy_source() and svd_policy_source() say
which applies (tuned:Apple M5 Pro, estimated:<name> (from Apple M5 Pro, ...),
env:... or user). A Mac nobody has measured gets an estimate: a measured
Mac’s timings refitted for a GPU as much weaker against its CPU as Geekbench
and the memory bandwidth say, with a margin, so that it misses some GPU wins
rather than sending work to a backend that loses
(how, and how close it comes).
METAL_LINALG_ESTIMATE_AS="Apple M4 Max:40" shows what another Mac gets.
Every threshold can also be overridden with an environment variable or
set_*_policy(); see seeing and changing the routing.
The formula lives in its own tap, c0rmac/homebrew-metal-linalg. Tap it, trust it, then install:
brew tap c0rmac/metal-linalg
brew trust c0rmac/metal-linalg # Homebrew 7 and later ask this of any third-party tap
brew install metal-linalg
This builds and installs libmetal_linalg.dylib, the headers under
include/metal_linalg/ and a CMake package. The compiled Metal shaders are
embedded in the library, so nothing is looked up on disk at run time and
nothing needs the Metal compiler. Updating, uninstalling and depending on it
from another formula are covered in
the tap’s README.
You need CMake 3.25 or later, a C++20 compiler (Xcode’s clang) and MLX
(brew install mlx). The Metal shader compiler is needed only to change a
shader: without it the build uses the compiled shaders committed under
shaders/prebuilt/. From Xcode 26 it is a separate download (xcodebuild
-downloadComponent MetalToolchain); after editing a shader with it,
cmake --build build --target update_prebuilt_shaders refreshes
shaders/prebuilt/.
git clone https://github.com/c0rmac/metal-linalg.git
cd metal-linalg
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DCMAKE_PREFIX_PATH=/opt/homebrew
cmake --build build -j
cmake --install build # or --prefix <dir>
Build options:
| option | default | effect |
|---|---|---|
METAL_LINALG_WITH_MLX |
ON |
the MLX API; OFF builds the buffer API and the C API alone, with no MLX anywhere |
METAL_LINALG_SHARED |
ON when built on its own, OFF as a subproject |
a shared or a static library |
METAL_LINALG_BUILD_TESTS |
ON when built on its own, OFF as a subproject |
the tests, benchmarks and tuning harnesses |
METAL_LINALG_BUILD_EXAMPLES |
that of METAL_LINALG_BUILD_TESTS |
the examples, run as tests |
METAL_LINALG_USE_PREBUILT_SHADERS |
OFF, and ON without a Metal compiler |
the metallibs in shaders/prebuilt/ instead of compiling the shaders |
Against the installed package:
find_package(MetalLinalg REQUIRED)
target_link_libraries(my_app PRIVATE metal_linalg::metal_linalg)
Or built as part of your own project, so that the library uses the shaders and routing tables of the checkout you point it at. As a subproject it is static, folded into your binary, and installs nothing of its own:
add_subdirectory(path/to/metal-linalg) # or:
include(FetchContent)
FetchContent_Declare(metal_linalg
GIT_REPOSITORY https://github.com/c0rmac/metal-linalg.git GIT_TAG v2.0.0) # or main, for the latest
FetchContent_MakeAvailable(metal_linalg)
target_link_libraries(my_app PRIVATE metal_linalg::metal_linalg)
A complete program: one batch, all three decompositions. CMakeLists.txt:
cmake_minimum_required(VERSION 3.25)
project(my_app LANGUAGES CXX)
find_package(MetalLinalg REQUIRED)
add_executable(my_app main.cpp)
target_link_libraries(my_app PRIVATE metal_linalg::metal_linalg)
main.cpp:
#include <metal_linalg/metal_linalg.h>
#include <mlx/mlx.h>
#include <cstdio>
namespace mx = mlx::core;
namespace ml = metal_linalg;
int main() {
mx::set_default_device(mx::Device::gpu);
// A batch of 1000 random 64 x 32 matrices.
mx::array A = mx::random::normal({1000, 64, 32});
// QR: Q [1000, 64, 32] with orthonormal columns, R [1000, 32, 32] upper triangular.
auto [Q, R] = ml::qr_accelerated(A);
// Thin SVD: U [1000, 64, 32], S [1000, 32] descending, Vt [1000, 32, 32].
auto [U, S, Vt] = ml::svd_accelerated(A);
// Symmetric eigendecomposition of C = A^T A: w [1000, 32] ascending, V [1000, 32, 32].
mx::array C = mx::matmul(mx::swapaxes(A, -1, -2), A);
auto [w, V] = ml::eigh_accelerated(C);
mx::eval({Q, R, U, S, Vt, w, V});
mx::array qr_err = mx::max(mx::abs(mx::subtract(mx::matmul(Q, R), A)));
mx::array svd_err = mx::max(mx::abs(mx::subtract(
mx::matmul(mx::multiply(U, mx::expand_dims(S, -2)), Vt), A)));
std::printf("max |QR - A| = %.1e, max |U diag(S) Vt - A| = %.1e\n",
qr_err.item<float>(), svd_err.item<float>());
}
cmake -S . -B build -DCMAKE_PREFIX_PATH=/opt/homebrew
cmake --build build
./build/my_app # e.g. max |QR - A| = 5.2e-06, max |U diag(S) Vt - A| = 6.7e-06
Every function takes any batch shape [..., M, N] and returns MLX arrays.
Unlike most MLX operations the decompositions are not lazy: each call
evaluates its input and runs its kernels before it returns. A few finishing
steps, such as undoing the input scaling, come back as ordinary lazy arrays,
so eval the results as usual. The functions are written for, and tested
with, the GPU as the default MLX device.
There are two Python packages, with the same solvers and routing; install the one for the arrays you use:
| you use | install | import |
|---|---|---|
MLX (mlx.core.array) |
pip install metal-linalg |
import metal_linalg as ml |
PyTorch (torch.Tensor) |
pip install metal-linalg-torch |
import metal_linalg_torch as mlt |
They are independent: the PyTorch package neither installs nor loads MLX, and the MLX package does not need torch. Both can be installed side by side.
pip install metal-linalg
Prebuilt wheels for Apple Silicon Macs on macOS 14 or later, Python 3.10 to
3.14. Each release is built against one MLX release and pins it (mlx==0.32.3
at present), since it shares MLX’s library and arrays; pip installs that
mlx with it. python/README.md covers building from
source, and against Homebrew’s MLX.
import mlx.core as mx
import metal_linalg as ml
a = mx.random.normal((1000, 64, 32))
Q, R = ml.qr(a)
U, S, Vt = ml.svd(a)
w, V = ml.eigh(a.swapaxes(-1, -2) @ a)
ml.eigh_backend(512, 1) # 'cpu': which backend a shape gets on this Mac
Inputs may be mx.array, NumPy arrays or nested lists; outputs are float32
mx.array. ml.eigvalsh and ml.svdvals return the values alone.
pip install metal-linalg-torch
One prebuilt wheel for Apple Silicon Macs on macOS 14 or later, for every Python from 3.10 and every PyTorch from 2.4: it calls the library’s C API and is compiled against neither, so it pins nothing and upgrading torch never breaks it.
import torch
import metal_linalg_torch as mlt
a = torch.randn(1000, 64, 32, device="mps") # or on the CPU
Q, R = mlt.qr(a) # like torch.linalg.qr
U, S, Vh = mlt.svd(a) # thin factors, like torch.linalg.svd(a, full_matrices=False)
L, V = mlt.eigh(a.mT @ a) # like torch.linalg.eigh
mlt.svd_backend(4096, 4096) # 'bidiag': which backend a shape gets on this Mac
The functions mirror their torch.linalg namesakes (arguments, result types,
the input’s device), support autograd with torch’s own formulas, and compile
with torch.compile: they are the custom operators
torch.ops.metal_linalg.*. Computation is in float32. Against torch.linalg
on an M5 Pro with PyTorch 2.13 (conda-forge’s, its CPU LAPACK from Accelerate) and
metal-linalg 2.16 (best of five, benchmarks/benchmark_torch.py;
the same tensors on MPS for torch’s MPS path and for this package, which uses them in place):
| torch, CPU | torch, MPS | metal-linalg-torch | |
|---|---|---|---|
| QR, 1024 × 128×128 | 202 ms | 32 ms | 4.4 ms |
| SVD, 256 × 128×64 | 69 ms | 71 ms | 5.1 ms |
| SVD, 4096 × 32×32 | 223 ms | 232 ms | 6.3 ms |
| eigh, 4096 × 16×16 | 33 ms | 35 ms | 2.0 ms |
| eigh, one 2048×2048 | 263 ms | 272 ms | 56 ms |
| SVD, one 4096×4096 | 3.60 s | 3.64 s | 339 ms |
| eigvalsh, one 4096×4096 | 1.78 s | 1.80 s | 125 ms |
| svdvals, one 4096×4096 | 1.95 s | 1.99 s | 193 ms |
It is ahead on every row. Of these calls PyTorch 2.13 runs only QR on the GPU
for MPS tensors, and this is 7x faster there (the QR batch runs on the
library’s Householder kernels); its SVD takes as long on MPS
as on the CPU, and eigh, eigvalsh and svdvals have no MPS kernels and go
through its CPU fallback. Against those, 13-35x for the other batches of
small matrices (the SVD of 256 matrices of 128×64 is a QR and then
golub_kahan on the GPU; the two batches of 4096 run on the GPU and the CPU
at once), 4.7x for eigh of one 2048×2048
and 10x for the SVD of one 4096×4096 with its vectors, and 10-14x for its
eigenvalues or singular values alone (the last three by a two-stage
reduction).
python-torch/README.md has the details: what differs
from torch.linalg, gradients, and MPS tensors.
The repository is a Swift package with two libraries: MetalLinalg, on
[Float], which needs nothing else, and MetalLinalgMLX, on mlx-swift’s
MLXArray. macOS 14 or later, on Apple Silicon. In Package.swift:
dependencies: [
.package(url: "https://github.com/c0rmac/metal-linalg.git", from: "2.0.0"),
],
targets: [
.target(name: "MyApp", dependencies: [
.product(name: "MetalLinalg", package: "metal-linalg"), // or "MetalLinalgMLX"
]),
]
or in Xcode, File > Add Package Dependencies, with the repository’s URL.
import MetalLinalg
// Two symmetric 64 x 64 matrices, row-major, one after the other.
let (w, v) = try eighAccelerated(a, batch: 2, n: 64) // w ascending, v's columns the vectors
import MLX
import MetalLinalgMLX
let a = MLXRandom.normal([1000, 64, 32])
let (q, r) = try qrAccelerated(a)
let (u, s, vt) = try svdAccelerated(a)
Errors are thrown as MetalLinalgError. docs/swift.md has
the whole API and how the package is built.
The C API is part of the library: install it as for C++, with Homebrew
or from source. To leave MLX out of the library altogether, build from
source with -DMETAL_LINALG_WITH_MLX=OFF.
Under the MLX API is a core that works on plain float buffers (row-major,
the matrices of a batch one after another) and has no MLX in it:
<metal_linalg/core.h> in C++, and the same through a C API,
<metal_linalg/c_api.h>:
#include <metal_linalg/c_api.h>
#include <stdio.h>
int main(void) {
/* Two symmetric 2 x 2 matrices, row-major, one after the other. */
const float a[8] = {2, 1, 1, 2, 4, 0, 0, 1};
float w[4], v[8];
if (metal_linalg_eigh(a, 2, 2, /*lower*/ 1, w, v, NULL) != METAL_LINALG_OK) {
fprintf(stderr, "eigh failed: %s\n", metal_linalg_last_error());
return 1;
}
printf("%g %g | %g %g\n", w[0], w[1], w[2], w[3]); /* 1 3 | 1 4 */
return 0;
}
cc -std=c99 main.c -I/opt/homebrew/include -L/opt/homebrew/lib -lmetal_linalg -o main
The conventions (layout, optional outputs, errors) are in docs/c-api.md.
Objective-C++ (.mm) calls the C++ API directly, on MLX arrays; any
Objective-C file can call the C API on its own buffers, with no MLX. See
docs/objective-c.md.
In C++. Each is a complete program in examples/, built with the
tests and run by ctest, so it stays correct.
orthonormal_bases.cpp: 10,000 sets of 4
vectors in R^16, orthonormalised in one call.
mx::array vectors = mx::random::normal({10000, 16, 4});
auto [Q, R] = metal_linalg::qr_accelerated(vectors); // Q [10000, 16, 4], Q^T Q = I
pca.cpp: the dominant direction of each of 1000 point
clouds, from the eigenvector of its covariance with the largest eigenvalue.
Eigenvalues come back ascending, so that is the last column.
// X: [1000 clouds, 500 points, 8 dims]
mx::array cov = mx::divide(mx::matmul(mx::swapaxes(X, -1, -2), X), mx::array(499.0f));
auto [variance, components] = metal_linalg::eigh_accelerated(cov);
mx::array principal = mx::take(components, 7, -1); // [1000, 8], one axis per cloud
eigvalsh_accelerated(cov) returns the eigenvalues alone, for about a third
less work.
nearest_orthogonal.cpp: projecting
100,000 noisy 3×3 matrices back onto the orthogonal group (the orthogonal
Procrustes problem). If M = U S V^T, the nearest orthogonal matrix is U V^T.
auto [U, S, Vt] = metal_linalg::svd_accelerated(observed); // observed: [100000, 3, 3]
mx::array nearest = mx::matmul(U, Vt);
svdvals_accelerated(a) returns the singular values alone, for about half
the work.
routing.cpp: which kernel a shape will get on this
machine, and why.
std::printf("%s, %u GPU cores, eigh policy from %s\n", metal_linalg::device_name(),
metal_linalg::gpu_core_count(), metal_linalg::eigh_policy_source());
// Apple M5 Pro, 20 GPU cores, eigh policy from tuned:Apple M5 Pro
metal_linalg::eigh_backend(32, 4096); // EighBackend::ql: a large batch of small matrices goes to the GPU
metal_linalg::eigh_backend(512, 64); // EighBackend::cpu: the CPU's cores win a batch of mid-size ones
metal_linalg::eigh_backend(2048, 1); // EighBackend::tridiag: one large matrix, mostly on the GPU
auto p = metal_linalg::eigh_policy(); // replace the measured policy at run time
p.gpu_min_batch = 1;
metal_linalg::set_eigh_policy(p); // eigh_policy_source() is now "user"
The same can be done without recompiling through environment variables, e.g.
EIGH_GPU_MIN_BATCH=1 or SVD_DEVICE=gpu; docs/tuning.md
lists them.
| header | contents |
|---|---|
<metal_linalg/metal_linalg.h> |
all of the below |
<metal_linalg/qr.h> |
qr_accelerated; QrPolicy, qr_policy(), set_qr_policy(), qr_policy_source(); qr_backend(m, n, batch) |
<metal_linalg/eigh.h> |
eigh_accelerated, eigvalsh_accelerated; EighPolicy, eigh_policy(), set_eigh_policy(), eigh_policy_source(); eigh_backend(n, batch), eigvalsh_backend(n, batch), eigh_uses_gpu, eigvalsh_uses_gpu |
<metal_linalg/svd.h> |
svd_accelerated, svdvals_accelerated; SvdPolicy, svd_policy(), set_svd_policy(), svd_policy_source(); svd_backend(m, n, batch), svdvals_backend(m, n, batch), svd_uses_gpu, svdvals_uses_gpu |
<metal_linalg/device.h> |
device_name(), gpu_core_count(): the GPU the policies were resolved for; cpu_threads(), set_cpu_threads(): how many cores the CPU paths spread a batch over |
<metal_linalg/core.h> |
the same on float buffers, without MLX: core::qr, core::eigh, core::svd; the policies, backends and options |
<metal_linalg/c_api.h> |
the C API: metal_linalg_qr, _eigh, _svd, the routing queries and policies |
Each header’s metal_linalg::detail namespace has the individual backends,
which always run their kernel, with options (tolerances, sweep bounds, launch
parameters) and a per-matrix info word; these are what the tests and tuning
harnesses use.
ctest --test-dir build --output-on-failure # test_qr, test_eigh, test_svd, test_core, test_c_api, the examples
./build/benchmark_qr # GPU against the CPU, per solver
./build/benchmark_eigh
./build/benchmark_svd
./build/sweep_svd --policy # the device and the policy in effect
The tests (116 QR, 319 eigh and 384 SVD checks through MLX, 208 on the buffer API and 66 on the C API) cover every backend directly and through the router, shapes around every kernel boundary, batches, transposed views, structured and rank-deficient input, magnitudes from 1e-30 to 1e+37, NaN inside a batch (eigh, SVD), and the routing policies without assuming any device’s values. The Python (MLX and PyTorch) and Swift packages have their own tests; see their guides.
| path | contents |
|---|---|
include/metal_linalg/ |
the public headers |
src/ |
host code: routing policies, the Metal runtime, one driver per backend, the MLX and C layers |
shaders/ |
the Metal kernels; prebuilt/ holds their compiled metallibs |
python/ |
the Python package for MLX: bindings, the metal_linalg module, its tests |
python-torch/ |
the Python package for PyTorch: the metal_linalg_torch module over the C API, its tests |
Package.swift, swift/ |
the Swift package: the C module map, the embedded shaders, the Swift API and its tests |
examples/ |
small self-checking programs, one per use case |
tests/ |
correctness tests |
benchmarks/ |
GPU-against-CPU benchmarks |
tuning/ |
measuring a Mac (run.py), combining runs (combine.py), and the sweeps behind them |
docs/ |
per-solver guides, the tuning guide, studies, and every submitted run under results/<device>/<id>/ |
ql kernel;
the two-stage reduction on an M5 Pro,
behind the band backends and bisection on the GPUThe most useful contribution is measuring your Mac: one command,
python3 tuning/run.py, then a pull request with the results folder.
CONTRIBUTING.md walks through both, including how to make
the pull request, and covers bug reports and code changes too.
MIT; see LICENSE.