metal-linalg: Batched QR, Eigendecomposition and SVD on Apple Silicon GPUs
QR, symmetric eigendecomposition and SVD for batches of matrices on Apple Silicon GPUs, for MLX and PyTorch: 2x to 10x faster than the CPU. Each call is routed to the fastest Metal kernel, or to the CPU, by a policy measured per Mac. C++, Python, Swift and C.



