Skip to content
← Work

Optimizing matrix multiplication

Successive optimization of matmul from naive loops through cache blocking to CUDA/cuBLAS, profiled with Nsight.

When
2024-11 – 2024-12
Source
MilburnJ/matrix_optimization ↗

Systems depth for ML-infrastructure work: what actually moves matmul performance once you look past the CPU - caches, prefetching, then the GPU.