Optimizing matrix multiplication
Successive optimization of matmul from naive loops through cache blocking to CUDA/cuBLAS, profiled with Nsight.
- When
- 2024-11 – 2024-12
- C++
- CUDA
- cuBLAS
- OpenBLAS
- Nsight Compute
Systems depth for ML-infrastructure work: what actually moves matmul performance once you look past the CPU - caches, prefetching, then the GPU.