Skip to content

GEMM Optimization

Performance numbers are teaching placeholders

The TFLOPS, speedup and memory figures on this page are reference values, not results reproduced by this repository on a pinned hardware/software stack. Re-measure on your own GPU before quoting them.

Detailed walkthrough of the 7-step GEMM optimization path.

Optimization Path

Summary

StepTechniqueTFLOPS (FP32)Speedup
1Naive0.51.0×
2Shared Memory2.04.0×
3Double Buffer3.57.0×
4Register Tiling6.012.0×
5WMMA50+100×
6MMA PTX60+120×
7Pipeline70+140×

References

Released under the MIT License.