Measured benchmark: RTX 3060 Laptop, 2026-08-17
This is the only benchmark page in the repository that reports repository-measured numbers. All other performance figures in teaching documents are placeholders unless they link here.
Environment
| Item | Value |
|---|---|
| GPU | NVIDIA GeForce RTX 3060 Laptop GPU, 6 GB, sm_86 |
| Driver / nvidia-smi | 591.44 |
| CUDA toolkit | 12.0.140 (nvcc) |
| Host | Ubuntu 24.04, GCC 13.3 |
| Build | Release, --use_fast_math, -lineinfo, -arch=sm_86 |
| WSL note | LD_PRELOAD=/usr/lib/wsl/lib/libcuda.so.1 was required on this WSL2 host |
01-sgemm-tutorial ladder
Command:
cd 01-sgemm-tutorial
make benchmark
./build/sgemm_benchmark -aWarmup: 5 runs. Timed runs: 20 per kernel, averaged with CUDA events.
| Kernel | 512^3 GFLOPS | 1024^3 GFLOPS | 2048^3 GFLOPS | 4096^3 GFLOPS |
|---|---|---|---|---|
| cuBLAS | 4242 | 5581 | 6390 | 5695 |
| Naive | 597 | 583 | 674 | 666 |
| Tiled 32x32 | 665 | 925 | 933 | 910 |
| Bank Conflict Free | 505 | 657 | 667 | 655 |
| Double Buffer | 496 | 678 | 683 | 667 |
| Tensor Core WMMA | 1626 | 1093 | 2768 | 4099 |
Key finding on this GPU: the simple tiled kernel is faster than both bank-conflict-free and double-buffer variants at 1024^3. Micro-optimizations must be profiled on the target hardware; they do not always improve.
WMMA correctness uses an FP16-input quantization-aware tolerance (atol = 0.0015 * sqrt(K)), because fixed small tolerances fail at K=2048+ even though the kernel is working as designed.
02-tensorcraft-core GEMM variants
Command:
cmake --preset default
cmake --build --preset default --target gemm_benchmark
./build/default/bin/gemm_benchmark --benchmark_min_time=0.1s --benchmark_repetitions=3Median wall time, 3 repetitions:
| Shape | Naive GFLOPS | Tiled GFLOPS | Double Buffer GFLOPS |
|---|---|---|---|
| 256^3 | 257 | 329 | 319 |
| 512^3 | 634 | 767 | 861 |
| 1024^3 | 641 | 913 | 952 |
| 2048^3 | 688 | 968 | 1057 |
For this teaching library, Double Buffer starts to help at larger sizes, but all custom kernels remain far below cuBLAS. The benchmark target exists to teach measurement, not to claim cuBLAS parity.
Known limitation
ncu profiling was attempted on this WSL2 host and failed with ERR_NVGPUCTRPERM (GPU performance counters not exposed). A reproducible runbook is provided in Performance Tuning; real traces must be captured on a machine where counters are enabled.