Version2.0.15
Revision396
Size0.2 MB
LicenseApache-2.0
Confinementclassic
Basecore20

A synthetic micro-benchmark that measures peak compute, bandwidth, and matrix throughput of GPUs and CPUs


clpeak — "Compute Latency PEAK". A synthetic micro-benchmark that measures the peak achievable performance of GPU compute devices. It exercises tight vector / MAD / MMA loops and vendor-SDK GEMM libraries (cuBLASLt on NVIDIA, MPS on Apple) to expose what the hardware is capable of — from raw ALU peaks to near-vendor-advertised matrix throughput.

clpeak began as an OpenCL-only tool and is now a multi-backend benchmark — OpenCL, Vulkan, CUDA, ROCm/HIP, Metal, oneAPI/SYCL, plus a native CPU backend — run back-to-back on the same hardware, so cross-stack differences (driver lowering, instruction scheduling, extension exposure) surface alongside the raw peak numbers.

Update History

2.0.14 (389)2.0.15 (396)
25 Jun 2026, 08:15 UTC
2.0.12 (379)2.0.14 (389)
23 Jun 2026, 14:15 UTC
1.1.2 (256)2.0.12 (379)
14 Jun 2026, 08:15 UTC
1.1.2 (256)
1 Apr 2026, 21:28 UTC

Published10 Jul 2019, 13:56 UTC

Last updated25 Jun 2026, 02:49 UTC

First seen1 Apr 2026, 21:28 UTC