Release 13.4 of the CUDA Samples supported by CUDA Toolkit 13.4. See Changelog for more information.
clock - Per-Block Kernel Timing with CUB Reduction
Description
A CUDA sample that demonstrates how to use the clock() function to accurately measure kernel execution time on a per-block basis. Each of the 64 blocks records its own start and end SM cycle counter, then performs a block-wide parallel min-reduction over 512 float elements using CUB's BlockReduce. The host collects all timestamps and computes the average elapsed clock cycles across all blocks.
Because blocks execute in parallel and out of order with no cross-block synchronization, each block independently measures its own execution time — this is the correct way to time GPU work at block granularity.
What You'll Learn
- Using
clock()inside a CUDA kernel to capture per-block SM cycle counts - Performing a block-wide parallel reduction with
cub::BlockReduceusing a custom binary operator - Loading multiple elements per thread (
ITEMS_PER_THREAD = 2) for CUB reductions - Understanding why only thread 0 holds the valid aggregate after a
BlockReduce - Querying device properties (
cudaDeviceGetAttribute) without helper libraries - Computing average elapsed clocks on the host from per-block timestamps
Key Concepts
- SM Clock Counter —
clock()reads the streaming multiprocessor's cycle counter; difference between two samples gives elapsed cycles for that block - CUB BlockReduce — warp-shuffle-based block-scope reduction; default constructor allocates shared memory internally via
PrivateStorage(), no explicitTempStorageneeded - Items per thread — each thread owns 2 elements; CUB's array overload of
Reducecombines them before the cross-thread reduction - Per-block timing — since blocks run independently, each block times itself; the host averages results across all blocks
Key APIs
CUDA Runtime
cudaSetDevice— select the active GPU for all subsequent CUDA callscudaDeviceGetAttribute— query device properties (compute capability, SM count) withoutcudaGetDevicePropertiescudaMalloc/cudaFree— allocate and release device memorycudaMemcpy— transfer data between host and deviceclock()— device-side SM cycle counter (returnsclock_t)
CUB
cub::BlockReduce<T, BLOCK_THREADS>— block-scope reduction templateBlockReduce::Reduce(T (&input)[ITEMS_PER_THREAD], ReductionOp op)— reduce multiple items per thread with a custom binary operator
Requirements
Hardware
- NVIDIA GPU with Compute Capability 7.5 or higher
Software
- CMake 3.20 or newer
- A C++17-capable host compiler
How to Build
See the top-level README for full build instructions, including how to build all samples or a single sample standalone.
How to Run
./clock
No command-line arguments are required. The sample always runs on device 0.
Expected Output
CUDA Clock sample
GPU Device 0: with compute capability 8.9 and Number of SMs 142
Average clocks/block = 1239.640625
The average clocks value varies by GPU and reflects how many SM cycles each block takes to complete the reduction. Blocks scheduled later on a busy GPU will show higher elapsed times.
Files
clock.cu— kernel implementation and host driverREADME.md— this fileCMakeLists.txt— build configuration