Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
This lecture covers GPU memory hierarchy, including global, local, shared memory, and caches. It explains the CUDA processing flow, GPU optimizations, and control-flow divergence. The instructor discusses strategies to optimize algorithms for GPUs, exploit shared memory, and coalesce memory accesses. Various techniques to efficiently use parallelism and resources on GPUs are explored, such as reduction operations and addressing bank conflicts. The lecture concludes with a focus on scalability with array size and a summary of optimizing code for GPUs.
This video is available exclusively on Mediaspace for a restricted audience. Please log in to MediaSpace to access it if you have the necessary permissions.
Watch on Mediaspace