ISPA: Exploiting Intra-SM Parallelism in GPUs via Fine-Grained Resource Management
ISPA: Exploiting Intra-SM Parallelism in GPUs via Fine-Grained Resource Management
Han Zhao,Weihao Cui,Quan Chen,Minyi Guo
TLDR
ISPA designs persistent and elastic block to solve the thread slot and shared memory contention between co-located kernels and adopts the register allocation method to manage the register contention.
Abstract
Emerging GPUs have multiple Streaming Multiprocessors (SM), while each SM is comprised of CUDA Cores and Tensor Cores. While CUDA Cores do the general computation, Tensor Cores are designed to speed up matrix multiplication for deep learning applications. However, a GPU kernel often either uses CUDA Cores or Tensor Cores, leaving the other processing units idle. Although many prior research works have been proposed to co-locate kernels to improve GPU utilization, they cannot leverage the Intra-SM CUDA Core-Tensor Core Parallelism. Specifically, ISPA designs persistent and elastic block to solve the thread slot and shared memory contention between co-located kernels. ISPA also adopts the register allocation method to manage the register contention. These resource management methods are applicable for both white-box kernels and
