/Reducing Latency In Highly Scalable Hpc Applications Via Accelerator-resident Runtime Management
Abstract

Methods, systems, and computer-readable media for runtime management using accelerator-resident managers. A first manager resident in a first accelerator accesses a runtime representation stored in shared memory accessible by resident managers of a plurality of accelerators. The runtime representation identifies dependencies among kernels of an application and assignments of the kernels to Accelerated Processing Units (APUs) of the plurality of accelerators. The first manager launches kernels assigned to APUs of the first accelerator and updates the runtime representation based on availability information stored in the shared memory for APUs of the first accelerator and APUs of a second accelerator. The first manager updates the runtime representation to change an assignment of at least one kernel from an APU of the second accelerator to an APU of the first accelerator for a later execution iteration.

Full Text

What is claimed is:

Methods, systems, and computer-readable media for runtime management using accelerator-resident managers. A first manager resident in a first accelerator accesses a runtime representation stored in shared memory accessible by resident managers of a plurality of accelerators. The runtime representation identifies dependencies among kernels of an application and assignments of the kernels to Accelerated Processing Units (APUs) of the plurality of accelerators. The first manager launches kernels assigned to APUs of the first accelerator and updates the runtime representation based on availability information stored in the shared memory for APUs of the first accelerator and APUs of a second accelerator. The first manager updates the runtime representation to change an assignment of at least one kernel from an APU of the second accelerator to an APU of the first accelerator for a later execution iteration.
Timeline
Filed
05/22/2026
Published
09/17/2026
Granted
Not Available
IPC Codes(4)
G06F 9/48:Program initiating; Program switching, e.g. by interrupt
G06F 9/38:Concurrent instruction execution, e.g. pipeline or look ahead
G06F 9/50:Allocation of resources, e.g. of the central processing unit [CPU]