Implementations disclose migrating and staging inference context data across protocol domain boundaries. An accelerator coupled to a switch sends a protocol request to a resource provisioning unit (RPU), the request associated with inference context data stored in a local memory of the accelerator and comprising a first physical address; the RPU translates the request to a CXL.mem M2S request comprising a second physical address, and sends the CXL.mem M2S request to a CXL memory device; wherein the inference context data is migrated between the local memory and the CXL memory device across different protocol domains. Other implementations describe a host reading inference context data from a CXL memory device and staging the data to an accelerator via an RPU that translates between CXL.mem and an accelerator interconnect protocol. The inference context data may include key-value cache data, model weights, or embeddings associated with inference workloads.
Full Text
What is claimed is: