Data Center ML Inference

Overview

With the proliferation of machine learning (ML) applications such as LLM-based ones, inference serving systems have become critical for the deployment of ML-powered applications. An inference serving system, typically deployed in a large-scale data center, takes existing pre-trained ML models and creates multiple instances of it to serve incoming inference requests at a high rate. The main challenge consists in minimizing the total resource consumption for serving a given number of inference requests while respecting the latency service-level objectives (SLOs) of the requests. Sometimes, the processing of an inference request can involve multiple pre-trained ML models chained or orchestrated in a dataflow graph, further complicating the problem. To address such a challenge, we are exploring how to leverage modern cloud computing paradigm namely serverless computing to enable more flexible resource allocation and investigate auto-scaling mechanisms to enable extreme elasticity for ML inference serving.

People

Funding

  • Deutsche Forschungsgemeinschaft (DFG)

Publications

No publications yet.