MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

February 24, 2023

Conference Paper

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Abstract

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Published: February 24, 2023

Citation

Baheri B., J. Tronge, B. Fang, A. Li, V. Chaudhary, and Q. Guan. 2022. MARS: Malleable Actor-Critic Reinforcement Learning Scheduler. In Prceedings of the 41st International Performance Computing and Communications Conference (IPCCC 2022), November 11-13, 2022, Austin, TX, 217-226. Piscataway, New Jersey:IEEE. PNNL-SA-170367. doi:10.1109/IPCCC55026.2022.9894315

Research topics

High-Performance Computing

PNNL

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Abstract

Citation

Research topics

PNNL Plays Host to Annual AI Workshop

PNNL Showcases AI Innovations at National Competitiveness Expo

VecPAC: A Vectorizable and Precision-Aware CGRA