Efficient Job Scheduling Using Reinforcement Learning Approach in Cluster Computing
Keywords:
Job Scheduling, Reinforcement Learning, Machine Learning, GPU Cluster Computing, Proximal Policy Optimization, XGBoost, HPCAbstract
Job scheduling in high-performance computing (HPC) clusters directly affects user waiting times and GPU hardware utilization. Conventional heuristic schedulers such as First Come First Serve (FCFS), Shortest Job First (SJF), and backfilling depend on static rules that cannot adapt to dynamic workloads. While machine learning (ML) prediction models and reinforcement learning (RL) schedulers have each shown capability, neither fully resolves the problem: prediction models cannot respond to real-time cluster state changes, while RL agents explore vast action spaces without prior experience. This paper offers a hybrid ML+RL scheduling framework for GPU cluster computing. An XGBoost regression model is trained offline on real job execution records from the MIT SuperCloud dataset, incorporating NVIDIA DCGM GPU hardware data logging. The accomplished predictor is implanted as a live per-step component inside a custom Gymnasium simulation environment, generating node-specific runtime predictions at each scheduling decision. A Proximal Policy Optimization (PPO) agent receives the full live state of all 225 cluster nodes together with top 4 predicted priority nodes, allowing the agent to balance ML supervision with real-time cluster awareness. Evaluated on 11,001 held out test jobs, the proposed scheduler reduces average waiting time by 98.9% (from 121.5 hours to 1.3 hours), decreases makespan by 84.6% (from 3,453 hours to 533 hours), and improves GPU utilization from 7.24% to 46.90%, outperforming all baselines including the SJF oracle with perfect runtime knowledge.