TY - GEN
T1 - Generalizing HVAC Control With Domain Randomized Reinforcement Learning
AU - Boitel, Pablo
AU - Zhang, Kun
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/22
Y1 - 2026/6/22
N2 - Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones through a universal, non-invasive thermostat interface. The controller acts on temperature setpoints from zone measurements and forecasts, while a recurrent policy supports online adaptation under partial observability. Our main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows, which models correlated and multimodal distributions of thermal-zone parameters while maintaining physical plausibility and controllability. This produces a realistic and progressively adaptive training curriculum that improves transfer across buildings. We evaluate NOMAD-RL against a constant-setpoint PID controller, RL without domain randomization, and MPC in single- and multi-zone settings. NOMAD-RL consistently outperforms the PID and non-randomized RL baselines, and approaches the performance of a well-tuned MPC, especially in the more challenging multi-zone case. These results highlight the potential of adaptive, physics-informed domain randomization for robust and transferable HVAC control.
AB - Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones through a universal, non-invasive thermostat interface. The controller acts on temperature setpoints from zone measurements and forecasts, while a recurrent policy supports online adaptation under partial observability. Our main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows, which models correlated and multimodal distributions of thermal-zone parameters while maintaining physical plausibility and controllability. This produces a realistic and progressively adaptive training curriculum that improves transfer across buildings. We evaluate NOMAD-RL against a constant-setpoint PID controller, RL without domain randomization, and MPC in single- and multi-zone settings. NOMAD-RL consistently outperforms the PID and non-randomized RL baselines, and approaches the performance of a well-tuned MPC, especially in the more challenging multi-zone case. These results highlight the potential of adaptive, physics-informed domain randomization for robust and transferable HVAC control.
KW - Control
KW - Domain randomization
KW - HVAC
KW - Reinforcement learning
UR - https://www.scopus.com/pages/publications/105043741447
U2 - 10.1145/3765611.3815359
DO - 10.1145/3765611.3815359
M3 - Contribution to conference proceedings
AN - SCOPUS:105043741447
T3 - ACM Sustainability Week Companion 2026 - Proceedings of the 2026 ACM Sustainability Week
SP - 383
EP - 386
BT - ACM Sustainability Week Companion 2026 - Proceedings of the 2026 ACM Sustainability Week
PB - Association for Computing Machinery, Inc
T2 - 2026 ACM Sustainability Week, ACM Sustainability Week Companion 2026
Y2 - 22 June 2026 through 25 June 2026
ER -