Zero-Touch Networks (ZTNs) represent a cutting-edge paradigm shift towards fully automated and intelligent network management, providing the automation and intelligence needed to handle the complexity, scale, and dynamic behavior of next-generation mobile systems, including the sixth generation (6G). Artificial intelligence (AI) functions as the brain and backbone of ZTNs. In particular, deep reinforcement learning (DRL) algorithms have demonstrated strong potential for improving operational efficiency and supporting intelligent decision-making. Zero-touch management is particularly critical in transformative technologies such as network slicing (NS), which enables the creation of isolated logical networks on a shared physical infrastructure. Each logical network is designed to meet distinct quality of service (QoS) requirements, which makes NS indispensable for both current and future mobile generations. However, realizing ZTNs within a NS framework is associated with several challenges, especially in the radio access network (RAN) domain. Among the many difficulties encountered in the RAN domain, the following three challenges stand out as particularly critical: (i) managing inter- and intra-slice resource allocation; (ii) ensuring the security of inter-slice resource sharing; and (iii) achieving coordinated and concurrent management of radio resources at both inter- and intra-slice levels. These challenges are further exacerbated by the dynamic nature of wireless channels, scarcity of radio resources, diverse service-level agreements (SLAs), massive device densification, random traffic arrivals, and various network imperfections such as imperfect channel state information (CSI), imperfect orthogonal frequency division multiple access (OFDMA), and hardware impairments (HWIs).
In this context, the present thesis makes several key contributions to address these challenges. First, we investigate intra-slice resource allocation for RAN slicing by introducing an efficient self-optimizing (SO) scheme for a multi-user multiple-input single-output (MU-MISO) system, named PABSO-DRL. The proposed PABSO-DRL scheme dynamically and jointly manages power allocation and beamforming to ensure high data rates for enhanced mobile broadband (eMBB) while concurrently ensuring high reliability required by ultra-reliable low-latency communications (uRLLC). The scheme is designed to handle heterogeneous QoS requirements using a multi-agent deep Q-network (DQN) approach, while accounting for imperfect CSI, incomplete OFDMA isolation, and time-varying dynamics of the RAN environment.
Next, we focus on investigating inter-slice resource allocation. To this end, we propose a secure self-optimizing scheme based on a cooperative multi-actor–critic (CoMA2C) framework, referred to as SO-CoMA2C, to manage multiple radio resources (power and bandwidth) across a set of heterogeneous slices based on the fluctuating traffic load and in the presence of HWIs in open RAN (O-RAN). The main goal of the proposed scheme is maximizing the spectral efficiency while ensuring the SLA of each slice. Furthermore, we ensure the security of the allocation by integrating the Advanced Encryption Standard (AES) algorithm into the proposed scheme.
Finally, we develop a hierarchical self-optimizing framework aimed at maximizing the long-term QoS and spectral efficiency of heterogeneous services. The proposed framework adopts a two-layer strategy implemented through the following two complementary slicing manage ment schemes: (i) a cooperative multi-actor–critic (CoMA2C) scheme that allocates power and bandwidth across heterogeneous slices at a large timescale and (ii) a multi-agent DQN (MADQN)scheme that manages power and beamforming for active users within each slice at a small timescale. This design accounts for HWIs, traffic fluctuations and channel variations. Furthermore, a promising rate-splitting multiple-access (RSMA)-based scheme is investigated to further enhance the performance within each heterogeneous service. Beyond performance enhancement, the proposed framework also addresses coordination efficiency by minimizing overheads. In particular, the inter-slice scheme is triggered only when substantial changes occur in the traffic loads of the hosted slices. This design reduces the overall system overhead in terms of memory consumption, training time, and related operational costs.
Simulation results based on realistic system model assumptions demonstrate that the proposed cooperative multiple DRL-based approaches framework outperform baseline methods and successfully meet diverse SLA requirements in dynamic O-RAN deployment environments. The results of extensive evaluations further show that our proposed schemes provide a pathway towards more comprehensive and predictive resource management, ensuring robust performance under uncertain network conditions and imperfect operational environments. Moreover, the results highlight flexibility, robustness, and possible scalability of the proposed design while achieving a lower computational overhead as compared to benchmark baselines. Overall, the findings validate the potential of employing multiple cooperative DRL agents to enable automated, scalable, and intelligent RAN slicing in next-generation wireless communication networks.
| Date | 30 Mar 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Kuljeet Kaur (Supervisor) & Georges Kaddoum (Co-supervisor) |
|---|
Sabr, O. (Author),
Kaur (Supervisor) &
Kaddoum (Co-supervisor),
30 Mar 2026Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering