Skip to main navigation Skip to search Skip to main content

Source rate control in videoconferencing application using state–action–reward–state–action temporal difference reinforcement learning

  • Ali Rezagholizadeh

Student thesis: Master's thesisMaster in Engineering: Information Technology Engineering

Abstract

The portion of IP video over the Internet has exceeded 80% and is still increasing. Therefore, it is crucial for video service providers to satisfy the video quality experienced by the end users. Moreover, the Internet has some features such as heterogeneous components and diverse traffic management algorithms that do not guarantee any quality of service. Thus, it is the application’s responsibility to control its flow such to meet the users’ expectations. Several works proposed the sending rate adjustment algorithms from source rate control to control the packet size with the aim of providing good quality of service (QoS) or quality of experience (QoE). A few works are applied and evaluated in the context of real-time interactive multimedia transmission, i.e., for real-time communication. Such algorithms can be categorized as hand-crafted controlling rules and automatic control. The rule-based methods showed a lack of generalization to other types of networks motivating the study of automatic methods. Existing automatic control algorithms applied in real-time interactive multimedia applications, such as videoconferencing, are based on Reinforcement Learning (RL). In this work, we propose a tabular RL method,i.e., on-policy State–Action–Reward–State–Action (SARSA) as a one-step Temporal Difference method, to control the source rate, while the other RL-based works on real-time communication apply function approximation (FA) as a way of learning the model. Although tabular RL approaches consume more memory space to store its parameters, the updating of parameters is not approximated like what is done in FA and this can lead to faster convergence to an optimum value. Moreover, the few existing RL methods in a real-time communication environment formulate reward using some objective QoS metrics which can only estimate the QoE that a user experiences. To the best of our knowledge, we represent the first tabular RL method as a rate adjustment control for real-time interactive multimedia transmission. We also propose a new perspective in formulating the rate control problem in RL. We use PSNR, a widely used full reference visual quality metric, to evaluate the video quality perceived by the user. In this work, after suggesting a way to apply and evaluate a method in a general network environment, we propose a classical-queuing-model-based network simulator for our experiments. We apply and evaluate the proposed method in the videoconferencing context using an H.264 video codec over the simulated network environment. Our proposed method is evaluated and compared with Bounded Neural Network (BNN) over two configurations of the network simulator, i.e., with low and high available bitrate. The results show that SARSA outperforms BNN with significantly better PSNR, packet loss, and bitrate consumption.
Date27 Apr 2022
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorStéphane Coulombe (Supervisor) & Ghyslain Gagnon (Co-supervisor)

Cite this

'