Skip to main navigation Skip to search Skip to main content

L’approche d’apprentissage automatique pour l’optimisation de la robotique en essaim dans l’entrepôt automatisé servi par la communication 5G

Translated title of the thesis: A machine learning approach for optimizing swarm robotics in 5G-enabled automated warehouses
  • Manh Tai Ho

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

The fifth generation wireless network (5G) provides high-speed, low-latency and high-reliability connections that can meet the requirements of the Industrial Internet of Things (IIoT) in industrial automation, especially for robotic control. In intelligent storage, robotics play an essential role in achieving intelligent logistics solutions that include organization, planning, control and intelligent execution of goods/items flow in the warehouse. Recent advance in wireless communications and battery technologies make it possible to replace many human workers with robotic systems in order to reduce labor costs, improve warehouse work efficiency, and increase reliability. However, the deployment of swarm robotics poses new challenges in terms of control in order to coordinate many types of resources in the warehouse to deliver 5G services for robotics and plan tasks for robots. In particular, efficient wireless resource management in a highly dynamic 5G network like in an automated warehouse is a challenging problem because extreme reliability and low latency with high mobility of robots are not efficiently solvable by traditional optimization approach. To this end, in this thesis, we tackle the two main challenges of an automated warehouse simultaneously : i) provisioning 5G services and ii) controlling swarm robotics. The main contributions of this thesis are as follows : 1. Firstly, we formulate the problem of provisioning 5G services to serve swarm robotics in the automated warehouse as a joint clustering of coordinated multi-points (CoMP) and ultra-reliable low-latency communication (URLLC) 5G beamforming. Traditional iterative optimization approaches are not efficient in solving this non-convex real-time problem due to their high computational time. We thus propose a CoMP clustering algorithm using the combination of game theory and deep reinforcement learning Proximal Policy Optimization (PPO )method to obtain an approximate stationary solution to the global optimal solution. 2. Secondly, we study the problem of controlling autonomous heterogeneous robotic systems in the automated warehouse. We formulate a non-convex long-term queue control optimization problem to minimize the task queue length in the warehouse. Traditional solutions based on optimization approaches are not effective in managing the stochastic nature of the goods/tasks flow and a large number of robots in the system. Therefore, we propose a task scheduling algorithm based on the PPO method to find an optimal task planning policy. Due to the system’s heterogeneity, we propose a federated proximal weighted learning algorithm to implement the decentralized PPO algorithm which improves the performance of distributed PPO agents deployed in different geographically distributed warehouses. Our simulation results demonstrate the effectiveness of our proposed algorithm compared to existing methods. 3. Finally, we propose a model for provisioning 5G services and controlling swarm robotics simultaneously in an automated warehouse. We aim to maximize long-term energy efficiency while meeting the energy consumption constraint of robots and the ultra-reliable low-latency communication (URLLC) requirements between the central controller and the swarm robotics. This optimization model is non-convex since the achievable rate and decoding error probability with short block length are neither convex nor concave. We propose a deep reinforcement learning approach that uses the Deep Deterministic Policy Gradient (DDPG) method and the Convolutional Neural Network (CNN) to obtain an optimal stationary control policy which consists of a number of continuous and discrete actions. The experimental results show that our proposed multi-agent DDPG algorithm outperforms existing solutions in the state-of-the-art in terms of error probability and energy efficiency.
Date20 Jul 2023
Original languageFrench
Awarding Institution
  • École de technologie supérieure
SupervisorMohamed Cheriet (Supervisor) & Kim Khoa Nguyen (Co-supervisor)

Cite this

'