Skip to main navigation Skip to search Skip to main content

Planification intelligente des flux entre centres de données

Translated title of the thesis: Information-agnostic inter-data center traffic scheduling
  • Meriem Amina Si Saber

Student thesis: Master's thesisMaster in Engineering: Engineering

Abstract

The important growth of data center users and inter-data center exchanges highlights the importance of efficient flow scheduling. However, the problem in the current inter-DC WANs is that the scenario where end users clearly specify the requirements for their transfers is unrealistic because of the absence of an interface that accomplishes this task. Consequently, and to maintain the level of performance that must be met, most existing works assume the availability of flow information and focus only on efficiency rather than generality, this environment is referred to as information-agnostic. The main idea behind this thesis is to overcome this challenge and propose an efficient alternative that understands traffic characteristics without requiring an expensive and non-scalable user interface. We thus propose an information-agnostic classification-based scheduling approach that differentiates inter-datacenter traffic classes in an online manner. First, we built a novel correlation-based classification module combining a cost-sensitive approach with a Bagged Random Forest ensemble algorithm (BRF), to address the interclass imbalance problem while meeting the critical time requirements of the different traffic classes. In order to calculate interflow correlations representing the rebalancing weights, we propose Reverse k-Nearest Neighbors (RkNN), a new algorithm that outperforms several data level, algorithm level and cost-sensitive strategies on four real-world datasets. The results reveal that the proposed algorithm outperforms most approaches in the different datasets in terms of precision, recall, F1 measure, AUC, and Kappa. The other algorithms resulted in either high precision with low recall or low precision and high recall causing congestion or resource over provisioning. The outcome of the classification module represents key parameters to characterize the entering traffic in the scheduling module described through an optimization problem, which besides guaranteeing better QoS performances measured by optimal Flow Completion Times, also targets cost minimization by proposing a cost-effective resource provisioning strategy for inter-data center networks. While other approaches result in high loss rates, our approach preserves the quality and the amount of exchanged traffic. Also, our approach outperforms existing approaches, particularly in the generalization aspect, since our scheduling approach is an online, multi-class scheduling method. The most time-consuming part of our method is tuning (which is still quite negligible).
Date13 Dec 2024
Original languageFrench
Awarding Institution
  • École de technologie supérieure
SupervisorMohamed Cheriet (Supervisor)

Cite this

'