The rising inflow of people living in megacities has demanded a smart approach to create a sustainable infrastructure to urban areas and provide more efficient services. Buildings automatically opening front doors, lights flicking on, heating and cooling systems adjusting by themselves, cameras tracking the traffic among other situations, are examples of countless arrays of sensors interacting with spaces and people. These sensors and systems dispatch huge volume of data into software platform hundreds of times per minute, demanding high velocity in processing and storing this information through the network, Internet of Things. This digital behaviour gives a foundation to the concept of smart city. While the data variety, volume and velocity are the Big Data definition, both are anchored on the development of Information and Communication Technology.
However, the available data are distributed and collectively aggregated and its fusion might reveal patterns that would not be possible if the data were analyzed separately. The main assumption in data fusion review is that the big picture from fused information allows the optimization of electricity flow through the power grid, supporting transportation networks moving, watching over people’s health and safety, and much more. Yet, according to the literature review, there are two major problems related to fusion data. The first issue is, any system (e.g., application and platform) is susceptible to produce data that might be inaccurate, insufficient, duplicated, incorrect, inconsistent, ambiguous, besides the missing values. The second, in most cases, fusion solutions focus on predefined mining strategy and supervised tasks.
To overcome the first issue, it is noted that missing values can significantly affect the result of analyses and decision making in any field and that the two major approaches to deal with this issue are statistical and model-based methods. Whereas the former brings bias to the analyses, the latter is usually designed for specific cases. To cope with the limitations of both methods, we present a stacked ensemble framework integrating the adaptive random forest algorithm, the Jaccard index, and Bayesian probability. Considering the challenge that the heterogeneous and distributed data from multiple sources represents, we have built a model that supports different data types: continuous, discrete, categorical, and binary.
Aiming to overcome the second limitation of data fusion, we introduce a fusion ensemble learning model for multiple and heterogeneous datasets stacking the Restricted Boltzmann Machine, which gather latent features of the unlabelled datasets and the matrix tri factorization algorithm to fuse the features in a block-matrix structure. Combining such techniques, we were able to design an efficient knowledge discovery tool. The evaluation of both proposed solutions, missing values imputation and fusion data has shown that our ensemble learning model produces encouraging and competitive results, overcoming the limitations previously found in the literature review.
| Date | 12 Feb 2021 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Mohamed Cheriet (Supervisor) |
|---|
Costa Carvalho, A. L. (Author),
Cheriet (Supervisor),
12 Feb 2021Student thesis: Master's thesis › Master in Engineering: Information Technology Engineering