The study of histological images is particularly important task for the diagnosis of breast cancer. This complex and time-consuming task could be assisted by the use of deep learning networks. Indeed, these deep networks are the state of the art in visual recognition, notably on histological images datasets. However, this type of model requires a large number of annotated images, which are not always available in histology. This work focused on assessing the impact of using data without annotation on the performance of current competitive models. In this project, we studied the behaviour of semi-supervised techniques Ladder Network, Virtual Adversarial Training, Mean Teacher and Deep Co-training to variations in labelled and unlabelled data. These techniques were evaluated on BACH and TUPAC datasets, composed of histological images to fight breast cancer. The results show that the contribution of non-annotated data is especially interesting when there is a small amount of annotated data. In addition, the order of performance of the different approaches depends on the database used. Then, different configurations with the Deep Co-training approach were studied, this approach being one of the most competitive techniques in this field. Tests have shown that adding additional models leads to better performance, although attention should be paid to corruption issues. Increasing the diversity of models allows, up to a certain limit, to improve the performance obtained. Finally, additional data generated by the BadGan approach were used. This last approach makes it possible to add an additional cost function that is beneficial, in the case studied, at the cost of increasing complexity.
| Date | 8 Oct 2019 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Éric Granger (Supervisor) |
|---|
Shorten, L. (Author),
Granger (Supervisor),
8 Oct 2019Student thesis: Master's thesis › Master in Engineering: Automated Manufacturing Engineering