Skip to main navigation Skip to search Skip to main content

Reduced supervision methods for medical image segmentation

  • Jizong Peng

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Medical image segmentation is an important pre-processing step in computer-aided diagnosis systems. Methods based on neural networks have demonstrated state-of-the-art performance on various segmentation tasks with different image modalities. Despite their unprecedented success, neural networks usually require a large amount of reliable densely-labeled data. However, obtaining this data is a laborious and costly process, which often requires medical experts, and annotations obtained by this process can be prone to errors. To mitigate the scarcity of denselyannotated data, a promising research direction is to exploit images with reduced supervision signals. These reduced types of supervision usually consist of image tags, points, scribbles or bounding boxes as annotations, however images without any form of supervision can also be leveraged. Recent works have also tried to combine these weak annotations with anatomical priors for regions of interest to guide the network prediction towards anatomically-plausible solutions. The main objective of this thesis is to develop accurate algorithms for medical image segmentation which can learn with reduced supervision. Specifically, we first propose a weakly-supervised segmentation algorithm that learns from scribbles and discrete anatomical constraints. Next, we present a segmentation framework, based on deep ensemble learning, that enables the collaborative training of multiple segmentation networks with a small set of labeled images and a larger amount of unlabeled ones. In another contribution of the thesis, we solve this problem by introducing an algorithm based on mutual information that uses unlabeled images to regularize the feature representation in the network and boost segmentation accuracy when few images are densely annotated. We then propose a method based on representation learning that exploits the information from unlabeled images with various medical meta-labels. As the last contribution, we demonstrate a boundary-aware information maximization method for dense representation pre-training, which acquires meaningful anatomical structure cues from unlabeled images and thus significantly improving segmentation accuracy given a small set of labeled images. This thesis has resulted in three journal publications, two papers in peer-reviewed international conferences, two short papers presented in medical imaging workshops, as well as one paper currently under review. The specific objectives of this thesis are presented below. As our first objective, we propose an efficient strategy for weakly-supervised segmentation to impose constraints or regularization priors on target regions. This segmentation method is among the first to employ discrete optimization with a neural network, which enables the network obtain a more accurate solution faster. Our proposed method is based on the alternating direction method of multipliers (ADMM) algorithm and trains a CNN with discrete constraints and regularization priors. The performance of this method is validated on the segmentation of medical images with few annotated pixels, as well as discrete constraints of the size and boundary regularity of segmented regions. Experiments on two benchmark datasets showed our method to provide significant improvements compared to existing approaches in terms of segmentation accuracy, constraint satisfaction and convergence speed. In our second objective, we focus on semi-supervised segmentation and propose an algorithm based on ensemble learning. This method trains multiple models with a reduced number of annotated images, as well as with non-annotated images used for exchanging information between the trained models. To enforce the diversity of models, an adversarial loss is also designed. The effectiveness of this method is assessed on three medical image segmentation tasks covering different modalities, where it boosts segmentation accuracy when very few labeled images are used. The impact of our diversity loss is studied by visualizing the images generated by the adversarial training. We also explore the performance gains obtained with an ensemble containing more than two models, showing that adding models can improve results at the cost of increased computations. In our third objective, a novel semi-supervised segmentation method is proposed. This method leverages the mutual information computed on categorical distributions to achieve both global representation invariance and spatial smoothness. In this method, we maximize the mutual information for intermediate feature embeddings that are taken from both the encoder and decoder of a segmentation network. A loss on global mutual information is employed on the encoder to enforce invariance towards geometric transformations. Likewise, a loss on the local mutual information is also used to promote spatial consistency in feature maps from the decoder, and thus to provide a smoother segmentation. The advantages of our method are evaluated on four challenging publicly-available datasets for medical image segmentation. Experimental results show our method to outperform recently-proposed approaches for semi-supervised segmentation and provide an accuracy near to full supervision while requiring very few annotated images. In our fourth objective, we aim to acquire a useful representation by employing unlabeled images. Specifically, we adapt standard contrastive learning to train the encoder of the network for different pre-defined tasks: determining if two images of a 3D MRI scan are from the same position, same subject, or were acquired at the same moment of the cardiac cycle. In order to mitigate the noise presented in these meta-labels, an effective self-paced learning strategy is then proposed in contrastive learning, which yields a more robust representation and thus performance improvements for the segmentation tasks. We verify the quality of the proposed method on five medical image segmentation datasets, indicating clearly the advantage of our proposed self-paced mechanism using the meta-labels. We present, in our last objective, a cluster-based method to learn discriminative representations for dense feature maps. This approach employs an improved mutual information loss to group dense embeddings into multiple balanced and confident clusters. A boundary-aware loss based on pixel-wise cross-correlation is also enforced to align the cluster boundaries to image edges, which regularizes different clusters to correspond to anatomical structures in the image. Our proposed losses complement the contrastive loss presented in the previous objective, and their combination leads to remarkable improvements for the downstream segmentation tasks. Experimental results obtained from two clinically-relevant benchmark datasets clearly indicate the advantage of our method over contrastive-based counterparts, leading to a segmentation precision close to that of full-supervision, given only a few densely-annotated examples.
Date9 Jun 2022
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorChristian Desrosiers (Supervisor) & Marco Pedersoli (Co-supervisor)

Cite this

'