Image segmentation is vital in many clinical and research applications, such as disease characterizations, surgical planning, diagnostic measurements, and shape analysis. However, manual delineation is time-consuming, may require expertise, and is subject to variability. Automated algorithms offer a solution to these limitations, thereby assisting clinical and research workflow. Recent deep learning-based techniques have successfully provided high-quality automated segmentation, generally using a substantial amount of labeled data. However, the labels can be ambiguous or unreliable. This thesis tackles these challenges with the primary objective of developing uncertainty-aware tools that can aid in training image segmentation networks. Particularly, the first objective proposes an intensity-based soft labeling strategy to tackle potential ambiguities in the annotation. The second objective presents an anatomically-aware uncertainty estimation to guide the segmentation network under limited supervision. The third objective proposes an attention-based representation for weakly supervised segmentation. The findings from these research objectives have resulted in three journals, two peer-reviewed conference publications, and a short conference article. The contributions of each research objective are summarized below.
In the first objective, we propose a Geodesic Label Smoothing (GeoLS) approach that captures image intensity details within the soft labeling process. The image intensities convey information that could clear potential ambiguities in the annotation. However, existing soft-labeling methods rely only on segmentation masks, ignoring the underlying image context associated with the label. We leverage the geodesic distance transform to capture the intensity variations between pixels. The generated maps modify the hard labels to obtain new intensity-based soft labels. The resulting geodesic soft labels better model spatial and class-wise relationships as they capture the variations of image gradients across classes and anatomy. The benefits of our intensity-based geodesic soft labels are assessed on three diverse sets of publicly accessible segmentation datasets. Our experimental results show that the proposed method consistently improves the segmentation accuracy compared to state-of-the-art soft-labeling techniques in terms of the Dice similarity and Hausdorff distance.
The second objective aims to estimate uncertainty by leveraging anatomically-aware representation during training of segmentation network under semi-supervised settings. Specifically, an anatomically-aware representation is first learned to model the available segmentation masks. The learned representation maps a segmentation prediction into an anatomically plausible segmentation. The deviation from the plausible segmentation aids in estimating the underlying pixel-level uncertainty maps. These maps filter the unreliable target regions to guide the segmentation network. The proposed method consequently estimates the uncertainty using a single inference from our representation, reducing the total computation during training compared to existing uncertainty-aware approaches. We evaluate our method on two publicly available segmentation datasets. Our anatomically-aware approach improves the segmentation accuracy over the state-of-the-art semi-supervised methods in terms of two commonly used evaluation measures.
Finally, the third objective proposes to learn an attention-based dynamic representation for medical image analysis. Particularly, a representation is learned by integrating an attention module into an embedding network. This integrated attention mechanism provides a direct visual insight into the discriminative features of the embedding network. Furthermore, a single metric learner is inadequate for learning a variety of object attributes in images, such as color, shape, or artifacts. Instead, multiple metric learners could aid in learning different aspects of these attributes in subspaces of an overarching embedding. However, number of learners is to be found empirically for each new dataset. We, therefore, present a dynamical subspace learner, which removes the need to know apriori the number of learners in the multiple learners approach. The benefits of our attention-based dynamic representation are evaluated in the application of weakly supervised segmentation, image clustering, and image retrieval. Our method provides an attention map directly during inference to illustrate the visual interpretability of the embedding features. These attention maps offer proxy labels, improving the segmentation accuracy by up to 15% in the Dice score compared to state-of-the-art interpretation techniques. Moreover, our method achieves competitive results compared to the multiple metric learner approach and significantly outperforms the classification network in terms of clustering and retrieval scores on three different public benchmark datasets.
The research work described in this thesis advances medical image segmentation across full, semi, and weak supervision. Our intensity-based soft labels enhance the segmentation, especially in challenging regions. Our anatomically-aware uncertainty estimation approach effectively uses limited annotation, reducing the need for extensive labeling. The attention-based representation approach provides structured data organization and visual interpretability, enabling segmentation with only image-level labels. This thesis presents new tools that assist clinicians and researchers by providing faster, consistent, and accurate delineation of target objects.
| Date | 3 Aug 2024 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Hervé Lombaert (Supervisor) & José Dolz (Co-supervisor) |
|---|
Adiga Vasudeva, S. (Author),
Lombaert (Supervisor) &
Dolz (Co-supervisor),
3 Aug 2024Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering