Skip to main navigation Skip to search Skip to main content

Representation learning for document image analysis with practical considerations

  • Sherif Abuelwafa

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

This thesis sets up reliable document image representation learning approaches that can stand up to the practical real-world challenges currently facing the document image analysis field. Particularly, two challenges are tackled, performing efficient analysis on large-scale datasets and adapting to the scarcity of labeled training data. The proposed approaches aim to improve the performance of the document image analysis processes when applied to real-world use-cases. For this purpose, we address the practical challenges in two main tasks of document image analysis, classification and semantic segmentation. Current document representation approaches usually focus on use-cases with an unrealistic assumption that any document representation can well generalize when applied on large-scale document datasets. Therefore, we first propose a document representation approach for the task of document classification that can generalize well for such large-scale datasets. The classification process in this task is based on the existence of a distinctive visual local object (e.g., footnote) within the document image, which is of high relevance to various use-cases in the document image analysis field. The proposed approach is applied to datasets that contain more than 32 million document images and show a consistent reliable performance across various datasets using less than 0.07% of the dataset’s samples for training. Many recent representation learning approaches are based on supervised feature learning, which requires a large amount of annotated training document images to obtain reliable performance. Meanwhile, in real-world use-cases, the available amount of labeled data is very limited and scarce, while a large amount of unlabeled data is often abundant. We, therefore, propose a document representation learning approach for the task of document classification, which is capable of learning features solely from unlabeled data, and without any dependence on hand-crafted features. Unlike our earlier work above, the classification process in this work is based on the global context of the document image. Our approach utilizes unlabeled data to learn a representation that is used later for document classification either with few labeled data or with no labeled data. The efficiency of the proposed approach and its associated performance boost is demonstrated with the obtained experimental results. Considering each previously classified document, we finally propose a document representation learning approach for the task of document semantic segmentation to obtain an additional interpretation of that document’s content and prepare it for further analysis tasks. This approach is capable of learning features from unlabeled data without requiring annotated data, datasetdependant heuristics techniques, or textual information. In addition, it tackles the common challenge of having high inter-class similarities between different semantic classes. Experiments on various public datasets demonstrate the effectiveness of our proposed approach by yielding better results than earlier approaches.
Date15 Dec 2021
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorMohamed Cheriet (Supervisor)

Cite this

'