Historical manuscripts contain precious information regarding human being’s cultures and knowledge in many different domains. These valuable resources need to be preserved, maintained and shared. To this end, nowadays, repositories of digitized documents have been created from these manuscripts and as a result, huge volumes of scanned document images are available. Retrieving information from extremely large digitized resources is the next concern. On the other hand, the diversity in structure and layout of these ancient manuscripts, as well as the deterioration that is usual in historical documents, make extracting information and analyzing them a challenging task that is unlikely to be done by human beings.
Besides text contents, historical documents also contain some typographical objects such as illustrations and diagrams which carry visual knowledge and support the document content by providing an abstract view of the concepts. These objects help to understand the text content more productively. Identifying these typographical objects gives us information regarding the structure of documents. Moreover, information about typographical objects would be beneficial in creating indexes and metadata for large repositories of digitized documents.
Due to the recent promising success of deep learning approaches in computer vision applications, in this thesis, a CNN-based approach has been used to detect illustrations and diagrams and classify the document images based on the presence of these typographical objects. The proposed model has been applied on large datasets of historical document images of ECCO and NAS. These two datasets contain over 32 Million and 500,000 ancient document images respectively.
Similarly to the other real-world applications, in our target datasets, we had access to only a restricted number of labelled data as training and test set. Furthermore, our training dataset is imbalanced and there is an unequal distribution of classes. To deal with these issues and also to alleviate the resulting overfitting, we have empowered our approach with regularization and augmentation techniques to improve the performance. The final model achieved promising results on the large datasets of ECCO and NAS.
| Date | 18 Aug 2021 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Mohamed Cheriet (Supervisor) |
|---|
Hajabedi, Z. (Author),
Cheriet (Supervisor),
18 Aug 2021Student thesis: Master's thesis › Master in Engineering: Engineering