The healthcare sector faces major interoperability challenges. This complexity stems from the diversity of data representation from multiple sources, such as electronic records, biomedical equipment or devices, and connected objects. Interoperability is a key element of digital transformation and is the subject of intensified research aimed at progressive standardization to facilitate exchanges between the various information systems used in patient care pathways.
The main objective of this thesis is to contribute to promoting the interoperability of standardized, high-quality information concerning a patient's medical conditions in the Frenchspeaking Canadian and Quebec context.
The first step presents an overview of the publications and state of the art concerning the HL7 FHIR interoperability standard, a standard specifically developed for the healthcare sector, which aims to ensure shared semantics in exchanges between different systems exchanging data. The aim here is to identify the data targeted by this semantization as well as the approaches used to date. This review presents a classification of the methodologies used at different stages of the data exchange process: extraction, annotation, modeling, mapping, and transformation. The conclusions of this study indicate that machine learning and natural language processing (ML/NLP), recently introduced for unstructured data, offer promising avenues. In addition, the use of exchange standardization based on terminologies such as SNOMED is becoming an essential approach to ensuring semantic interoperability in the future.
Secondly, a model promoting the semantic interoperability of data concerning the patient’s medical conditions, particularly allergies, is proposed. A hybrid framework is proposed that uses both natural language processing techniques and rules to extract and transform free-form French clinical texts and then convert them to a standardized model based on the international HL7 FHIR standard and SNOMED CT terminology. To do this, the FRASIMED corpus of open French-language medical data was used and mapped to the international patient summary model using the MedSpacy platform. The experimental results are promising and could help improve the semantic interoperability of unstructured data in French.
Subsequently, the following three-step process was carried out toward making NLP-to-FHIR systems production-ready for clinical deployment: (1) analyze current methodologies for verification and validation of NLP-to-FHIR pipelines; (2) conduct root-causes analysis identifying potential sources of bias in these systems; and (3) develop a comprehensive, fivephased research roadmap that research teams can follow to systematically address verification and validation challenges ((a) Governance and Privacy Framework; (b) Dataset Availability and Diversity; (c) Data Selection and Annotation Methodology; (d) Linguistic Development; and (e) Validation, Bias Detection, and Refinement).
In conclusion, unstructured French-language data concerning a patient's health status can be converted and matched using standardized healthcare terminology to ensure simple, highquality exchange between the various systems involved in patient care. A methodology framework is required for verification and validation process for a successful deployment in clinical setting.
| Date | 14 May 2026 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Alain April (Supervisor) & Alain Abran (Co-supervisor) |
|---|
Amar, F. (Author),
April (Supervisor) &
Abran (Co-supervisor),
14 May 2026Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering