On 24 August, I presented the goal and work of my PhD on “Automating Structuring Clinical Texts” at the BIHR Summer Day. Preparing the talk gave me a welcome opportunity to step back from the technical details and reflect on a larger question: where are we today, and where do we want to go?
What is BIHR?
The Belgian Integrated Health Record (BIHR) is a national program to make health and wellbeing data available to citizens and healthcare professionals in a secure, standardized and context-aware way. It is not one centralized patient record. Instead, it aims to connect existing systems and trusted data sources so that the right information can be used at the right time, directly within the applications people already work with.
That ambition matters. Better access to reusable data can improve individual care, collaboration between institutions, population health management, research and policy. But it also comes with a difficult design question: how do we create useful structure without increasing the administrative burden on healthcare professionals?
Where should structuring happen?
Healthcare information systems have already moved from paper to digital documents. We are now entering a second transition, from documents towards individual data points. That can make information more findable, accessible, interoperable and reusable. It can help a healthcare professional find an answer without reading a patient record that has grown over an entire lifetime, support population health management and improve collaboration across institutions.
But structuring also has a cost. More forms can mean more administrative work and less time for care. When I look at unstructured clinical text, I see three possible directions:
At one end, we can keep the source text as it is and use search or AI agents to reconstruct the necessary context for every question. That flexibility is valuable, and I do not think this approach is necessarily wrong. Its disadvantage is that the structure created to answer a question is usually discarded afterwards, meaning the same work may have to be repeated and opportunities for cost-efficient reuse are lost.
At the other end, we can try to capture everything in forms without free text. That would theoretically solve the structuring problem at its source, but it creates a data-entry and motivation problem. Medicine is too complex to fit entirely into forms, and healthcare professionals do not naturally think that way.
The middle path I am exploring is to preserve the clinical narrative while building a persistent, standardized layer on top of it: structure once, then reuse. AI can help produce that layer, but it is only valuable when it remains faithful to the source text, is normalized enough to query and is explicit enough to support reasoning.
Healthcare professionals should not have to think about data structure while documenting care. The structuring should be handled for them.
This is the moonshot behind one of my projects. The aim is to connect information across all clinical notes in the patient record, while retaining where every fact came from. A question such as “Was pleural effusion present?” should then be answerable across the record without losing the evidence or clinical context behind the answer.
What a pharmacy pilot taught us
The second, more pragmatic project, developed with Dirk Broeckx, focused on normalizing medication dosage instructions into the FHIR-based BeDosage format. Pharmacies currently write these instructions in many different ways. Our idea is to create a set of clear descriptions for each medication group, with a validated structure behind them that software can use automatically.
A pilot with one pharmacy started from 40,412 unique free-text instructions. Within Virtual Medicinal Product groups, 33,667 texts were included and 98.7% could be structured automatically in just over ten hours of AI-agent processing, at a cost of roughly €300.
Those numbers are encouraging, but the limitations are equally important. Expert validation remains necessary, and dosage instructions have a long tail of rare patterns: only 9.5% of the texts shared their meaning with another instruction. The exercise also revealed a second opportunity beyond interoperability: clearer instructions for patients and caregivers.
Where we go from here
My main reflection after the BIHR Summer Day is that the future will not be built by choosing between clinical text and structured data. We need both. Standards such as FHIR, SNOMED CT and LOINC give information a shared language; AI can help produce that structure without forcing every healthcare professional into increasingly rigid forms.
Getting there will require more than good models. It will require expert validation, traceability, trust and close collaboration between citizens, healthcare professionals, public authorities, researchers and industry. If we get that combination right, we can make health data more reusable while giving time back to care.
Download the presentation slides
Thank you
A very big thank you to Dirk Broeckx, both for our collaboration on the dosage project and for his continued work on BIHR. I am equally grateful to Ilke Montag and Jan De Maeseneer. Together, Dirk, Ilke and Jan helped found BIHR and continue to turn its ambition into practical progress.
Finally, thank you to Karlien Erauw and Agoria for making the BIHR Summer Day possible and creating a place where people from across Belgian healthcare could exchange ideas about what comes next.