My PhD has one overarching goal: reduce the administrative burden on caregivers. That goal has remained stable throughout my journey, even though the work required to reach it has looked very different from what I first imagined.
At the start, my requirements seemed simple: give me access to clinical data, provide a secure processing environment, and I can start building. I quickly learned that this was much easier said than done.
Even as a medical doctor, I do not have a therapeutic relationship with the patients whose data could support my research. Access therefore comes with strict legal, ethical, and organizational requirements. Those safeguards exist for good reason: patients’ privacy must be protected. The problem is that researchers and builders still lack a strong, reusable path through the practical bottlenecks.
A shared bottleneck
Clinical notes contain enormous research value, but they also contain names, dates, addresses, contact details, and other identifying information. Without a reliable way to remove or replace those identifiers, much of that text cannot safely be reused for secondary purposes such as research, quality improvement, or medical AI development.
Too often, every hospital or research group has to solve the same problem again. That slows down useful work, duplicates costs, and leads to solutions that are difficult to validate beyond a single institution.
De-identification initially appeared on my roadmap as one step towards accessing the data I needed. As the work grew, I realized it should not remain a private tool built for a single project. It could become shared infrastructure that helps future researchers avoid the same pain and frustration.
Introducing MedDeID
MedDeID is an open-source, local-first framework for clinical-text de-identification. It supports Dutch and English workflows and can be used through Python, a command-line interface, batch processing, Docker, or an HTTP service. The wider tooling covers the path from annotation and synthetic-data generation to model training, evaluation, inference, and pseudonymisation.
Local control is central to the project. Organizations can run MedDeID on premises, so sensitive clinical text does not have to leave their environment. Public models make it possible to start with synthetic training data, while institutions that are able to annotate their own data can train and validate models under their own governance.
The first results are encouraging. On an independently annotated benchmark of 300 Dutch hospital notes, a compact model trained on hospital data detected 98.9% of identifying text while redacting 0.24% of text outside the annotated identifiers. A model trained only on synthetic data detected 96.1%. On a separate set of 100 primary-care notes, the synthetic model achieved higher recall than the hospital-trained model: 90.3% compared with 87.0%. The full methods, limitations, and results are available in the MedDeID preprint.
These numbers are a starting point, not a claim that the problem is finished. Reliable deployment requires local validation, careful threshold selection, operational safeguards, and continued testing across different types of clinical text.
The ecosystem I want to help build
I want to see a future in which healthcare organizations across Flanders and the Netherlands can collaborate confidently with researchers and companies—improving the quality and efficiency of patient care without compromising patient privacy.
The most useful version of that future is decentralized. Each participating organization can maintain a local test set that never leaves its environment. New models can then be evaluated across institutions without centralizing sensitive data. This would tell us whether an update creates real value for the Dutch-language healthcare ecosystem, rather than only improving one benchmark.
I have spent the past year developing and testing the tools and models needed to make large-scale use possible, and I am making the work publicly available. But a robust ecosystem cannot be carried by one person. There are three concrete ways to contribute:
- Test the software and models. Tell me where predictions fail, which bugs you encounter, or what functionality is still missing for your workflow.
- Validate MedDeID on your own data. Local test sets across hospitals, primary care, and other settings are essential for understanding where new versions genuinely improve practice.
- Support the roadmap. AI and software development require sustained time and resources. Even relatively small contributions from stakeholders can help turn an open-source research project into dependable shared infrastructure.
Clinical-text de-identification is a problem we all share. We can each rebuild a partial solution in isolation, or combine our expertise and solve it more reliably and cost-effectively together.
You can read the documentation, try the live demo, explore the source code, or read the preprint. If you want to test, validate, fund, or build on MedDeID, contact me at stig.hellemans@uantwerpen.be.