Healthcare AI · Compliance · Annotation
Medical Image Annotation With Documented Provenance
The evidence record a notified body or an auditor expects when AI trained on annotated medical images reaches the market, and why that record cannot be assembled after the fact.
In short
The August 2026 deadline does not apply to medical device AI
- Diagnostic imaging software that performs clinical analysis is high risk under Annex I of the EU AI Act, not Annex III. It follows a later clock.
- The obligation date for these systems was 2 August 2027, and it is now moving to 2 August 2028 under the Digital Omnibus. The widely quoted August 2026 date is for a different category.
- The duty that bites is data governance: proving how your training data was built. For medical images, that means provenance.
- Provenance is the one artefact you cannot backfill. If it was not captured while the labelling happened, it cannot be written in later.
What documented provenance means in medical image annotation
Documented provenance in medical image annotation is the complete, verifiable record of how every label on every image was produced. It captures where each image came from and under what lawful basis, which version of the labelling instructions applied, who annotated the image and with what qualifications, how disagreements between annotators were resolved, and how the result was checked. In short, it is the evidence that lets an auditor or a notified body reconstruct, after the fact, exactly how a training dataset was built and whether the labels can be trusted.
For AI that is itself a medical device, or that sits inside one, this record is not optional polish. It is the substrate of the technical documentation that the Medical Device Regulation and the EU AI Act both expect, and it is the part a vendor cannot produce later if it was not captured while the work was happening.
Provenance is the one artefact you cannot backfill
Most compliance gaps can be closed late. A model can be retrained. A data governance policy can be written this quarter to cover a process that began last year. A risk assessment can be drafted in a week. Provenance is the exception. You cannot reconstruct who annotated a particular scan in 2026, under which version of the guideline, with what adjudication, and at what measured level of agreement, unless you recorded it at the time.
This is why the lead time matters more than the headline date. A team that starts capturing a defensible provenance record now will have years of clean evidence when an assessment arrives. A team that waits will have a dataset it cannot fully account for, and retrofitting a credible history into an annotation programme is somewhere between expensive and impossible. The record has to be built as the labels are made, not described afterwards.
Where the obligations actually sit, and when
Two regulatory frameworks apply to AI driven medical imaging software, and they run on different clocks. Conflating them is the most common error in the market, so it is worth separating them cleanly.
The Medical Device Regulation and the IVDR apply now
If your imaging software performs a clinical function, it is a medical device. The Medical Device Regulation and the In Vitro Diagnostic Regulation already require data governance, clinical evidence, traceability, and surveillance once a device is on the market. These duties are in force today, independent of the AI Act, and they already reach the quality and documentation of the data used to develop the device.
The EU AI Act adds a second, parallel layer
Software as a medical device that performs clinical analysis is automatically classified as high risk under Article 6(1) of the AI Act, because it falls under the product safety legislation listed in Annex I and requires conformity assessment by a notified body. That classification triggers specific duties. The two most relevant to annotation are Article 10 on data and data governance, which expects training, validation, and test data to be relevant, representative, and examined for bias, and Article 15 on accuracy and robustness. Both feed the technical documentation described in Annex IV.
The timing, stated correctly
The widely repeated August 2026 deadline does not apply to medical image AI. That date belongs to the standalone systems listed in Annex III and to the separate AI literacy duty under Article 4. Medical devices follow the later Annex I clock. The European Commission's own medical device guidance confirms that Article 6(1) only takes effect for these systems on the later date, so an AI medical device is not classified as high risk under the Act before then.
Regulatory status note
The dates here reflect the EU AI Act as enacted and the Digital Omnibus on AI, for which EU institutions reached a provisional political agreement in May 2026, with formal adoption still pending at the time of writing. Treat the original AI Act dates as the binding planning baseline until the amendment is adopted, and check the current position before relying on a specific date. Last reviewed 29 June 2026.
What a defensible provenance record must contain
The following is the working schema we hold annotation programmes to. It is deliberately specific, because the difference between a record that survives scrutiny and one that does not is detail, not intent. Each element should be captured per image or per dataset version, not asserted at the level of the project as a whole.
| Provenance element | What the record must capture |
|---|---|
| Image source and rights | The origin of each image, the lawful basis for using it, the licensing terms, and any restriction on redistribution or on use for model training. |
| Patient privacy handling | The method used to remove or mask identifiers, who performed it, and how the removal was verified rather than assumed. |
| Guideline version | The exact version of the labelling instructions in force when each annotation was made, together with the full change history of those instructions. |
| Annotator identity and competence | A pseudonymous but resolvable identifier for each annotator, their relevant qualifications such as a board certified radiologist, and their onboarding and training records. |
| Task assignment and timing | Which annotator handled which image, when, and under which task definition. |
| Adjudication and agreement | How disagreements were resolved through consensus, expert adjudication, or majority vote, and the measured level of agreement between annotators. |
| Quality assurance sampling | The proportion of work checked a second time, who checked it, the error categories tracked, and the measured error rate over time. |
| Tooling and environment | The annotation platform and version used, including any model assisted labelling and the human review applied on top of it. |
| Immutable audit log | A tamper evident record of every create, edit, and delete action, with the actor, the timestamp, and the prior value. |
| Dataset versioning and lineage | A fixed version identifier for every dataset release, and the lineage linking the training, validation, and test splits back to source images. |
| Representativeness and bias slices | Coverage across patient demographics, scanner and device types, acquisition settings, and clinical sites, so that bias can be measured rather than presumed absent. |
Where annotation programmes get caught
The failures we see most often are not exotic. They are ordinary record keeping shortcuts that look harmless until an assessment asks a precise question.
- Guideline drift with no version history, so the team cannot say which rule produced which label.
- Annotator anonymity with no competence record, so there is no way to show qualified people did the work.
- Spreadsheet audit trails that any user can edit silently, which means the trail proves nothing.
- No agreement metrics, so quality is asserted in prose rather than measured in numbers.
- Demographic and device blind spots that surface only after deployment, when they are hardest to fix.
Questions and direct answers
Is medical image annotation covered by the EU AI Act?
When do the EU AI Act obligations for AI medical devices apply?
Does the August 2026 deadline apply to medical image AI?
What evidence proves annotation provenance?
Can provenance be reconstructed after a dataset is built?
How do the MDR duties and the AI Act duties differ here?
Next step
See where your provenance record is already defensible
If you are building AI on annotated medical images, the most useful first move is a short Compliance Review of your current annotation process against the schema in this guide, before a proof of concept on a sample of your data. That tells you where your record already holds and where it has gaps you would rather find now than during an assessment.
Request a Compliance ReviewSources
- Regulation (EU) 2024/1689, the EU AI Act: Articles 4, 6, 10, 15, and 113, and Annexes I, III, and IV. EUR-Lex.
- European Commission, Medical Device Coordination Group guidance on the interplay between the AI Act and the Medical Device Regulation and the IVDR.
- The Medical Device Regulation (MDR) and the In Vitro Diagnostic Regulation (IVDR), on data governance, clinical evidence, traceability, and surveillance after market launch.
- European Commission, Digital Omnibus on AI: proposal and provisional political agreement, 2026. Pending formal adoption at the time of writing.
Last reviewed 29 June 2026. Compliance dates are subject to the formal adoption of the Digital Omnibus and may change. This guide is general information, not legal advice.