Healthcare AI · Compliance · Annotation

Medical Image Annotation With Documented Provenance

The evidence record a notified body or an auditor expects when AI trained on annotated medical images reaches the market, and why that record cannot be assembled after the fact.

ISO 27001:2022 Audit log coverage Versioned guidelines Measured agreement

In short

The August 2026 deadline does not apply to medical device AI

  • Diagnostic imaging software that performs clinical analysis is high risk under Annex I of the EU AI Act, not Annex III. It follows a later clock.
  • The obligation date for these systems was 2 August 2027, and it is now moving to 2 August 2028 under the Digital Omnibus. The widely quoted August 2026 date is for a different category.
  • The duty that bites is data governance: proving how your training data was built. For medical images, that means provenance.
  • Provenance is the one artefact you cannot backfill. If it was not captured while the labelling happened, it cannot be written in later.

What documented provenance means in medical image annotation

Documented provenance in medical image annotation is the complete, verifiable record of how every label on every image was produced. It captures where each image came from and under what lawful basis, which version of the labelling instructions applied, who annotated the image and with what qualifications, how disagreements between annotators were resolved, and how the result was checked. In short, it is the evidence that lets an auditor or a notified body reconstruct, after the fact, exactly how a training dataset was built and whether the labels can be trusted.

For AI that is itself a medical device, or that sits inside one, this record is not optional polish. It is the substrate of the technical documentation that the Medical Device Regulation and the EU AI Act both expect, and it is the part a vendor cannot produce later if it was not captured while the work was happening.

Provenance is the one artefact you cannot backfill

Most compliance gaps can be closed late. A model can be retrained. A data governance policy can be written this quarter to cover a process that began last year. A risk assessment can be drafted in a week. Provenance is the exception. You cannot reconstruct who annotated a particular scan in 2026, under which version of the guideline, with what adjudication, and at what measured level of agreement, unless you recorded it at the time.

This is why the lead time matters more than the headline date. A team that starts capturing a defensible provenance record now will have years of clean evidence when an assessment arrives. A team that waits will have a dataset it cannot fully account for, and retrofitting a credible history into an annotation programme is somewhere between expensive and impossible. The record has to be built as the labels are made, not described afterwards.

Where the obligations actually sit, and when

Two regulatory frameworks apply to AI driven medical imaging software, and they run on different clocks. Conflating them is the most common error in the market, so it is worth separating them cleanly.

Aug 2026 Dec 2027 Aug 2028 STANDALONE SYSTEMS · ANNEX III was 2026 now Dec 2027 MEDICAL DEVICE AI · ANNEX I was Aug 2027 now Aug 2028 the date everyone quotes
Figure 1. The August 2026 date belongs to standalone Annex III systems and the AI literacy duty. Medical device AI sits on the later Annex I clock.

The Medical Device Regulation and the IVDR apply now

If your imaging software performs a clinical function, it is a medical device. The Medical Device Regulation and the In Vitro Diagnostic Regulation already require data governance, clinical evidence, traceability, and surveillance once a device is on the market. These duties are in force today, independent of the AI Act, and they already reach the quality and documentation of the data used to develop the device.

The EU AI Act adds a second, parallel layer

Software as a medical device that performs clinical analysis is automatically classified as high risk under Article 6(1) of the AI Act, because it falls under the product safety legislation listed in Annex I and requires conformity assessment by a notified body. That classification triggers specific duties. The two most relevant to annotation are Article 10 on data and data governance, which expects training, validation, and test data to be relevant, representative, and examined for bias, and Article 15 on accuracy and robustness. Both feed the technical documentation described in Annex IV.

The timing, stated correctly

The widely repeated August 2026 deadline does not apply to medical image AI. That date belongs to the standalone systems listed in Annex III and to the separate AI literacy duty under Article 4. Medical devices follow the later Annex I clock. The European Commission's own medical device guidance confirms that Article 6(1) only takes effect for these systems on the later date, so an AI medical device is not classified as high risk under the Act before then.

Regulatory status note

The dates here reflect the EU AI Act as enacted and the Digital Omnibus on AI, for which EU institutions reached a provisional political agreement in May 2026, with formal adoption still pending at the time of writing. Treat the original AI Act dates as the binding planning baseline until the amendment is adopted, and check the current position before relying on a specific date. Last reviewed 29 June 2026.

What a defensible provenance record must contain

The following is the working schema we hold annotation programmes to. It is deliberately specific, because the difference between a record that survives scrutiny and one that does not is detail, not intent. Each element should be captured per image or per dataset version, not asserted at the level of the project as a whole.

Representativeness and bias slices Dataset versioning and lineage Immutable audit log Adjudication, agreement and quality checks Annotator identity, competence and guideline version Image source, rights and privacy handling captured at the time
Figure 2. A provenance record is built in layers, every one captured as the work happens, never reconstructed afterwards.
Provenance elementWhat the record must capture
Image source and rightsThe origin of each image, the lawful basis for using it, the licensing terms, and any restriction on redistribution or on use for model training.
Patient privacy handlingThe method used to remove or mask identifiers, who performed it, and how the removal was verified rather than assumed.
Guideline versionThe exact version of the labelling instructions in force when each annotation was made, together with the full change history of those instructions.
Annotator identity and competenceA pseudonymous but resolvable identifier for each annotator, their relevant qualifications such as a board certified radiologist, and their onboarding and training records.
Task assignment and timingWhich annotator handled which image, when, and under which task definition.
Adjudication and agreementHow disagreements were resolved through consensus, expert adjudication, or majority vote, and the measured level of agreement between annotators.
Quality assurance samplingThe proportion of work checked a second time, who checked it, the error categories tracked, and the measured error rate over time.
Tooling and environmentThe annotation platform and version used, including any model assisted labelling and the human review applied on top of it.
Immutable audit logA tamper evident record of every create, edit, and delete action, with the actor, the timestamp, and the prior value.
Dataset versioning and lineageA fixed version identifier for every dataset release, and the lineage linking the training, validation, and test splits back to source images.
Representativeness and bias slicesCoverage across patient demographics, scanner and device types, acquisition settings, and clinical sites, so that bias can be measured rather than presumed absent.

Where annotation programmes get caught

The failures we see most often are not exotic. They are ordinary record keeping shortcuts that look harmless until an assessment asks a precise question.

  • Guideline drift with no version history, so the team cannot say which rule produced which label.
  • Annotator anonymity with no competence record, so there is no way to show qualified people did the work.
  • Spreadsheet audit trails that any user can edit silently, which means the trail proves nothing.
  • No agreement metrics, so quality is asserted in prose rather than measured in numbers.
  • Demographic and device blind spots that surface only after deployment, when they are hardest to fix.

Questions and direct answers

Is medical image annotation covered by the EU AI Act?
Yes, when the annotated images train AI that is itself a medical device or sits inside one. Such software is automatically high risk under Article 6(1), which brings the data governance and accuracy duties that the annotation record has to support.
When do the EU AI Act obligations for AI medical devices apply?
The original date for these Annex I systems was 2 August 2027. Under the Digital Omnibus, for which EU institutions reached a provisional agreement in May 2026, it moves to 2 August 2028. Until that amendment is formally adopted, treat the original date as the planning baseline.
Does the August 2026 deadline apply to medical image AI?
No. That date covers the standalone systems in Annex III and the separate AI literacy duty under Article 4. Software as a medical device follows the later Annex I clock, not the August 2026 date.
What evidence proves annotation provenance?
The schema above: image source and rights, privacy handling, guideline version, annotator competence, task and timing, adjudication and agreement, quality sampling, tooling, an immutable audit log, dataset lineage, and representativeness slices, each captured per image or per dataset version.
Can provenance be reconstructed after a dataset is built?
Not credibly. You cannot recover who annotated an image, under which guideline, with what adjudication, unless it was recorded at the time. This is the single strongest reason to start capturing the record before a deadline rather than after one.
How do the MDR duties and the AI Act duties differ here?
The Medical Device Regulation and the IVDR are in force now and already govern data quality and traceability. The AI Act adds a parallel layer, chiefly Article 10 and Article 15, on the later Annex I clock. A medical device AI has to satisfy both, which is why a single, well structured provenance record that serves both frameworks is the efficient path.

Next step

See where your provenance record is already defensible

If you are building AI on annotated medical images, the most useful first move is a short Compliance Review of your current annotation process against the schema in this guide, before a proof of concept on a sample of your data. That tells you where your record already holds and where it has gaps you would rather find now than during an assessment.

Request a Compliance Review

Sources

  1. Regulation (EU) 2024/1689, the EU AI Act: Articles 4, 6, 10, 15, and 113, and Annexes I, III, and IV. EUR-Lex.
  2. European Commission, Medical Device Coordination Group guidance on the interplay between the AI Act and the Medical Device Regulation and the IVDR.
  3. The Medical Device Regulation (MDR) and the In Vitro Diagnostic Regulation (IVDR), on data governance, clinical evidence, traceability, and surveillance after market launch.
  4. European Commission, Digital Omnibus on AI: proposal and provisional political agreement, 2026. Pending formal adoption at the time of writing.

Last reviewed 29 June 2026. Compliance dates are subject to the formal adoption of the Digital Omnibus and may change. This guide is general information, not legal advice.