If you built your compliance plan around 2 August 2026, that date just moved. On 7 May 2026, the Council of the EU and the European Parliament agreed — provisionally — on the Digital Omnibus on AI. The headline: the rules for high-risk AI systems are pushed back by more than a year.

It's tempting to read that as breathing room and move on. We'd push back on that — not to invent urgency, but because of what got delayed and why. The delay is real and welcome. For most teams, the right response is still to keep going. Here's what changed, what Article 12 still demands, and why the training-data layer is the one part of this you want to start early.

What the Omnibus actually changed

The new agreement delays high-risk obligations on two tracks.

Stand-alone high-risk systems — the Annex III list covering recruitment, credit scoring, education, biometrics, law enforcement, and essential services — now apply from 2 December 2027, about sixteen months later than the original 2 August 2026 date. High-risk AI built into regulated products (medical devices, machinery, vehicles) moves to 2 August 2028.

One caveat, stated plainly: this is a provisional agreement. It still has to be formally adopted and published in the Official Journal, expected in the coming weeks. The major law firms tracking it, and the EU's own institutions, are already planning around the new dates — so we are too — but the text isn't final yet.

And here's the part that matters most: only the deadline moved. The risk tiers, the assessments, the obligations themselves — all intact. When 2 December 2027 arrives, Article 12 will read exactly as it does today.

Diagram of Article 12 requirements: the two load-bearing words automatic and lifetime, the three logging purposes, and the minimum six-month retention floor.
Figure 2. What Article 12 still requires. Automatic means the system records events itself; lifetime means from deployment to decommissioning. The logs serve three purposes, kept for at least six months.

What Article 12 still requires

Article 12 catches teams off guard because it isn't about how your model behaves. It's about whether you can prove what it did.

It requires high-risk AI systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system." Two words carry the weight:

  • Automatic — the system creates the records itself. A document you write afterwards, describing what you think happened, doesn't count.
  • Lifetime — from the day you deploy the system to the day you retire it. Not from the day your compliance programme finally switched on.

The logs exist for three jobs: spotting when a system might pose a risk or has been changed significantly, supporting monitoring after launch, and helping deployers keep watch day to day. Related rules say you keep these logs for at least six months, and longer if the system needs it.

Does this apply to you? For most organisations running AI in real, decision-affecting situations that fall into a high-risk category, yes. And getting it wrong is expensive: penalties for high-risk non-compliance reach €15 million or 3% of global annual turnover, whichever is higher.

A quick note on who's responsible. Article 12 lands mainly on the provider — the company that builds the system. Deployers have their own separate duties. But plenty of enterprises both build and deploy, and making significant changes to a system can pull provider duties onto you. So "we just deploy it" is rarely a clean way out.

The gap most teams miss: a log proves what happened in production, not where your data came from

This is the under-discussed part, and the reason a delay isn't a reason to wait.

Article 12 is written about the system running — the decisions it makes, the events it logs in real time. Most tools sold around it focus there. That work matters, but it isn't the whole exposure, and it isn't the part that takes longest to get right. The real lead time sits in the data layer.

When a regulator, auditor, or customer's procurement team digs into a high-risk model, the questions go back to the training data. How was this label made? Who labelled it? Who checked it? Where did people disagree, and how was that settled?

That's a data provenance question — provenance meaning the documented history of where your data came from and how it was handled. And provenance has one inconvenient feature: you either captured it while you were labelling, or you didn't. You can't reconstruct a believable audit trail for training data after the fact. A log you start keeping in 2027 tells a regulator nothing about a dataset you labelled in 2026. That's the quiet weakness in a lot of annotation pipelines — they were built to produce labels, not evidence about those labels.

That made sense when buyers asked about speed and cost per label. But the question has shifted: not "who labels fastest," but "who can prove the label holds up." Hardly anyone in the annotation market positions around compliance — which, if you're the one buying, is exactly where the advantage sits. The teams that treat these extra eighteen months as preparation, not a reprieve, will have a defensible evidence base when the date lands.

Side-by-side comparison: runtime logging covers the deployed system and satisfies Article 12, while data provenance must be captured at labelling time or it is lost.
Figure 3. A runtime log proves what a deployed system did — not how its training data was built. Provenance is captured at labelling time or not at all; a 2027 log says nothing about a dataset labelled in 2026.

What a defensible annotation stack has to show

Trace Article 12's logic back to the data, and a short checklist falls out. A pipeline that can stand up to scrutiny needs to show how its logs connect to where the data came from:

  • A tamper-evident log of every labelling action — a record of who did what and when, written as it happened and impossible to quietly edit later. This is the line between a record an auditor accepts and one they question.
  • A clear chain of custody for the dataset — the full path from raw data to finished label, exportable in a form auditors recognise.
  • Separation of duties — labeller, reviewer, and auditor as distinct, enforced roles, so quality sign-off isn't someone marking their own homework.
  • A real quality score — a measured number, plus a record of how disagreements got resolved.
  • Retention you can rely on — logs kept long enough for the legal minimum and your model's full life.

None of this is exotic. But most teams only notice the gap when someone asks for the evidence and it isn't there — and by then the data is already months old.

Where LabelFort fits

This is the problem LabelFort was built for, not bolted onto afterwards. It's Predusk AI's data annotation platform for teams whose training data has to survive outside review, and its design follows Article 12's logic by tying every log back to the data's history.

Every label, every reviewer action, every configuration change goes into a record that can't be quietly altered, and the full chain of custody for any project exports in formats auditors already accept — so when the question comes, the answer is a download, not a fire drill. Quality isn't claimed; it's measured. We score agreement between annotators using two standard statistical measures (Cohen's Kappa and Krippendorff's Alpha) against a threshold above 90%, checked on your own data. Roles are kept separate through distinct workspaces. And while a Visual Language Model handles a first pass to speed things up, no auto-generated label reaches your dataset without a human checking it — so you get the speed without losing the human accountability a regulator expects. Audit-log coverage is 100% by design.

To be precise, because overclaiming here would defeat the point: LabelFort governs the training-data layer — the evidence for how your datasets were built. It works alongside, not instead of, the runtime logging your live system needs for its own Article 12 duties. Its job is to make the data half defensible from the very first label.

Compliance posture

Five frameworks. One annotation backbone that passes Legal, Security, and Procurement.

ISO 27001:2022
CERTIFIED
SOC 2
CERTIFIED
HIPAA
COMPLIANT
GDPR
COMPLIANT
DPDP
READY

The honest bottom line

The deadline moved. The work didn't. Article 12 doesn't ask whether you mean to keep records; it asks whether your systems produce them automatically, across their whole life — and whether that same discipline reaches back to the data underneath. The sixteen-month delay is a genuine gift of time. The only question is whether you spend it building the evidence base or putting it off. For training data, putting it off is the expensive choice, because you can't backfill provenance.

If you want to know whether your annotation evidence would hold up under that scrutiny, that's a conversation worth having now — while there's still time to fix whatever it surfaces.

Next step

Evaluate LabelFort against your regulator's checklist.

Every Predusk engagement starts with a Compliance Review: we map your risk surface to LabelFort's controls and scope an evidence-grade proof of concept — on your data, under your constraints, before any data moves.