Blog/Quality Assurance

What Happens During an EU AI Act Audit: A Step-By-Step Breakdown

Businesswoman using a laptop at a desk

Summarize with:

Once a company decides it needs an EU AI Act compliance audit, the next question is usually the same: what does that actually involve? Not the regulatory theory, the practical sequence. What gets reviewed, in what order, and what you walk away with at the end. If you're still working out whether your organization needs one in the first place, this breakdown of the signals that mean you're in scope is worth reading first.

A technical audit runs in five stages, typically over eight to twelve weeks. The scope follows the obligations set out in Articles 9 to 17 and Annex IV of the Act. Here's what happens at each stage.

TL;DR

30-second summary

What actually happens during an EU AI Act compliance audit, stage by stage?

  • Stage 1 maps every AI system, and it's rarely the list a company expects. Independent research consistently finds enterprises discover two to three times more AI systems than anticipated — recommendation engines embedded in e-commerce platforms, resume-screening features bundled into HR tools, chatbot plugins added two product cycles ago. This stage requires input from engineering, procurement, marketing, and ops, since no single team has the full picture.
  • Stage 2 classifies each system's risk tier and role, and role is decided per system, not once for the company. A company that sees itself purely as a deployer of third-party AI tools can discover that heavy customization of one integration actually makes it a provider for that specific system, with correspondingly heavier obligations. High-risk classification takes the most audit time and is the tier most often misjudged.
  • Stage 3 checks every in-scope system against its actual obligations, not just written policy. This means verifying a documented risk management process exists rather than a policy stating one does, data governance records showing what data trained the system, and a human oversight design specifying how someone can actually override outputs. Third-party and open-source model gaps transfer directly to the organization building on top of them.
  • Stage 4 tests whether evidence would survive a market surveillance request, which is different from whether it exists. A common finding is a gap between a compliance document written after the fact and the system's actual operational reality, a risk management policy nobody on the engineering team has ever followed doesn't hold up under scrutiny, and it's exactly the kind of gap regulators are trained to find.
  • Stage 5 produces a prioritized remediation roadmap mapped to statutory dates, not a generic to-do list. Since the Digital Omnibus deferred high-risk conformity assessment deadlines to December 2027 and August 2028, most organizations have real runway to work through findings properly rather than racing a deadline — worth using deliberately, since conformity assessment itself typically takes six to twelve months.

Bottom line: The audit sequence matters as much as its content. Guessing at classification before the inventory is complete, or skipping to remediation without testing whether evidence would actually hold up, produces a document that looks thorough and misses what matters. The deliverable is a board-readable audit report plus a working compliance register that doubles as a standing answer pack for customer AI questionnaires.

Stage 1: AI system inventory

Every audit starts by mapping every AI system your organization builds, buys, or has embedded in third-party tools. This is rarely the same list as what's formally documented. Independent research consistently finds that enterprises discover two to three times more AI systems than they expected going in, things like a recommendation engine embedded in an e-commerce platform, a resume-screening feature bundled into an HR tool, or a chatbot plugin added to a support widget two product cycles ago.

This stage alone often takes longer than people expect, because it requires input from more than one team. Engineering knows what's in the codebase. Procurement knows what's under contract. Marketing and ops often know about AI features nobody else flagged, because they're the ones using them day to day.

Stage 2: Risk classification

Once the inventory exists, each system gets classified under the Act's four tiers (prohibited, high-risk, limited, minimal), and your role per system, provider or deployer, is defined, since obligations depend on both (see our breakdown of the four tiers and the deadline timeline if you need the fuller picture).

Role isn't decided once for the whole company. It's decided per system, and it's where audits often turn up surprises. A company that thinks of itself purely as a deployer of third-party AI tools can discover that one integration, because of how heavily it was customized, actually makes them a provider for that specific system, with a correspondingly heavier set of obligations. High-risk classification is where the audit spends the most time, since it covers systems like hiring tools, credit scoring, and biometric identification, and it's also the tier most often misjudged. A tool that looks like ordinary software, a CV-ranking feature, an automated underwriting model, can fall squarely into this tier once its actual function is examined closely.

Not sure whether customizing a third-party AI model has made you its provider?

We help organizations classify each AI system by role and risk tier, catching the misclassifications that carry the heaviest obligations.

Stage 3: Gap assessment

Man taking off his glasses while working on a laptop

Every in-scope system is checked against its actual obligations: risk management, data governance, technical documentation, logging, human oversight, transparency, robustness, and post-market monitoring.

For high-risk systems especially, this means checking for a documented risk management process, not just a policy stating one exists; data governance records showing what data trained or feeds the system and how quality and bias were assessed; technical documentation describing the system's design and intended purpose; and a human oversight design that specifies how a person can intervene in or override the system's outputs. Third-party and open-source models get checked here too. A vendor's compliance gap becomes your compliance gap the moment you build on top of their model, since your organization typically inherits obligations tied to that model's use even if you didn't train it.

Stage 4: Evidence review

Regulators do not accept intentions. This stage tests whether your documentation, logs, and oversight records would actually survive a market surveillance request, which is a different question from whether they exist.

A common finding here is a gap between a compliance-oriented document written after the fact and the system's actual operational reality. A risk management policy that describes a process nobody on the engineering team has ever followed doesn't hold up under scrutiny, and it's exactly the kind of gap regulators are trained to look for. This is also where logging gets tested directly: does the system actually produce the logs its documentation claims, are they retained for the required period, and if a human reviewer is supposed to be able to override a decision, can they actually do so in practice.

Stage 5: Remediation roadmap

The audit ends with a prioritized plan mapped to the statutory dates: what to fix first, what can wait, and what it will cost. Not every finding carries the same urgency, and a useful roadmap says so explicitly, separating systems missing a control right now from systems with a documentation gap that's real but lower-stakes.

Since the Digital Omnibus deferred the high-risk conformity assessment deadlines to December 2027 and August 2028, most organizations now have real runway to work through this list properly rather than racing a deadline. That's worth using deliberately, since conformity assessment itself typically takes six to twelve months once a system is ready for it, and gaps found during a rushed pre-deadline scramble cost far more to fix than the same gaps found with a year or two of lead time.

What you walk away with

The deliverable is an audit report your board can read, plus a working compliance register your team owns afterward. The same material doubles as a standing answer pack for customer AI questionnaires: filled in once, reused in every procurement review that asks about your AI systems' compliance status.

Where to start

None of this needs to happen all at once, but it does need to happen in order. Guessing at classification before the inventory is complete, or skipping straight to a remediation plan without testing whether evidence would actually hold up, produces a document that looks thorough and misses the things that matter.

TestDevLab's EU AI Act compliance audit follows the five-stage process described above, run over eight to twelve weeks. It's built to tell you not just what the Act requires, but whether your systems, and your evidence, can prove you meet it.

FAQ

Most common questions

How long does an EU AI Act compliance audit take?

A technical audit typically runs eight to twelve weeks across five stages: system inventory, risk classification, gap assessment, evidence review, and remediation roadmap. The inventory stage alone often takes longer than expected because it requires input from multiple teams. Engineering knows what's in the codebase, procurement knows what's under contract, and marketing or operations teams often know about AI features nobody else flagged. The scope follows the obligations set out in Articles 9 to 17 and Annex IV of the Act.

Why do companies discover more AI systems than they expected during an audit?

Independent research consistently finds enterprises discover two to three times more AI systems than anticipated going into an inventory. These are typically embedded or bundled features that were never formally documented as AI. For example, a recommendation engine inside an e-commerce platform, a resume-screening feature bundled into an HR tool, or a chatbot plugin added to a support widget. No single team holds the complete picture, which is why the inventory stage requires cross-functional input rather than a single engineering audit.

How is risk classification determined for each AI system?

Each system is classified under the Act's four tiers — prohibited, high-risk, limited, or minimal — and the organization's role for that specific system, provider or deployer, is determined separately. Role is decided per system, not once for the whole company, and audits frequently uncover cases where heavy customization of a third-party integration has converted a deployer into a provider for that specific system. High-risk classification takes the most audit time and is the tier most often misjudged, since tools that look like ordinary software, such as CV-ranking features or automated underwriting models, can fall squarely into it once their actual function is examined.

What does the evidence review stage of an EU AI Act audit actually test?

Evidence review tests whether documentation, logs, and oversight records would survive a market surveillance request, which is a different question from whether they exist on paper. A common finding is a gap between a compliance document written after the fact and the system's actual operational reality. A risk management policy describing a process the engineering team has never followed does not hold up under regulatory scrutiny. This stage also tests logging directly, confirming systems produce the logs their documentation claims, retain them for the required period, and that human oversight mechanisms actually function as designed rather than existing only in policy.

What does an organization receive at the end of an EU AI Act audit?

The deliverable is an audit report readable by a board, alongside a working compliance register the internal team owns going forward. This includes a prioritized remediation roadmap mapped to statutory deadlines, distinguishing systems missing a control right now from systems with a lower-stakes documentation gap. The same material typically doubles as a standing answer pack for customer AI questionnaires, since it can be reused across procurement reviews that ask about AI system compliance status rather than being rebuilt for each request.

A policy that looks compliant isn't the same as one that is

TestDevLab's EU AI Act compliance audit follows the full five-stage process over eight to twelve weeks, built to tell you not just what the Act requires, but whether your systems and evidence can prove you meet it.

Summarize with:

QA engineer having a video call with 5-start rating graphic displayed above

Save your team from late-night firefighting

Stop scrambling for fixes. Prevent unexpected bugs and keep your releases smooth with our comprehensive QA services.

Explore our services