Skip to main content
v2026.11,858 entries · CC-BY 4.0

Third-Party AI Auditing: What It Is and Who Does It

Third-party AI auditing means an independent party assessing an AI system, or the organization running it, against a compliance framework like ISO/IEC 42001 or the EU AI Act — not the same as using AI tools to automate financial audits.

Written and maintained by CASRAI Editorial Board

Last updated

Third-party AI auditing is the practice of having an organization with no role in building or operating a given AI system assess that system, or the practices of the organization that built or deployed it, against a defined external framework — a regulation, an industry standard, or a published internal policy. It is a governance and compliance discipline, not a performance benchmark.

It is easy to confuse with a different, unrelated practice area that happens to share the same search terms: "AI in auditing," meaning the use of AI and machine-learning tools inside financial or IT audit engagements to automate parts of the audit itself, such as transaction testing or anomaly detection in ledger data. That is audit automation — using AI as a tool inside a conventional financial audit. This page is about the opposite direction: auditing AI systems and the organizations responsible for them, using human auditors and structured assessment processes.

What a Third-Party AI Audit Actually Assesses

The subject of a third-party AI audit is conformance, not capability. An audit does not generally ask "how accurate is this model" or "how well does it perform on a benchmark." It asks whether the system, and the organization operating it, satisfy the specific requirements set out in whichever framework the audit is run against: has a risk-management process actually been followed and documented; does the required technical documentation exist and match what was deployed; were the testing, human-oversight, and record-keeping obligations the framework specifies actually met.

That distinction matters because it is possible for a system to perform well on its own terms and still fail an audit — if the organization can’t produce the documentation, testing records, or oversight evidence the framework requires. Conversely, an audit does not certify that a system is safe or accurate in any absolute sense; it certifies that a defined set of process and documentation requirements were met.

Third-Party Audit vs. Internal Review

An internal review is carried out by people inside the organization that built or deployed the system — an internal audit function, a compliance team, an ML governance group. Internal review is useful for continuous self-monitoring and catching problems early, and most organizations that eventually undergo a third-party audit have already run several rounds of internal review to prepare for it.

What internal review cannot provide is independence. Its findings are only as credible, to an outside party, as the reviewer’s independence from the thing being reviewed — and an internal team, by definition, has none. A third-party audit is performed by an entity with no stake in the outcome, and its output (a certificate, an audit report, a conformity assessment) is built to be relied on by someone who is not in a position to simply take the audited organization’s word for it: a regulator, a customer, an insurer, a business partner running vendor due diligence.

The Frameworks a Third-Party AI Audit Can Be Run Against

There is no single "AI audit" standard. In practice, third-party AI auditing happens against one of a small number of concrete frameworks, and which one applies changes who performs the audit and what it produces.

ISO/IEC 42001 is an international standard specifying requirements for an AI management system (AIMS) — the same structural approach ISO/IEC 27001 takes for information security management. An organization can be certified against ISO/IEC 42001 by an accredited certification body; this is the closest AI-specific equivalent to a conventional management-system certification audit.

The EU AI Act’s conformity assessment regime (Article 43) applies to high-risk AI systems as defined in the Act. For the largest category of high-risk systems (Annex III, point 1), a provider can choose between an internal-control-based assessment or an assessment that involves a notified body, depending on whether harmonised standards or common specifications were fully applied; where no relevant standard exists yet, third-party involvement is required. For the remaining high-risk categories (Annex III, points 2–8), the Act currently specifies internal control only, with no notified-body involvement, unless the European Commission later extends third-party assessment by delegated act. AI systems used for law enforcement, migration, asylum, or border-control purposes follow a different path again: assessment by the relevant market surveillance authority rather than a private notified body.

The NIST AI Risk Management Framework (AI RMF) is different in kind from the first two: it is explicitly voluntary, and NIST itself does not certify or audit organizations against it. What exists instead is a market of third-party assessments — audit and advisory firms that will evaluate an organization’s AI governance against the RMF’s structure, often as one component of a broader engagement that also covers standards like SOC 2 or ISO/IEC 27001. An organization can point to a favorable AI RMF-aligned assessment as evidence of mature practice, but there is no NIST certificate behind it.

Who Performs Third-Party AI Audits

For ISO/IEC 42001, the audit is carried out by an accredited certification body — a firm that has itself been accredited to issue ISO/IEC 42001 certificates by a national accreditation body such as UKAS (UK), ANAB (US), or RvA (Netherlands). Several established certification and assurance firms (BSI, Schellman, DNV, KPMG, and NQA among them) now offer accredited ISO/IEC 42001 certification alongside their existing ISO 27001 and SOC 2 practices.

For the EU AI Act’s notified-body track, the auditor is a conformity assessment body that has been formally designated as a notified body by an EU member state’s national notifying authority. A notified body audits both the provider’s quality management system and its technical documentation, and issues the certificate that underpins the system’s CE marking.

A separate group is worth naming precisely because it gets confused with the two above: the frontier-model evaluator ecosystem. Organizations like METR, an independent research nonprofit that evaluates frontier models’ dangerous capabilities and has run frontier-risk assessments in partnership with labs including Anthropic, OpenAI, Google DeepMind, Meta, and Amazon; the UK’s AI Security Institute (renamed from the AI Safety Institute in February 2025); and the US Center for AI Standards and Innovation (CAISI, renamed from the US AI Safety Institute in June 2025, and operating within NIST) all evaluate frontier models directly — dangerous-capability testing, red-teaming, pre-deployment evaluation. That is a genuinely different activity from a compliance audit: it assesses what a model can do, not whether an organization’s paperwork and process satisfy a named framework. The two ecosystems increasingly overlap in practice, but they are not the same audience performing the same task.

What a Third-Party AI Audit Produces

The output depends on the framework. An ISO/IEC 42001 audit that succeeds produces a certificate, in the same pattern as other ISO management-system certifications, with the certification body continuing to run periodic surveillance audits to keep it valid. An EU AI Act conformity assessment produces an EU declaration of conformity and CE marking, and — where a notified body was involved — a certificate from that body plus an audited technical documentation file that regulators can request. A NIST AI RMF-based assessment produces no certificate at all, since NIST doesn’t issue one; it typically produces a written report benchmarking the organization’s governance practices against the framework’s structure, delivered by whichever firm ran the engagement.

None of these outputs are a statement that the underlying AI system is safe in some general sense. They are a statement that a defined, external party checked a defined set of requirements and found them met, on the date of the audit.

Where This Fits in the Frontier AI Safety Landscape

Third-party AI auditing sits inside a wider vocabulary problem: "audit," "assessment," "evaluation," and "conformity assessment" get used by different frameworks, labs, and regulators to mean related but distinct things, often without cross-reference to how another framework uses the same word. CASRAI’s NIKOLAI is an open, versioned dictionary of elements for frontier-AI safety documentation that crosswalks exactly this kind of terminology — including auditing and evaluation vocabulary — against primary-source frameworks from labs, standards bodies, and regulators. NIKOLAI itself is not an auditor, an evaluator, or a certification body; it does not assess any system or issue any designation. What it does is let someone reading an ISO/IEC 42001 certificate, an EU AI Act technical documentation file, and a METR evaluation report line up the terms each one uses against a common reference.

Frequently Asked Questions

Is third-party AI auditing the same thing as "AI auditing" in accounting or finance?

No. "AI auditing" in an accounting context usually means using AI tools inside a financial or IT audit to automate tasks like transaction testing. Third-party AI auditing, as covered on this page, means independently auditing an AI system or organization’s compliance with a governance framework. The two share a search term and nothing else.

Does NIST certify AI systems under the AI RMF?

No. The NIST AI RMF is explicitly voluntary guidance, and NIST does not operate a certification or audit program for it. Third-party firms assess organizations against the framework’s structure, but there is no NIST certificate to obtain.

Is third-party assessment mandatory under the EU AI Act?

Only for a subset of high-risk AI systems. Under Article 43, the largest high-risk category can require notified-body involvement when no harmonised standard or common specification has been fully applied; most other high-risk categories currently use internal control only, with no notified body involved.

What’s the difference between a third-party AI auditor and a frontier-model evaluator like METR?

An auditor checks whether an organization’s process and documentation satisfy a named compliance framework. An evaluator like METR, the UK AI Security Institute, or the US Center for AI Standards and Innovation tests what a model can actually do — capability and red-team testing performed directly against the system. They’re complementary, but they answer different questions for different audiences.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Third-Party AI Auditing: What It Is and Who Does It

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →