Our previous blog post described the functions of assurance, meaning the distinct activities, processes, and interactions that together make up the assurance lifecycle. It traced how research, standards, policy, monitoring, and testing come together to assure AI and highlighted which functions are well-established and which are only emerging. In this post, we want to zoom into the operational heart of AI assurance: the Testing, Evaluation, Verification, and Validation (TEVV) process.
Testing: Gathering Evidence through Measurement
Every assurance process begins with measurement, the ground-level function through which AI providers and third-party assurers apply testing methodologies to an AI system, its components, or its underlying data, in order to generate evidence of how the system actually behaves.
The scope of what is measured and the methodologies used is broad and can encompass both qualitative and qualitative metrics and techniques. In practice, this means bias audits that compare outcomes across demographic groups, adversarial probes to assess system resilience, scenario exercises with users to understand human-AI interactions or formal verification using mathematical proof techniques.
The goal is not measurement for its own sake, but to produce evidence that can support meaningful judgments about an AI system's quality and the implications of its deployment. Before testing begins, two questions must be answered: What system behaviours and properties actually matter in the intended operational context? And Which testing procedure can accurately measure the intended concept? The answers to these questions determine what gets tested and what conclusions can be drawn about the system’s real-world behavior.
Evaluation: Verifying and Validating AI
Raw evidence from testing only becomes useful through evaluation. Evaluation describes the synthesis step where results are compared against requirements to determine whether a system does what it was built to do (verification) and whether it is fit for its intended purpose and context (validation). This is where technical findings meet operational, legal, and governance concerns to inform deployment readiness decisions.
The governing questions at this stage are rarely whether a system has passed a given test, but rather to what extent the results obtained are predictive of future behaviour, and what standard of adequacy should apply in this deployment, for this population of users. These are matters of judgment rather than calculation and require integrating risk appetite, contextual constraints, and the expectations set by regulation and organisational mandates in the final decision.
From Internal Testing to External Certification
An internal evaluation, however rigorously conducted, remains a self-reported claim. It is at this stage that the cycle deliberately extends beyond the organisation. Conformity assessment bodies, including independent auditors, notified bodies, and in some cases second-party assessors such as customers or investors conducting due diligence, examine the evidence produced against the standards established in Pillar 1. Where the system and its supporting documentation satisfy these requirements, certification follows, and that certification becomes the signal on which the wider market relies in order to extend trust without independently repeating the underlying work of testing.
This external check matters for reasons that go beyond compliance. An organisation assessing its own system faces an inherent conflict of interest: the same party that has a stake in a system’s approval cannot be relied upon to grade it with full impartiality. Even absent an explicit conflict of interest can commercial pressure to meet deployment timelines subtly shape how ambiguous results are interpreted. External reviewers are not only insulated from that pressure, they are also better positioned to identify blind spots, since they approach the system without the assumptions and habits of thought that accumulate over the course of building it.
External verification also carries a legitimacy that internal attestation cannot supply on its own, signalling to regulators, insurers, and the public that a claim has been examined by an independent party. A certification, however, is only as credible as the body issuing it, which is why accreditation exists as a further layer above conformity assessment. Accreditation is a meta-assurance function, typically administered by a single national body, that verifies the competence and impartiality of the assessors themselves. Absent accreditation, a claim of certification carries limited evidentiary weight.
Communicating What Was Found
Once results have been validated, they must be packaged, in formats such as model cards, system cards, technical files, and instructions for use, and directed to the parties who require them: senior managers weighing a deployment decision, end users who need to understand the scope of a system's appropriate use, downstream actors in the supply chain, and regulators who require documentation for enforcement purposes.
Where communication is conducted with appropriate rigour, it allows trust to travel along the AI value chain, enabling other actors in the assurance ecosystem to act on evidence they did not themselves generate. But as described previously, communication is the function most often treated as a residual reporting obligation. If test results aren’t communicated effectively to those that need them, either because they aren’t shared or they are shared in inaccessible formats and languages, the most rigorous test and evaluation efforts are useless.
The challenge of effective communication has grown significantly over recent years as AI use has spread beyond purely technical domains. The stakeholders who need to understand assurance evidence today include procurement officers, legal teams, affected communities, and policymakers, to name just a few. These groups have far more varied backgrounds, qualifications, and expertise than those who engaged with AI systems one or two decades ago. Ensuring TEVV-related documentation is available, accessible, understandable, and scrutable to these different stakeholders requires more resources dedicated to communication to make assurance meaningful beyond the organisation that conducted it.
Restarting the TEVV cycle
It would be a misreading of the map to treat the four processes described so far as a linear sequence that concludes upon deployment. The communication of findings may feed directly back into a further round of testing and evaluation, and post-deployment monitoring, encompassing the tracking of drift, the assurance that model updates have not eroded safety, and the identification of real-world harms, restarts the cycle. Findings from this ongoing monitoring may travel further upstream still, prompting new research questions, revised standards, or updated mandates within Pillar 1. Assurance is thus not a gate that a system passes through once, but a cycle it remains within for the duration of its lifecycle.
The Hidden Functions Shaping AI TEVV
Two further functions operate beneath the TEVV cycle, contributing structurally to all activities without applying directly to any particular system or test result. Professionalisation activities for assurers ensures that the people performing testing, evaluation, and audit are actually qualified to do so. The quality of every step in the TEVV cycle depends entirely on the competence of the people carrying it out. Stakeholder Engagement is equally essential. Complaints, incident reports, and engagements with affected users provide the market surveillance signals that allow regulators and deployers to identify harms, investigate failures, and enforce accountability. These inputs feed back into the wider ecosystem, shaping regulatory responses, informing new standards, and updating the risk mandates that set the scope for the next round of TEVV.
What AI Assurance Is
This post concludes a series that set out to answer a seemingly simple question: what is AI assurance, actually? Across six blog posts, we have examined the scope of AI assurance, the specific objects that must be assured, the actors that populate the assurance ecosystem, the interactions that bind them into a functioning network, the functions that support the assurance cycle, and finally, in this post, the operational heart of AI assurance: the AI testing, evaluation, verification, and validation process.
Taken together, these blog posts show that the simple question that started the series begets a complex answer. AI assurance is not one activity carried out by one actor on one thing. It is a distributed, continuous system in which scientific research, standardisation, testing, evaluation, certification, communication, monitoring, redress, and many more processes are mutually dependent, and in which the competence of practitioners and the voice of affected stakeholders are as consequential as the technical methods employed. Understanding assurance in these terms is a precondition for building the ecosystem well, and we hope this series, and the underlying assurance maps, have provided a useful primer for doing so.