Logo
FrontierNews.ai

Why AI Audits Are Missing the Real Problem: The Case for a 'Power Test'

Current AI audits focus on whether systems work consistently across groups, but they often skip a crucial question: Is the system predicting the right thing in the first place? A landmark 2019 study of a commercial algorithm used by US health systems illustrates the problem. The system was designed to identify patients who should receive additional care, but it actually predicted future healthcare costs rather than medical need. Because less money was historically spent on Black patients with comparable illnesses, Black patients assigned the same risk score were considerably sicker than white patients. Replacing cost with a closer measure of need would have increased the proportion of Black patients selected for additional support from 17.7 percent to 46.5 percent.

The algorithm was not simply inaccurate. It was competently predicting the wrong thing. This distinction matters enormously as audits become a central instrument of AI regulation worldwide. New York City requires bias audits for certain automated employment decision tools. The European Union's AI Act provides for conformity assessments and fundamental-rights impact assessments for high-risk systems. The US National Institute of Standards and Technology (NIST) framework encourages organizations to govern, map, measure, and manage AI risks.

What's Wrong With Today's AI Audits?

Most technical audits ask important but limited questions: Does a system perform consistently? Do error rates differ across groups? Are documented controls working? These questions are necessary, but they do not address whether the system's target is legitimate. A statistically balanced system can still support a punitive welfare policy, an exclusionary hiring process, or a surveillance practice that should not have been automated in the first place.

This gap creates what researchers call "bias laundering." The term does not imply that auditors intentionally conceal discrimination. Rather, it describes a narrower institutional risk: political choices about what should be predicted are translated into technical targets, assessed through compliance metrics, and returned to the public with the authority of an audit. In the healthcare example, treating expenditure as a proxy for medical need transformed a social inequality into the model's objective, which an audit could then certify as technically sound.

How Should AI Audits Actually Work?

Some governance frameworks already contain pieces of a better approach. The NIST AI Risk Management Framework asks organizations to document intended purposes, consult relevant external actors, and make a go-or-no-go decision about whether an AI system is appropriate. UNESCO's Ethical Impact Assessment similarly asks whether AI adoption is justified, identifies relevant stakeholders, and examines positive and negative consequences. The EU AI Act's fundamental-rights impact assessment requires covered deployers to describe the context of use, identify affected groups, assess risks, and establish mitigation and redress arrangements.

However, these frameworks remain fragmented. Some are voluntary. Some apply only to particular organizations or categories of systems. Others require documentation without clearly establishing who has authority to challenge the purpose of the system or the contractual arrangements behind it.

Experts argue that a comprehensive "power test" should accompany traditional fairness tests. This power test would examine four critical dimensions:

  • Objective Selection: Who decided what the AI system should predict, which alternatives were considered, why automation was chosen as the solution, and what would happen without it?
  • Data and Infrastructure Control: Who owns or controls the relevant data, model, computing infrastructure, and intellectual property? For public-sector systems, procurement contracts should guarantee access to documentation, change logs, incident reports, and independent testing.
  • Benefit and Burden Distribution: How are benefits, burdens, and savings distributed? A system may reduce administrative costs while transferring investigation, delay, documentation, or appeal costs to applicants, workers, patients, or welfare recipients. These effects should be treated as part of system performance rather than as externalities.
  • Practical Recourse: Can affected people obtain relevant information, challenge data and classifications, reach a responsible human decision-maker, and receive timely remedies? Affected communities should also have standing to influence system objectives before deployment, not merely report harms afterward.

The power test does not require a new regulatory agency or an additional audit industry. It can be incorporated into existing impact assessments, procurement reviews, conformity procedures, and public transparency records. The United Kingdom's Algorithmic Transparency Recording Standard already asks public bodies to publish information about system ownership, rationale, deployment context, data, risks, and accountability. The EU's fundamental-rights assessment template could make affected-group participation mandatory for consequential public uses and add questions about vendor dependence, benefit distribution, and the no-AI alternative.

What Practical Steps Can Regulators Take?

  • Standardized Templates: Develop proportionate requirements that reduce compliance costs, particularly for smaller organizations, while ensuring rigor for systems that materially shape employment, education, health, welfare, credit, migration, policing, or access to public services.
  • Auditor Independence: Establish minimum methodological standards, approve qualified auditors, require disclosure of financial relationships, and trigger reassessment after material system changes or serious incidents. Providers should not be able to define the benchmark, select the evidence, and determine what counts as a successful result without external scrutiny.
  • Confidential Annexes: Allow regulators to require a confidential annex where genuine security or trade-secret concerns prevent full publication, while still publishing a meaningful public summary that citizens can understand and scrutinize.

A power test is not the same as politicizing technical review. Auditors should not replace legislators, regulators, courts, workers, or affected communities in making policy decisions. Rather, auditors should verify that the institution deploying an AI system has genuinely considered whether automation serves the public interest, who benefits and who bears the costs, and whether affected people have meaningful routes to challenge or correct the system.

As AI systems increasingly shape consequential decisions in healthcare, employment, welfare, and criminal justice, the question of what these systems should predict becomes as important as whether they predict it accurately. The emerging audit regime needs to answer both.