Logo
FrontierNews.ai

New Zealand Police Restart Whisper AI After Officers Broke Rules, Raising Questions About Accuracy in Law Enforcement

New Zealand police have restarted use of OpenAI's Whisper speech-to-text model for criminal investigations after a troubled trial period, but the tool's significant accuracy problems on Indigenous and minority languages are raising fresh concerns about whether AI should be trusted in high-stakes law enforcement work. The police high-tech crime group has now been using the approved Whisper AI for nine months following a restart in September 2025, according to Official Information Act documents released to RNZ.

What Went Wrong During the Initial Trial?

The first Whisper trial ran for six months through March 2025, but officers significantly misused the system by breaking explicit restrictions. Police had banned the use of Whisper on Māori and Pacific language audio because the model is 45 percent inaccurate on those languages. Despite this clear prohibition, officers processed English-language audio during the trial period, violating the approved use case. An internal police document captured the problem bluntly: "Trial misuse demonstrated limitations of relying on behavioural controls alone".

The misuse was serious enough that police shut down the entire project in March 2025. However, the shutdown created an unexpected problem. Police found that officers were turning to unapproved AI models instead, which prompted leadership to reconsider the ban. In September 2025, police restarted Whisper with additional safeguards in place.

How Accurate Is Whisper Really?

The accuracy concerns extend beyond language-specific failures. A Cornell University study published in May 2024 found that Whisper hallucinated approximately one percent of the time, meaning it fabricated content that was never spoken in the original audio. The study documented cases where Whisper invented racial commentary, violent rhetoric, and even imagined medical treatments that did not exist in the source material.

OpenAI acknowledged the ongoing challenge. "Addressing hallucinations is an ongoing area of research," the company told RNZ. A company spokesperson added: "Speech recognition systems are not perfect and should be evaluated carefully for their intended use case". OpenAI has released newer variants, including GPT-Realtime-Whisper, which the company claims reduce hallucinations and improve accuracy across various languages.

What Safeguards Are Now in Place?

Police implemented multiple layers of control after restarting the program. Only trained and approved staff can now access Whisper, and all users must complete mandatory AI awareness training. Critically, police stated that Whisper output is not used as evidence in court proceedings. Instead, transcripts generated by the tool are treated as preliminary investigative material that requires verification by a qualified person before any evidentiary use.

Detective Superintendent Keith Borrell, director of the national criminal investigations group, explained the current approach: "Text is added to the outputs to make it clear that transcripts must not be relied upon for critical decision making or evidential purposes without verification by an appropriately qualified person".

Steps to Implement AI Safeguards in Law Enforcement

  • Mandatory Training: Require all officers using AI tools to complete formal training on the technology's limitations, accuracy rates by language, and proper use cases before gaining access to the system.
  • Technical Restrictions: Implement hard technical blocks that prevent the tool from being used on prohibited language categories, rather than relying solely on policy and officer compliance.
  • Verification Requirements: Establish a mandatory verification step where a qualified human reviewer must independently confirm any AI-generated transcripts before they can be used in investigations or presented to courts.
  • Regular Audits: Conduct systematic audits of all AI tool usage to detect misuse patterns and ensure compliance with approved use cases, with results documented and reviewed by oversight bodies.
  • Community Consultation: Engage with affected communities, particularly Indigenous and minority groups, before deploying speech recognition tools that may have documented accuracy gaps for their languages.

Why Does This Matter for Criminal Justice?

The stakes in law enforcement are uniquely high. Transcription errors or hallucinations could lead officers down investigative dead ends, contaminate evidence chains, or unfairly target individuals based on fabricated statements. The fact that Whisper is 45 percent inaccurate on Māori and Pacific languages is particularly concerning in New Zealand, where these communities have historical reasons to distrust police systems.

Abdur Razzaq of the Federation of Islamic Associations noted the lack of community input: "We should have been consulted". This reflects a broader tension in public-sector AI deployment: the productivity gains promised by vendors must be weighed against measurable accuracy risks, especially when those risks fall disproportionately on underrepresented groups.

Police acknowledged the high stakes in an internal March 2025 presentation about generative AI use: "The consequences of misuse or errors were high for the public". The same document warned that "there is no specific legal framework for the use of AI in New Zealand," making due diligence even more critical.

Police

What Happens Next?

Police policy requires six-monthly audits of all generative AI use, yet as of the RNZ inquiry, no formal audit of Whisper had been completed since the trial ended in March 2025. Police stated that "every use is logged against the user for any future auditing purposes," but the absence of completed audits raises questions about oversight.

Police

The Justice Ministry is also exploring speech-to-text technology for court transcription and said both agencies are "in regular contact to share and learn from their respective approaches". This suggests Whisper's use in law enforcement may expand to judicial settings, making accuracy and hallucination concerns even more pressing.

The New Zealand police experience illustrates a recurring pattern in public-sector AI adoption: behavioral controls and training alone are insufficient without technical guardrails, independent evaluation, and clear rules about how AI outputs can be used in high-stakes decisions. As more agencies consider deploying speech recognition tools, the lessons from this trial offer a cautionary roadmap.