Logo
FrontierNews.ai

Pakistan's AI Judge Tool Closes 6.3% More Cases Without Sacrificing Quality

A large-scale trial in Pakistan found that a specially designed AI tool built on OpenAI's GPT-4 language model boosted case closures by 6.3% with no measurable drop in judgment quality. The experiment, conducted across roughly half of Pakistan's trial judges, demonstrates how large language models (LLMs), which are AI systems trained on vast amounts of text to understand and generate human language, can address real-world bottlenecks when properly designed and deployed with adequate training.

Pakistan's judiciary faces a crisis of scale. With a backlog of 2.26 million cases and fewer than two judges per 100,000 people, compared to 22 in the European Union and 8 in Brazil, the system was desperately in need of help. Economist Sultan Mehmood of the New Economic School in Moscow, along with collaborators, designed JudgeGPT to tackle this problem by combining GPT-4 with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes.

How Does JudgeGPT Actually Work?

The tool uses a technique called retrieval-augmented generation (RAG), which allows the AI model to search through a database of Pakistani legal documents and return results with footnotes linking directly to relevant cases and laws. This approach addresses a critical weakness in general-purpose AI chatbots: they often "hallucinate," or confidently generate false information. By tethering the model to a verified database, JudgeGPT grounds its responses in real legal precedent.

  • Legal Research Acceleration: Judges can query the system with plain language and receive relevant case law instantly, eliminating hours of manual searching through precedents and statutes.
  • Document Summarization: The tool can condense lengthy legal documents into concise summaries, extracting the key points judges need to make informed decisions.
  • Judgment Drafting Support: JudgeGPT assists with structuring and writing judicial opinions, though the researchers found that roughly one-fifth of judges asked the tool to make substantive decisions with minimal human input, a practice the training sessions worked to discourage.

What Did the Trial Actually Show?

Between 2024 and the study period, 1,559 trial judges were offered access to JudgeGPT. The researchers divided them into three groups: those who received six 90-minute training sessions on how LLMs work and their limitations, those who received only general technology training, and a control group with no training. The results were striking.

Judges who completed the specialized training logged into JudgeGPT an average of 56 times and submitted 212 prompts over the study period. By contrast, judges with only generic training logged in just 10 times and submitted 25 prompts. Those with no training used the tool for about a month and then abandoned it entirely. The median district saw a 6.3% increase in resolved cases, and the more trained judges in a district, the larger the effect.

To assess whether faster case resolution came at the cost of quality, the researchers asked OpenAI's GPT-5-mini model to compare pairs of judgments from the same judge before and after training. The AI chose post-training judgments 59% of the time. Two experienced Pakistani lawyers evaluated 90 judgment pairs and agreed with the AI model's assessment 70.6% of the time, compared to 73% agreement with each other. Appeal rates also fell slightly, suggesting that judges were not rushing through cases.

"We do find an increase in cases resolved, and we don't find any corresponding decrease in decision quality," said Sultan Mehmood, economist at the New Economic School in Moscow.

Sultan Mehmood, Economist at the New Economic School in Moscow

The financial impact is equally compelling. A trained judge resolved an average of 38.5 more cases per month than the baseline, translating to roughly $38.50 saved in judicial costs for every dollar spent running the tool. While a 6.3% productivity boost might sound modest, it represents meaningful progress in a system where delays have created immense human suffering.

Why Does Training Matter So Much?

One of the study's most important findings is that simply handing judges a powerful AI tool does not guarantee adoption or effective use. Training proved vital. Judges who understood how LLMs work, their limitations, and the risks of bias and hallucinations used the tool persistently and more thoughtfully. Those without training either abandoned it or relied on it inappropriately, asking it to make substantive legal decisions rather than using it as a research and drafting aid.

"Just giving people the technology does not necessarily make them use it persistently," noted Sultan Mehmood.

Sultan Mehmood, Economist at the New Economic School in Moscow

This finding has implications far beyond Pakistan. As AI tools proliferate across professional sectors, the quality of implementation depends not just on the technology itself but on how well users understand its capabilities and constraints. The researchers emphasized that training lowered the proportion of inappropriate delegation, where judges asked the tool to make decisions rather than assist with research.

What Do Experts Say About AI in the Courtroom?

The Pakistan trial is the first large-scale independent assessment of ongoing judicial use of AI. Other countries, including Brazil and India, are rolling out AI tools for judges, and prominent U.S. law professor Eric Posner has conducted single-case studies comparing LLM judgments to human ones. However, the Pakistan experiment stands out for its rigorous methodology and real-world scale.

David Autor, an economics professor at the Massachusetts Institute of Technology, praised the work: "It's not easy to do large-scale field experiments in civil service, but especially where the stakes are so high." He noted that while a 6.3% productivity boost is not overwhelming, it is credible and likely to improve as the tool is refined and more widely adopted.

However, experts caution that efficiency is not the only measure of a justice system. John Zeleznikow, a professor of law and technology at La Trobe University in Australia, observed that while the trial successfully addressed speed and case volume, "what's not that clear is whether what you call the quality of justice is better." He emphasized that AI can be useful only if judges remain diligent about evaluating and verifying the tool's output.

"There are risks for using these AIs for sure, even with all these safeguards, but at some point you have to just put the judges in as strong a position as you can," said Elliott Ash, an associate professor of law, economics, and data science at ETH Zurich.

Elliott Ash, Associate Professor of Law, Economics, and Data Science at ETH Zurich

The Pakistan trial demonstrates that when AI tools are purpose-built for a specific context, grounded in reliable data, and paired with meaningful training, they can deliver measurable benefits without compromising core values. As courts worldwide grapple with backlogs and resource constraints, this experiment offers a blueprint for responsible AI deployment in one of society's most critical institutions.