Pakistan's AI Judge Assistant Cleared 38 Extra Cases a Month Per Judge. Here's What Actually Worked.

A real-world trial gave 1,559 Pakistani judges a custom AI legal tool. Productivity rose 6.3%. The secret ingredient was not the software.

AI2Day Newsdesk5 min read
Photoreal news-editorial overhead shot of a darkened government data center aisle with cool blue server rack indicator lights stretching to vanishing point, fai
Share

Key points

  • A trial across 1,559 Pakistani judges found a custom AI legal tool raised cases resolved by 6.3% over nine months in 2024.
  • Pakistan's courts carry a backlog of 2.26 million cases and have fewer than two judges per 100,000 people, compared to 22 in the EU.
  • Trained judges completed roughly 38.5 more cases per month, saving an estimated $38.50 in judicial costs for every $1 spent running the tool.
  • Judges who received six training sessions logged into the tool 56 times on average and sent 212 prompts, versus 10 logins for those who got only generic training.
  • About one in five prompts involved judges asking the AI to make decisions or write opinions with little human input, raising ongoing quality concerns.

Pakistan's courts are drowning. More than 2.26 million cases sit in a backlog, and the country has fewer than two judges for every 100,000 people. The European Union manages 22 judges per 100,000; Brazil runs 8. Something had to give.

So economist Sultan Mehmood of the New Economic School in Moscow, together with collaborators including Elliott Ash of ETH Zurich, built a tool called JudgeGPT and ran what may be the first large-scale independent test of AI inside a working court system. Their findings, first reported by IEEE Spectrum AI, land at an awkward moment: judges in several countries have already been caught quietly using commercial AI chatbots in ways courts never approved.

What is JudgeGPT and how does it work?

JudgeGPT is a custom AI assistant built specifically for Pakistani courts. The team layered OpenAI's GPT-4, the large language model (an AI system trained on vast amounts of text to generate and analyse language) that also powers many commercial chatbots, on top of a database of 128,292 Pakistani judicial opinions and 943 statutes.

The key fix for accuracy was a technique called retrieval-augmented generation, or RAG. Think of it as giving the AI a private library to check before it answers. Instead of guessing at case law, the tool searches the database, pulls the relevant documents, and shows its sources in footnotes. Ash puts it plainly: "It turns out that actually the way to fix hallucinations isn't just more intelligent models. It's to attach the models to a tool that can do a search and verify the sources." Hallucinations, in AI terms, means the model confidently inventing facts that don't exist.

Commercial chatbots tested earlier had failed badly on Pakistani legal queries, regularly fabricating case citations. JudgeGPT was built to stop that.

Did it actually make judges better or just faster?

Faster, measurably. Better, probably, though with caveats.

By the time 487 judges had completed the training programme, the median district recorded a 6.3% jump in resolved cases. Appeal rates fell slightly, a sign that speed was not obviously coming at the cost of accuracy. Each trained judge resolved around 38.5 more cases per month than the baseline.

Metric Result
Judges in trial 1,559
Cases in backlog 2,260,000
Productivity increase 6.3%
Extra cases per trained judge per month 38.5
Cost saving per $1 spent $38.50
Post-training logins per judge 56
Post-training prompts per judge 212

To assess quality, the team used OpenAI's GPT-5-mini (a lighter, faster version of the latest GPT model) to compare pairs of judgements from the same judge, before and after training. The AI picked the post-training judgement 59% of the time. Two experienced Pakistani lawyers agreed with those picks 70.6% of the time, and with each other 73% of the time, suggesting the AI quality score is a reasonable but imperfect proxy.

David Autor, an economics professor at MIT, called the result credible. "The 6.3% productivity boost is not overwhelming," he said, "but it's credible and likely to improve as the tool is more widely used."

What is the honest concern here?

Training mattered more than the software itself. Judges who received six 90-minute sessions covering how the AI works, its limits, and the risk of bias and hallucinations used the tool consistently for the full nine months. Those who got only generic technology training logged in a handful of times. Those who got nothing tried it for about a month and stopped.

The bigger worry is delegation. Roughly one in five prompts asked JudgeGPT what the right decision should be, or told it to write a legal opinion with minimal judge input. That crosses a line most legal systems draw firmly. Training reduced that proportion, but it did not eliminate it.

John Zeleznikow, a professor of law and technology at La Trobe University in Australia, frames the gap well: "What's not that clear is whether what you call the quality of justice is better."

For the judge who spoke to the researchers anonymously, though, the daily reality is simpler. They have carried more than 1,000 active cases for over a decade. "For research, it's just one prompt away," they said, "whereas before I had to search for the precedents and laws for hours."

The takeaway: If your workplace is piloting an AI tool, push hard for structured training, not just a login and a password. This study shows that access without education produces a month of curiosity and then abandonment. Six sessions changed everything.

© 2026 AI2Day