A Human in the Loop Is Not Enough: “Humanwashing”, AI and the Future of Judicial Decision-Making
We keep being reassured that artificial intelligence will not replace judges because there will always be a “human in the loop”.
A judge will remain responsible.
A human will review the recommendation.
The algorithm will advise, not decide.
That sounds reassuring.
But what if the presence of the human is enough to make a system look fair — even where the human does very little?
That is the uncomfortable question raised by a fascinating 2026 study published in the Journal of Legal Analysis.
The researchers gave thousands of participants fictional sentencing, bail and arbitration cases decided in different ways:
- by a human decision-maker;
- by an algorithm;
- or through a hybrid process in which an algorithm made a recommendation and a human retained final authority.
The familiar result appeared first.
Purely algorithmic adjudication was perceived as less procedurally fair than human decision-making.
But then something much more interesting happened.
Once a human judge was put back into the process, the fairness gap largely disappeared.
And it made surprisingly little difference whether that human was described as reviewing the evidence thoroughly or only briefly.
The authors call the resulting risk:
humanwashing.
The analogy is deliberate.
Just as “greenwashing” can make something appear environmentally responsible without changing what is underneath, humanwashing could allow nominal human oversight to make an algorithmic process appear meaningfully human when most of the substantive work has already been done by the machine.
For courts, that is not simply a technology problem.
It is a constitutional one.
A judge’s signature is not the same thing as judicial judgment.
Credit and source
This article was prompted by thoughtful commentary from attorney and mediator Rasmus H Wandall, who drew attention to the concept of humanwashing and its implications for digital courts.
The underlying research is:
Benjamin M. Chen, Yoan Hermstrüwer, Pascal Langenbach, Alexander Stremitzer and Kevin Tobia, “Mitigating the judicial human–AI fairness gap”, Journal of Legal Analysis, Volume 18, Issue 1 (2026), pp 207–242.
The analysis below is JSH Law’s own interpretation of what that research may mean for justice-system design in England and Wales.
The finding in one paragraph
Across two experiments involving 7,651 US participants, purely algorithmic decisions were generally perceived as less fair than human decisions. But when a human retained final authority after receiving an algorithmic recommendation, the perceived fairness gap largely disappeared.
Most importantly, increasing the stated intensity of human review from brief to thorough did surprisingly little to improve fairness ratings. The human’s presence mattered more to perceived legitimacy than the depth of the human’s involvement.
Eight things to understand before using this research
1. The study measured perceived fairness.
It did not determine whether human, hybrid or algorithmic decisions were objectively legally correct or substantively fair.
2. These were experimental vignettes.
Participants were evaluating fictional sentencing, bail and consumer-arbitration scenarios rather than experiencing real proceedings.
3. The participants were in the United States.
The findings cannot simply be assumed to transfer unchanged into England and Wales.
4. The algorithm was described as highly capable.
Participants were told that the algorithm had outperformed humans in predictive accuracy.
5. The human retained final authority in the hybrid condition.
This was not an autonomous robot handing down a judgment with a ceremonial human observer.
6. “Minimal review” still meant some human review.
The experiment contrasted a brief examination with a thorough one.
7. The demographic findings were nuanced.
Hybrid procedures closed the fairness gap completely for some groups but less completely for others, with the strongest difference associated with the intersection of race and political ideology.
8. Humanwashing is a normative risk identified by the authors.
The study does not establish that current courts are already humanwashing decisions.
What did the researchers actually do?
The paper, published on 24 July 2026, reports two large preregistered experiments involving a combined 7,651 participants.
Participants evaluated decision-making procedures in three fictional legal settings:
- criminal sentencing;
- bail;
- and consumer arbitration.
In the first experiment, participants encountered one of five decision-making arrangements:
- Human High — a human carried out thorough review;
- Human Low — a human carried out brief review;
- Hybrid High — an algorithm made a recommendation and a human thoroughly reviewed the case before making the final decision;
- Hybrid Low — an algorithm made a recommendation and a human briefly reviewed the case before making the final decision;
- Robot — the algorithm made the decision.
Participants then rated procedural fairness and outcome fairness.
The researchers also measured perceptions including:
- accuracy;
- comprehensiveness;
- transparency;
- and empathy.
Importantly, participants were told the algorithm used machine learning and had demonstrated superior predictive accuracy to human decision-makers.
So this was not an experiment in which people were invited to compare a competent human with a visibly unreliable machine.
The human–AI fairness gap
The starting finding will surprise almost nobody.
People generally perceived purely algorithmic adjudication as less procedurally fair than human adjudication.
This phenomenon has appeared in previous research and is sometimes called the judicial human–AI fairness gap.
That makes intuitive sense.
A court is not merely a prediction engine.
People expect a judge to:
- listen;
- understand context;
- evaluate competing evidence;
- apply law;
- exercise discretion;
- give reasons;
- and take responsibility.
Even where an algorithm may be statistically accurate, accuracy is not the only characteristic citizens associate with justice.
Procedure matters.
Then the researchers put the human back
The hybrid procedure changed perceptions significantly.
The algorithm generated a recommendation.
But the human retained final authority.
That was enough, in the aggregate results, to eliminate the perceived fairness gap between purely human and hybrid adjudication.
This matters because it appears to offer institutions a very attractive compromise.
Use AI to gain:
- speed;
- scale;
- consistency;
- data processing;
- and predictive capability.
Then keep a human at the end.
Efficiency plus legitimacy.
Problem solved.
Except the next result makes the apparent solution much less comfortable.
People barely distinguished between brief and thorough human review
This is, to me, the most consequential finding in the study.
The researchers deliberately varied how deeply the human decision-maker was said to engage with the case.
In the high-review condition, the human carefully and thoroughly examined the evidence.
In the low-review condition, the human conducted only a brief examination.
Participants understood the difference.
The researchers checked that.
Yet procedural-fairness ratings were largely insensitive to it.
Most strikingly, a hybrid process involving minimal human scrutiny received fairness ratings statistically equivalent to a wholly human process involving intensive judicial review.
That should make anyone interested in judicial AI pay attention.
People may care enormously that a human judge is present — while being surprisingly poor at distinguishing meaningful judgment from thin human validation.
What is “humanwashing”?
The authors use the term to describe a potential institutional danger.
Imagine that most of the substantive work is carried out algorithmically.
The system:
- reads the evidence;
- categorises the issues;
- scores risk;
- produces a recommended outcome;
- and perhaps even generates supporting reasons.
A human then reviews the result briefly and approves it.
Formally:
a human made the decision.
Substantively:
the machine may have determined the decisional frame before meaningful human reasoning began.
If the presence of the human nevertheless restores public perceptions of fairness, institutions acquire a dangerous incentive.
They can gain the legitimacy associated with human adjudication without necessarily preserving the substance of human adjudication.
That is humanwashing.
But this study does not show that humanwashed decisions are actually fair
This distinction is essential.
The research examined perceived procedural fairness.
It did not test whether the resulting decisions:
- were more accurate;
- were less biased;
- better applied the law;
- gave better reasons;
- were more appeal-proof;
- or produced better substantive outcomes.
Indeed, that is precisely why the finding is potentially troubling.
The appearance of legitimacy may become disconnected from the quality of the underlying process.
A process can feel reassuring because a judge’s name appears at the end of it.
That tells us nothing by itself about how independently the judge reasoned.
Nor was this a study of real judges
That limitation matters too.
The participants were members of the public evaluating written scenarios.
They were not:
- judges making decisions;
- litigants experiencing actual proceedings;
- lawyers observing judicial reasoning;
- or parties trying to challenge an algorithmic recommendation.
Real proceedings involve stakes, relationships, oral advocacy, uncertainty and institutional context that a vignette cannot reproduce.
And perceptions of AI and courts may differ significantly between jurisdictions.
The authors themselves recognise those limitations.
So this should not be reported as:
“Research proves judges can rubber-stamp AI and nobody notices.”
It does not.
The defensible conclusion is narrower:
in these experiments, the presence of human final authority substantially restored perceived fairness even where the stated human review was minimal.
That raises a serious design question.
The legitimacy effect was not uniform
The second experiment explored race and political ideology in greater depth.
Here the story became more complicated.
Hybrid procedures reduced the fairness gap for both Black and White participants.
But they did not eliminate it equally.
The effect was complete among White participants but only partial among Black participants.
Further analysis suggested that the more consistent moderator was political ideology, with the divergence concentrated particularly among Black liberal participants.
This is more precise than saying:
“marginalised groups were not reassured by human oversight.”
The research does not support that broad generalisation.
But it does support something still important:
legitimacy is not distributed uniformly.
The same institutional design may reassure one community more than another.
And communities with different experiences of the justice system may interpret claims of human oversight differently.
Why this matters now in England and Wales
This research has arrived at exactly the point at which artificial intelligence is moving deeper into the justice system.
The Ministry of Justice has already announced work including:
- AI legal assistants;
- AI-assisted case analysis;
- AI-supported court listing;
- speech and transcription technology;
- data-linking tools;
- and wider AI-enabled justice services.
Its September 2026 update says the department is moving from initial foundations towards scaling successful applications across the justice system.
The Government is also expressly linking AI to:
- reducing delay;
- court capacity;
- better access to justice;
- and more efficient public services.
None of the published material I have reviewed suggests that Family Court judges are currently being asked to outsource welfare determinations to AI.
That distinction matters.
Administrative support is not algorithmic adjudication.
A listing assistant is not a robot judge.
A research tool is not a welfare decision.
But the boundary between assisting a decision and shaping a decision deserves very careful attention as systems become more capable.
The judiciary is already drawing that distinction
The current AI guidance for judicial office holders in England and Wales stresses:
- accuracy;
- bias;
- confidentiality;
- hallucinations;
- and the judge’s continuing personal responsibility.
Lord Justice Birss has emphasised that judicial use of AI must remain consistent with the integrity of the administration of justice and the rule of law.
But an especially important warning came in June 2026 from Dame Victoria Sharp, President of the King’s Bench Division.
She distinguished between AI which assists judicial work and AI which begins to influence judicial reasoning.
She also identified the danger of automation bias — the human tendency to defer to systems which appear technical, neutral and authoritative.
That warning connects directly with humanwashing.
The danger is not only that an AI produces a visibly absurd answer.
It may produce a highly polished, seemingly neutral starting point which quietly shapes the judge’s analysis.
The most dangerous AI recommendation may be the plausible one
A spectacular hallucination is easy to spot.
A fake case citation can be checked.
A fabricated statute can be rejected.
But imagine an AI system produces:
- a polished summary;
- a neat issue list;
- a risk classification;
- a coherent chronology;
- and a recommended outcome.
Nothing is obviously absurd.
But perhaps:
- one piece of evidence has been weighted too heavily;
- another has disappeared;
- a disputed allegation has been described too confidently;
- a cultural assumption has influenced the analysis;
- or the model has framed the available choices too narrowly.
The judge then begins from that document.
Even if the judge retains final authority, the system may already have influenced:
what looks important.
what looks peripheral.
what looks risky.
and what looks like the obvious answer.
That is the point at which “decision support” becomes constitutionally interesting.
The Ministry of Justice has already said “human in the loop” is not enough
One of the most encouraging parts of the current UK approach is that the Justice AI Unit has already articulated almost exactly this problem.
In May 2026, AI Fellow Professor Phil Brown asked what a genuinely human-centric approach to justice AI should mean.
His answer went beyond merely retaining a human somewhere in the workflow.
The principle was:
AI should support human judgment, not replace it.
He also argued that success should not be measured simply through:
- time saved;
- tasks completed;
- or efficiency.
The relevant outcomes are human ones:
- better decisions;
- trust;
- safer communities;
- sustainable professional workloads;
- and meaningful public-service outcomes.
That is exactly the right direction.
Human-centric AI is not AI plus a human signature.
It is AI designed so that meaningful human judgment remains real.
What could humanwashing look like in the Family Court?
Again, I want to be very clear.
I am not suggesting this is currently happening.
The study did not examine Family Court proceedings and I have found no evidence that judges in private children cases are currently being given algorithmic recommendations about whether a child should live with or spend time with a parent.
But family justice is precisely the jurisdiction in which we should think about the boundary before technology reaches it.
Imagine a future AI system capable of:
- summarising a 1,500-page bundle;
- identifying disputed facts;
- categorising allegations of domestic abuse;
- highlighting risk indicators;
- summarising Cafcass material;
- analysing historic orders;
- identifying apparently contradictory evidence;
- and suggesting issues for determination.
Much of that could be genuinely useful.
Now imagine it also produces:
“Recommended welfare outcome.”
At that point the question changes.
Has the technology helped the judge understand the evidence?
Or has it already framed the result the judge must now consciously resist if they disagree?
Family justice is particularly vulnerable to apparently neutral summaries
Family cases are messy.
Evidence can include:
- contradictory accounts;
- partial records;
- children’s reported wishes;
- professional opinions;
- historic allegations;
- findings;
- admissions;
- inferences;
- and events whose meaning depends heavily upon context.
A summary therefore requires choices.
What gets included?
What gets omitted?
How is uncertainty expressed?
Which allegation appears first?
Which fact is described as established?
Which relationship dynamic becomes the organising story?
No summary is entirely neutral.
That is true when a human writes it.
It remains true when software does.
Efficiency does not remove interpretation.
A reasoned judgment must still contain the judge’s reasons
This is perhaps the most important constitutional boundary.
A judgment is not simply an output.
Its reasons perform several functions.
They tell the parties:
- what evidence was accepted;
- what was rejected;
- which legal test was applied;
- how competing considerations were balanced;
- and why the outcome followed.
Reasons also permit:
- appellate scrutiny;
- professional accountability;
- public understanding;
- and discipline in the decision-making process itself.
If an AI recommends an outcome and then generates a convincing explanation for it, there is a profound danger:
the explanation may describe why the decision could be justified, rather than why the judge actually reached it.
Those are not necessarily the same thing.
A post-hoc rationale is not a substitute for judicial reasoning.
So what does meaningful human oversight actually require?
“Human review” is too vague.
If judicial AI develops further, we will need operational standards for what meaningful review actually means.
I would expect at least the following:
1. The human understands what the AI did
A judge cannot supervise a process they cannot meaningfully interrogate.
2. The recommendation is contestable
The human must be capable of rejecting it without friction or institutional penalty.
3. Source material remains visible
The recommendation must not replace access to the evidence upon which it depends.
4. Uncertainty remains visible
AI should not turn disputed or ambiguous matters into apparently settled propositions.
5. There is enough time to review
A theoretical right to reject an algorithm means little if workload makes meaningful scrutiny impossible.
6. The human can identify disagreement
Systems should permit — and perhaps record — where professional or judicial judgment departs from algorithmic output.
7. The system is auditable
It should be possible to reconstruct what information was used and what recommendation was generated.
8. Responsibility remains identifiable
The person exercising legal authority must remain accountable for the decision rather than blaming the software which informed it.
The JSH Law Anti-Humanwashing Test
If a justice-system decision uses AI while retaining human final authority, I would ask eight questions.
What is the AI actually being used to do?
Does it organise information or recommend an outcome?
Can the human see the evidence and assumptions underneath the output?
Can the human genuinely reason away from the recommendation?
Was there enough time for meaningful scrutiny rather than nominal approval?
Can an affected person challenge relevant AI-derived assumptions?
Are the eventual reasons genuinely the decision-maker’s?
Who ultimately answers for the decision?
If the only meaningful answer is:
“A human clicked approve,”
that is not enough.
There is another question: should people know AI was involved?
Transparency becomes particularly important where AI moves closer to adjudication.
If an AI system merely helps manage courtroom availability, disclosure to individual litigants may have little relevance.
If it meaningfully influences:
- risk classification;
- case prioritisation;
- evidential summaries;
- recommended issues;
- or substantive decisions,
the transparency question becomes much harder to avoid.
People cannot meaningfully challenge a factor affecting a decision if they do not know it exists.
The Ministry of Justice’s current AI programme expressly identifies transparency as one of its governance foundations.
That principle should become stronger as the potential impact of the system increases.
Human oversight must be designed around the risk of deference
Human oversight is often described as though adding the human automatically reduces risk.
Sometimes it will.
But people are not perfect safeguards against machines.
Humans:
- anchor;
- defer;
- become fatigued;
- trust polished outputs;
- follow defaults;
- and operate under workload pressure.
A system which produces a confident recommendation may therefore influence the human precisely because it looks objective.
Good governance should anticipate that behaviour.
It should not simply assume:
human present = risk solved.
This is also why efficiency targets matter
Imagine an AI system saves a court thirty minutes per case.
That sounds positive.
Then caseload assumptions change.
Judges are expected to decide more cases.
The time previously available for independent review disappears.
The system still formally says:
“The judge retains final responsibility.”
But the operating environment now makes genuine independent scrutiny increasingly difficult.
This is how institutional design can turn meaningful oversight into ceremonial oversight without anybody explicitly deciding to do so.
Humanwashing therefore cannot be prevented solely through a policy statement.
Workload matters.
Time matters.
Incentives matter.
Interface design matters.
Institutional culture matters.
The objective should not be to make people trust AI courts
Rasmus Wandall’s commentary makes an important normative point here.
If adding a human judge causes people to perceive an algorithmically assisted process as fair, then public trust is not necessarily proof that the process deserves trust.
That does not mean public perceptions are irrelevant.
Procedural legitimacy matters profoundly.
But confidence should follow institutional quality.
It should not become a design metric which can be optimised independently of institutional quality.
The goal should not be to design AI courts that people are persuaded to trust. It should be to design justice systems worthy of trust.
Those are different objectives.
JSH Law: use AI to support judgment, not manufacture legitimacy
JSH Law’s work on responsible legal AI focuses on the boundary between computational assistance and human responsibility.
For litigants in person, AI can assist enormously with:
- organising evidence;
- building chronologies;
- identifying repetition;
- comparing documents;
- reducing duplication;
- and preparing material for human review.
But the same principle applies at every level of the justice system.
AI can process. Humans must remain responsible for judgment, significance and consequence.
The argument against humanwashing is not an argument against judicial AI
I think this distinction is important.
There are many uses of AI in courts which may be enormously beneficial.
AI may help:
- find relevant authorities;
- manage large bundles;
- identify scheduling conflicts;
- translate material;
- transcribe hearings;
- locate duplicated information;
- organise administrative data;
- and reduce routine clerical work.
None of that requires nostalgia for paper-based inefficiency.
But as AI moves from:
processing information
towards:
influencing judgment,
governance must become more demanding.
The more consequential the recommendation, the more meaningful the human involvement has to be.
The institution matters more than the interface
The phrase “human in the loop” sounds reassuring because it evokes something familiar.
A person.
A judge.
Accountability.
Judgment.
But the Chen, Hermstrüwer, Langenbach, Stremitzer and Tobia study exposes an important weakness in that reassurance.
We may attribute legitimacy to the human’s presence without knowing very much about what the human actually did.
That creates an institutional temptation.
Automate most of the process.
Keep the judge at the end.
Preserve the appearance of human adjudication.
Gain efficiency without losing public confidence.
That is precisely why the design principle must be stronger.
Not:
“Keep a human in the loop.”
But:
“Preserve meaningful human judgment.”
Those things are not synonymous.
And nowhere will the distinction matter more than in courts deciding questions involving:
- liberty;
- risk;
- credibility;
- family life;
- domestic abuse;
- children’s welfare;
- and fundamental rights.
The constitutional safeguard is not that a human appears somewhere in the workflow.
It is that a human decision-maker genuinely listens, evaluates, reasons, decides and remains answerable for the result.
That is what judicial responsibility means.
And as AI becomes more capable, preserving it will require deliberate institutional design rather than a checkbox labelled:
Human oversight: ✓
Related JSH Law analysis
- When AI Stops Answering and Starts Acting: Agentic AI, Legal Responsibility and the Next Professional Risk
- AI May Democratise Legal Drafting. Can the Courts Survive the Volume?
- If AI Does the Junior Work, Who Trains the Lawyers of 2035?
- The Jury Trial U-Turn Is Not the End of Court Reform
- Can AI Tell the Truth? Why the Difference Matters in Law
- The JSH Law Six-Question Check
Primary and authoritative sources
- Chen, Hermstrüwer, Langenbach, Stremitzer & Tobia — Mitigating the judicial human–AI fairness gap, Journal of Legal Analysis (2026)
- Courts and Tribunals Judiciary — Artificial Intelligence Judicial Guidance, October 2025
- Justice AI Unit — What Does a Human-Centric Approach to AI Mean for the Ministry of Justice?, May 2026
- Ministry of Justice — AI tech ambition to deliver smarter justice, June 2026
- Ministry of Justice — AI Action Plan for Justice: One Year On, September 2026
- President of the King’s Bench Division — Without Fear or Favour: Judicial Independence, Past, Present and Future, 2026

© 2026 JSH Law Ltd. All rights reserved.


