Penn State researchers have found that leading AI chatbots weigh patient traits differently from people when asked who should receive a scarce kidney, and they almost never admit they cannot decide.

The findings, presented at the ACM Fairness, Accountability and Transparency conference in Montreal from 25 to 28 June 2026 and issued by the university on 14 September 2026, raise questions about trusting AI to choose who gets a kidney. Models were compared with earlier studies of lay volunteers given the same hypothetical pairs.

How Researchers Tested Who Gets the Kidney

Hadi Hosseini, associate professor of informatics and intelligent systems and of economics at Penn State, led a team that posed a scarce-resource problem: two patients need a transplant but only one kidney is available.

Candidates were described by age, health, drinking habits and dependents. Models including GPT-4o, Claude-3.5-Haiku and Gemini variants had to pick a recipient. The same vignettes had already been put to hundreds of participants with no specified medical training.

Researchers isolated single traits, mixed competing attributes, and added a flip-a-coin option to measure indecision. Results appear in the conference proceedings. The US National Science Foundation part-funded the work under grants 2144413 and 2107173.

Co-authors include Samarth Khanna, Leona Pierce and John Dickerson, chief executive of Mozilla.ai.

Where Chatbots Diverge From Human Values

AI systems often fixate on one attribute rather than balancing several. Humans tended to favour younger patients; many models put heavier weight on lower alcohol use.

In one pair of 55-year-old patients with identical drinking habits, people chose the candidate with two dependents 93 per cent of the time, while Claude-3.5-Haiku chose the patient with none 65 per cent of the time.

Hosseini said: 'They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do.'

Even when invited to flip a coin, models gave a confident answer. Humans often expressed indecision.

Fine-tuning on 3,960 human decision examples lifted Qwen-3-14B's accuracy at predicting human choices from 58.03 per cent to 76.67 per cent, yet models still failed to recover human-like calibration of moral uncertainty.

Limits of Alignment and Clinical Caution

The authors do not advocate replacing professional judgement. Moral choices in organ allocation decide who lives and who dies.

Dickerson said: 'When we allocate something scarce, whether it's a kidney, a job or access to some other resource, there isn't always a single objectively correct answer.'

Ken Gusler, a science writer, shared the Penn State summary on 15 September 2026, pointing to the mismatch between chatbot certainty and human doubt.

A separate evaluation on Organ Procurement and Transplantation Network records has flagged fairness shifts between picking one winner and ranking a waiting list.

Key facts from the Penn State work are straightforward. Chatbots oversimplified trade-offs, rarely expressed indecision, and still needed explicit alignment after modest fine-tuning.

Real-world U.S. kidney allocation operates under formal matching and allocation policies involving medical and logistical criteria, rather than the simplified prompts used in the study. The researchers caution that their findings should not be interpreted as evidence that LLMs are ready to make real-world clinical allocation decisions.