REFINe-ing AI’s Role in Nephrology Diagnosis

#NephJC Ten POsts Chat

September 1st, 2026, 9 pm EST

Kidney Int Rep,  2026 Jun 23;11(9):106673., doi: 10.1016/j.ekir.2026.106673.eCollection 2026 Sep.

Randomized Controlled Trial of Large Language Model–Assisted Diagnostic Accuracy in Nephrology 

Raphaël Bentegeac, Philippe Amouyel, Bastien Le Guellec, Wisit Cheungpasitport, Mehdi Maanaoui, Aghiles Hamroun

PMID: 42502668
DOI:  
10.1016/j.ekir.2026.106673

Introduction

Modern Artificial Intelligence (AI) encompasses a spectrum of computational approaches, from traditional machine learning to deep neural networks capable of processing increasingly complex data (see NephMadness 2026, AI region). Among the most rapidly evolving applications of deep learning are large language models (LLMs) that are trained on vast amounts of information and can interpret, synthesize, and generate natural language.  Since the release of ChatGPT in late 2022,  medicine has been grappling with a collective identity crisis.  In the last few years, there has been an expanding body of research evaluating the performance of different LLMs across a wide range of medical tasks, from answering clinical questions and interpreting complex case vignettes to diagnostic reasoning and medical decision-making. A recent systematic review (Chen SF et al, Nat Med, 2026)  found that LLMs outperformed human comparators in 33% of studies, with better relative performance in knowledge-based assessments than in real-world clinical settings, highlighting the persistent gap between benchmark performance and clinical utility. Essentially, AI programs are very good at multiple-choice exam questions.

Nephrology has not been an exception to this trend. In an evaluation of 975 questions from the Nephrology Self-Assessment Program (NephSAP) and Kidney-Self-Assessment Program (KSAP), GPT-4 achieved 74% accuracy, although remaining below the 77% passing threshold and the average performance of nephrology examinees (Miao J et al, Clin J Am Soc Nephrol, 2024). Moreover, beyond knowledge-based assessments, the potential applications of LLMs in nephrology have rapidly increased to include patient management, clinical documentation, decision support, and medical training. Evidence supporting the safety and effectiveness of LLMs in real-world nephrology settings remains limited (Yongzheng Hu et al, Ren Fail, 2025). 

Uses and limitations of AI in Nephrology, VA by Cristina Popa

Against this vast background, the REFINe (Reasoning Enhancement With Feedback From a Generative AI in Nephrology) was conducted. Rather than asking whether a LLM could independently solve nephrology questions, REFINe addressed a more clinically relevant question: Can access to a high-reasoning LLM improve physicians diagnostic accuracy when faced with complex, open-ended nephrology cases?

REFINe tested whether AI can better diagnosticians of physicians- a more relevant question for the future of AI-assisted medicine. 

Methods

REFINe was a prospective, randomized, open-label, parallel-group superiority trial conducted entirely online between November 2025 and March 2026. English or French-speaking residents and board-certified physicians were randomized 1:1 to diagnose complex nephrology cases with or without LLM assistance. Each case evaluation was one unit of analysis. The review board of lille University Hospital approved the study. The study group planned to enroll 100 participants and recruited 97. They report no change to the criteria, the outcomes or the analyses after the trial started. 
The cases came from the Make Your Diagnosis series in Kidney International, published between August 2007 and January 2025. Two nephrologists (N.F. and A.H.) rewrote 245 of these cases. The aim was to stop the LLMs from recognizing cases used during training. Study designers changed every sentence and every number but kept the clinical logic and the final diagnosis. Ten cases were then chosen at random for the trial. 
It is worth noting what those ten cases detailed, because this shapes how we should interpret percentages in the paper. There were four glomerular and vascular cases, two tubular, interstitial and cystic disorders, two electrolyte and acid disorders, one case of acute kidney injury in critical care, and one transplant case. Seven of the ten included kidney or allograft pathology. One included urine microscopy. Two were text only. One was pediatric. This question inventory tells us more about the difficulty than any adjectives or descriptors could. 

Participants were residents or board-certified doctors who spoke English or French. They were recruited through professional societies and conferences. The main exclusion criteria was previous contact with the specific cases. The team used the Pocock and Simon minimization algorithm (80% fixed and 20% random) to balance five features: certification status, speciality, years of experience, how often the doctor used LLMs, and academic status. The process was automatic, and the allocation sequence was hidden from the study staff. This is a stronger allocation method than most similar AI studies have used. 

Each participant received the 10 cases, in random order, on a dedicated platform. They could not go back to review a previous case. In the control group, doctors gave up to three diagnoses and a confidence score, then moved directly to the next case- no AI input at any point. In the AI group, participants completed the same first step, and the platform showed them a fixed GPT-5 answer with short, structured reasoning: they could revise their diagnoses and confidence once before moving on. The AI answer came from a two step prompt with no examples given to the model. The first step returned exactly three ranked diagnoses. The second step turned them into case specific reasoning with probabilities and a short summary of the key differences. All GPT-5 answers were produced in October 2025 and then frozen (this was a fixed ranked list with probabilities, not a conversation). 

Figure S1. Study design, from Bentegeac R, et al, Kidney International Reports, 2026

Then two board-certified nephrologists scored the main outcome. They worked independently and did not know from which group each answer was obtained. They agreed on about 97% of answers before any discussion, and the rest were solved by consensus. For emphasis, the whole result of this trial rested upon the “yes” or “no” judgments about free text answers by the adjudicators. 

The sample size came from a simulation. It assumed 50 participants per arm, 10 cases each, 50% accuracy in the control group, and 70% in the AI group. With 100 participants, the power was calculated to be above 99%. The main analysis used mixed-effects logistic regression, with a fixed effect for the AI study arm and random intercepts for the participant cases. It was adjusted for gender, country, speciality, academic status, experience, confidence in AI, how often the doctor used LLMs, and how many AI tools they used. Every randomized participant was analysed in the assigned group. The subgroup analyses were planned in advance. The analyses by number of cases completed were done afterwards. The R 4.3.2 software environment was used for data analysis, statistical modelling, and graphics.

REFINe Population

Table of methods summary created from Bentegeac R, et al. Kidney Int Rep, 2026.

Results

The 97 participants produced 556 case evaluations. The median age was 38 yo (IQR 34 to 45). Three out of four said nephrology was their speciality, and 69% did not work in academic centres. Nine countries were represented, but 87% to 93% of participants were French, depending on the arm. The two groups were balanced at baseline. About four in ten participants had never used an LLM. 

Primary Outcome

Table of results created from Bentegeac R, et al. Kidney in Rep, 2026.

Secondary Outcomes

Doctors only took half of the good advice!

There were 178 answers that were wrong at first where the AI suggestion prompted the correct answer. Of those, 85 (47.8%) were corrected, and just under half were not. In the AI arm, doctors changed their top-3 in 47.4% of answers (from 41.3% to 53.6%).

Harm was rare

Only 1 of 69 answers that were right at first became wrong after a bad AI suggestion (1.5%; 95% CI: 0.04 to 7.9). In this format, doctors were not simply following the machine. 

No subgroup effect reached significance (all interaction p > 0.05). The analyses by number of cases completed gave similar effect sizes, but they also show something the main paper does not discuss.

One pattern stands out despite the null result: monthly LLM users showed by far the largest AI benefit (OR 10.51, 95% CI 3.18-24.67), well above never users or daily users, and the LLM-usage interaction (p=0.008) was the closest to significance of any subgroup tested (supp figure 3).

Table of results created from Bentegeac R, et al. Kidney in Rep, 2026.

The adjusted odds ratios were 4.3 (1.99 to 9.28) for doctors who completed at least five cases, and 3.94 (1.70 to 9.13) for those who completed all ten. 

Before the trial, the team tested eleven models on all 245 cases. This took about three hours and cost around EUR 300. GPT-5 was the best performer in the pre-trial benchmark (77.1%), so it was the model used in the AI arm.

Supplement figure S2 adapted from Bentegeac R, et al. Kidney in Rep, 2026.

DeepSeek V3.2 and MedGemma 27B were tested with text-only inputs because those setups could not accept images. Their scores therefore, mix model ability with a missing input. 


Funding: The authors received no specific funding for this work. 


Data availability statement: Participant consent did not include permission for public data sharing. The analysis code will be made publicly available in a Git repository.

Discussion

From AI versus nephrologists to AI with nephrologists

LLMs have already shown that they can readily solve board-style questions. This study has moved the field of study from multiple-choice questions to complex, integrated clinical vignettes requiring the synthesis of history, laboratory data, pathology, urine microscopy and images. This trial addresses a very practical and pressing question: Does giving clinicians access to an LLM actually improve diagnostic reasoning in nephrology? In this RCT of fixed scenarios, the answer appears to be yes. While the results seem like a slam dunk endorsement of our machine overlords (we will add your biological and technological distinctiveness to our own; resistance is futile), we need to read in between the lines. Top-3 and Top-1 diagnostic accuracy improved by around 20% with AI, and the effect was preserved among participants who completed at least 5 cases, but it could have been an even greater improvement if the physicians had trusted the AI more. 
Clinical nephrology is open-ended, and new clinical information is integrated in ways particular to our specialty. Our diagnoses rely on the incorporation of heterogeneous data types, including a urinary sediment, light microscopy and immunofluorescence pattern, blood gases, and a serology panel that half fits multiple etiologies (diseases rarely read textbooks and don’t deal in absolutes). There is rarely a diagnosis that relies solely on an individual, decisive test. Answers may depend on trajectory rather than on a point value, since the same creatinine of 2.4 mg/dL means one thing over three days and something entirely different over three years. Finally, nephrologists draw on a long tail of low-prevalence entities, precisely the region in which a probabilistic model may either shine or possibly exaggerate the exotic. 

The human - AI gap may be the real story

There were 178 occasions where the clinician's initial answer was wrong, but GPT-5 gave the correct diagnosis. Yet only 85 (47%) of those answers were eventually corrected. This is probably one of the most important findings of the study. The problem is not always whether AI can correctly determine the diagnosis, but rather whether physicians are willing and/or able to recognize their mistakes. When it comes to correct diagnosis, humans are subject to errors in logic, including anchoring, premature closing, and confirmation bias (Mutlack Z, et al. J Clin Med 2025). AI assistance brought clinicians substantially closer to the apparent performance ceiling of the model, but nowhere near all the way (maybe the real trial was about us, not LLMs). That is where this paper becomes less about how clever GPT-5 is and more about how we, as clinicians, respond when AI disagrees with us. Do we trust it, challenge it, ignore it, or simply refuse to change despite the evidence? 

Why does a tool capable of very high diagnostic accuracy not automatically transfer that accuracy to the clinician sitting in front of it? The study workflow and the AI interaction here were very artificial and may have contributed to the gap. GPT-5 produced exactly three ranked diagnoses with concise reasoning and probability estimates. The answer was pre-generated and frozen. There was no chat, no follow-up, no chance to ask, “Why not my diagnosis?” or “Which one feature in this case makes you think that?” That matters because most of us do not use AI like a static answer key. We question it, push back, ask for discriminating features, ask it to argue against itself, and sometimes keep asking until either we are convinced or the model starts hallucinating creatively enough that we close the window. So, the question is not simply whether GPT is accurate. It is how much of that accuracy can be transferred to a clinician in a safe and useful way.

What has been examined previously

Other studies with the use of LLMs and physicians diagnosing complex cases have had mixed results. Goh and colleagues found no benefit at all from giving doctors an LLM, while the model on its own beat both groups (physicians alone and AI assisted physicians) by 16 points, similar to the current study (Goh E, et al. JAMA Netw Open, 2024). Everett and colleagues did improve accuracy with collaborative workflows (Everett SS, et al. NPJ Digit Med, 2026). We know LLMs keep evolving (quickly), but human integration may take longer due to trust and fear of machine hallucinations. Presentation of source materials might help physicians be more accepting of AI prodding changes, especially if the logic is solid and the primary references could be reviewed. Again, a conversation versus a mandate seems to work better in altering a doctor’s clinical reasoning.

Many previous studies scored the quality of the reasoning against a rubric (Goh E et al, JAMA Netw Open, 2024| Everett SS et al, npj Digit Med, 2026| Qazi IA, et al. Nature Health, 2026). REFINe scored a binary “yes” or “no”: was the right diagnosis somewhere in the list. The nephrology study mixed a percentage with a completeness rating. 

We don't want doctors to blindly accept whatever an AI says. If AI gives a diagnosis that doesn't fit the patient, it should be questioned, because currently AI has its own set of logistical errors (hallucinations, prejudice due to training errors, and context loss). There’s another side to this; even if the AI is actually right, our own clinical experience and preconceived notions can make us dismiss it. The trial results show that the ability to decide when to listen to AI may be just as important as the AI's accuracy. The authors make a similar point, suggesting that AI literacy and how the interface feels to us may be important steps in human/AI integration to achieve the best results.

Considering the reverse scenario, among the responses that were initially correct, very few became wrong after viewing the AI suggestion. This means the participating clinicians didn’t simply abandon a correct diagnosis whenever the AI disagreed with them. That is reassuring, yet we should be cautious, because the study was small and the AI was presented in a very controlled format somewhat dissimilar from real world clinical reasoning.

It is important to note, the control group in this study did not have access to the usual clinical resources. That is quite different from trials in which doctors could use Google, UpToDate, or other conventional resources. In addition, LLMs are evolving at an incredible rate, making what’s true in machine logic yesterday almost obsolete in just a few weeks. We cannot simply say that REFINe proves that LLMs are better than physicians (this sounds like a headline AI would write). It just shows that adding this particular AI workflow to unaided clinical reasoning improved diagnostic performance. In clinical practice, AI is being integrated in a somewhat haphazard way, slowing potential real world improvements. However, AI will be a useful tool for clinicians who choose to enhance diagnosis and treatment strategies with an eye on confirming machine-generated recommendations.

The number we are highlighting is the distance between 57.5% (physician enhanced with AI) and 77.1% (AI alone), and the fact that half of the correct suggestions were simply not used. That gap is not an AI model problem. It is a problem about how we read, trust, and wrestle with these tools. Qazi and colleagues showed just how much this can be moved: an AI-literacy curriculum alone lifted diagnostic assistance accuracy from 42.6% to 71.4%, a 27.5 point gain - arguably a bigger lever than the model itself (Qazi IA et al, Nature Health, 2026). It is becoming quite apparent that AI literacy is no longer optional for the modern physician. Mastery of AI (a potentially daily consulted assistant) sits closer than the urine sediment and biopsy interpretation in core kidney education than most of us are comfortable admitting. Similar to the acceptance of the electronic medical record, AI will be transformational once everyone gets on board.

Trial comparison infographic by Daniel Ramirez and Assad SM

Strengths

Methodologically, there is quite a lot to like. Group allocation used the Pocock-Simon minimization algorithm, with 80% deterministic and 20% random assignment. This balanced certification status, specialty, years of experience, LLM usage frequency and academic status. Randomization was automated and allocation remained concealed from study personnel during enrollment. The primary analysis used mixed-effects logistic regression with random intercepts for both participant and vignette, which is important because each clinician contributed multiple cases and each vignette had a different intrinsic level of difficulty. The model also adjusted for gender, country, specialty, academic status, experience and AI-use characteristics. Although the intervention itself could never be blinded, the diagnostic outcome adjudication was. This is a stronger allocation method than most similar AI studies have used. 

Weaknesses

The sensitivity analysis is also reassuring as the benefit did not disappear when the analysis was restricted to people who completed more of the task. At the same time, 97 participants generated only 556 completed evaluations, despite being assigned up to 10 cases each, and that’s before counting the 46 signups (32% of all registrants) who never completed a single case and were excluded from analysis all together. That raises a question: were difficult cases more likely to be abandoned or did participants simply decide that ten consecutive renal diagnostic puzzles were enough nephrology for one evening? Completion was voluntary and uncompensated, so informative non-completion remains worth thinking about even though the sensitivity analyses are reassuring.

Baseline confidence was similar between control and AI groups, but confidence increased significantly after exposure to AI (p<0.001). That makes intuitive sense if someone or something has just agreed with you, produced three polished diagnoses, and assigned probabilities to them. You naturally feel more certain. But there is an important caveat here. The confidence has to move in the correct direction, not just increase. If AI makes a correct clinician more confident, that is useful; if it makes a wrong clinician more confident, that may be dangerous.

The control arm also had no access to conventional resources. That makes the trial clean experimentally, but it is not quite how nephrology is practiced. If we come across a rare genetic tubulopathy, an unusual biopsy or a toxic alcohol case, or a lot many numbers, we may look something up. That is not cheating; that is how nephrology is currently practiced. The clinically relevant future comparison is therefore probably usual resources versus usual resources plus LLM. 

Conclusion

This is the first randomized trial in nephrology to show that a structured LLM workflow improves diagnostic accuracy, with very few errors introduced by the AI. Although this study may not currently have real world implications, the growing breadth and speed of LLM evolution will make it an essential tool for all future physicians.

One last detail, buried in the disclosures: the authors used Claude Sonnet 4.5 and GPT-5 to help write and edit this manuscript. A paper about human-AI diagnostic collaboration was itself a product of human-AI collaboration, which feels about right.   

Summary by

Anca Elena Stefan
Nephrology specialist, Assistant Lecturer
Romania

Assad S M

Assistant Professor, Nephrology

Christian Medical College, India

Srinivasavaradan Govindarajan
Assistant Professor (Ped Neph)
VMMC and Safdarjung Hospital, India

Luis Daniel Ramírez-Calvillo

Nephrology Fellow
Instituto Nacional de Cardiología Ignacio Chavez, México

Reviewed by

Akshaya Jayachandran, Brian Rifkin, Cristina Popa

Special thanks to guest reviewer & author

Wisit Cheungpasitporn (The AI guy)

Header created by AI and prompts from Assad SM