Skip to content
AI.info

The Pulse

FDA Gives Cognita $1.29M to Test LLM Juries for Radiology

Cognita Imaging receives a $1.29 million FDA research contract to evaluate AI-generated radiology reports with multiple large language models. The 18-month project will examine about 1 million patient exams and send disputed cases to radiol

FDA Gives Cognita $1.29M to Test LLM Juries for Radiology

AI.info Team ·

A million exams will test whether AI can grade AI

The U.S. Food and Drug Administration has awarded Cognita Imaging a $1.29 million research contract to test whether multiple large language models can evaluate AI-generated radiology reports at a scale beyond conventional reader studies. The 18-month project took effect on June 22, 2026, and will examine approximately 1 million patient exams.

Cognita, a wholly owned subsidiary of Mosaic Clinical Technologies, calls the proposed method “LLMs-as-a-jury.” Rather than asking one model to judge another model’s report, the framework will compare the assessments produced by several models before directing clinically important disagreements to radiologists.

The announcement was published by Mosaic on September 16. The company says the work is intended to address a specific regulatory problem: complete radiology reports are harder to evaluate than systems that identify one finding or flag one region of an image.

Read the source announcement from Mosaic Clinical Technologies.

Akshay Chaudhari leads the FDA project

Akshay Chaudhari, Ph.D., Cognita co-founder and an associate professor of radiology and biomedical data science at Stanford University, leads the work as principal investigator. Louis Blankemeier, Ph.D., Cognita’s co-founder and chief executive officer, serves as co-investigator.

“Once an AI system starts writing the entire report, the evaluation problem changes. A few hundred cases may tell you whether a model works in a narrow setting. They do not show every way it can fail in practice. We want to test whether a group of language models, with radiologists reviewing the difficult cases, can give us a clearer and more repeatable picture at real-world scale.”

Akshay Chaudhari, Ph.D., Cognita co-founder and principal investigator

The project will evaluate both human-drafted and AI-generated reports. Researchers will study results across patient groups, care settings, imaging equipment and diseases, including uncommon findings that may not appear often enough in smaller validation samples.

Radiologists will investigate disagreements

Automated scoring will not serve as the final authority. Radiologists will review clinically important disagreements and determine whether the problem originated in the AI-generated report, the language-model jury or the original radiologist’s report.

Nina Kottler, M.D., chief medical AI officer at Mosaic Clinical Technologies, said the study is designed to expose failures that can disappear inside carefully selected datasets.

“Radiologists know the edge cases matter. A model can look strong in a carefully selected study and still struggle at another hospital, on different equipment or with a rare presentation. This work lets us look for those differences at a scale that would be very difficult to reach through reader studies alone, while keeping radiologists involved where their judgment matters most.”

Nina Kottler, M.D., chief medical AI officer, Mosaic Clinical Technologies

Cognita will also recreate smaller validation cohorts from the larger dataset. The comparison is intended to show what a study may miss when it relies on a limited number of cases, particularly when rare diseases, unusual presentations or differences between imaging systems affect report quality.

From GREEN to regulatory guidance

The FDA-funded work builds on Cognita’s GREEN project, an open-source tool that compares reference radiology reports with AI-generated reports and identifies clinically meaningful differences. The company plans to use clinical expertise, technology infrastructure and radiology workflows from Mosaic and the broader Radiology Partners ecosystem.

Under the contract, Cognita will provide the FDA with software code and practical guidance for constructing LLM juries. Deliverables will also include analyses comparing large and small validation cohorts, studies of discrepancies reviewed by radiologists, and reports describing the project’s findings and limitations.

The FDA awarded the contract through its Broad Agency Announcement program for advanced research and development in regulatory science. The agreement covers an evaluation method, not a commercial product authorization: Cognita and Mosaic state that it does not represent FDA endorsement, approval or clearance of any Cognita technology.

The test is whether consensus catches what samples miss

LLM-based evaluation can process far more reports than a traditional expert panel, but the project is built around the possibility that automated judges will also make mistakes. Cognita’s design therefore places human review around the cases where the models disagree or where the clinical consequences may be significant.

Success will depend on whether the jury’s consensus tracks radiologists’ judgments across the full range of exams, rather than only producing agreement among language models. The contract’s concrete test is scheduled across approximately 1 million patient exams, with the FDA set to receive code, cohort analyses, discrepancy studies and reports on limitations over the 18-month research period.

Source

Mosaic Clinical Technologies

Explore

More articles