← Glossary Term
RadLE
A benchmark testing AI models on real radiology cases.
RadLE is a benchmark that measures how well AI models read medical images like X-rays and CT scans, using cases vetted by radiologists. Version 2 added something important: testing whether a model knows when it is unsure.
In medicine, a model that says “I’m not certain, ask a human” is far more useful than one that guesses confidently. Benchmarks that reward honest uncertainty are a quiet but important shift.
Mentioned in