HealthNotableSingle-sourceDeveloping
5.2
Medical AI developers challenge current safety and accuracy benchmarking standards
STAT NewsLO·US·1 day ago
The evaluation of clinical chatbots from OpenEvidence and Doximity faces significant methodological hurdles regarding accuracy and reliability. It remains unclear how standardized benchmarks will translate to clinical safety, a critical factor for adoption in biopharma and healthcare settings.