Benchmarking clinical LLMs from OpenEvidence and Doximity proves complicated
As clinical chatbots proliferate in healthcare settings, the difficulty of establishing reliable performance benchmarks raises questions about how providers and patients can evaluate these AI tools before adoption.
AI · groundedMixed sourcing