Oossa

Study finds clinicians use AI assistants differently than benchmark tests

Analysis of 127,833 queries shows most real‑world use involves documentation and info retrieval, which benchmarks largely miss.

NoteBy Published by Oossa: 1 min read

Researchers examined 127,833 requests made to a hospital’s large language model (LLM) assistant by 6,342 doctors, nurses and other providers over eight months. They found that 36 % of the queries were about documentation and administration and 29 % were for knowledge retrieval, while only 4 % asked for diagnoses. More than a third of the requests could not be answered well as posed. When they compared these real‑world tasks to 58 public benchmarks, the benchmarks covered only about a third of the task mix and omitted documentation work entirely.

Why it matters

Hospitals may need to redesign AI evaluation to reflect everyday documentation and information‑search tasks that clinicians actually rely on.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.