Researchers examined 127,833 requests made to a hospital’s large language model (LLM) assistant by 6,342 doctors, nurses and other providers over eight months. They found that 36 % of the queries were about documentation and administration and 29 % were for knowledge retrieval, while only 4 % asked for diagnoses. More than a third of the requests could not be answered well as posed. When they compared these real‑world tasks to 58 public benchmarks, the benchmarks covered only about a third of the task mix and omitted documentation work entirely.
Why it matters
Hospitals may need to redesign AI evaluation to reflect everyday documentation and information‑search tasks that clinicians actually rely on.