# Study finds clinicians use AI assistants differently than benchmark tests

> Analysis of 127,833 queries shows most real‑world use involves documentation and info retrieval, which benchmarks largely miss.

Oossa · 2026-10-09 · https://oossa.com/en/study-finds-clinicians-use-ai-assistants-differently-than-benchmark-tests

Researchers examined 127,833 requests made to a hospital’s large language model (LLM) assistant by 6,342 doctors, nurses and other providers over eight months. They found that 36 % of the queries were about documentation and administration and 29 % were for knowledge retrieval, while only 4 % asked for diagnoses. More than a third of the requests could not be answered well as posed. When they compared these real‑world tasks to 58 public benchmarks, the benchmarks covered only about a third of the task mix and omitted documentation work entirely.

## The facts

- 127,833 queries from 6,342 clinicians across 35 specialties
- Benchmarks shared only 31 % of the real‑world task mix

## Why it matters

Hospitals may need to redesign AI evaluation to reflect everyday documentation and information‑search tasks that clinicians actually rely on.

## Sources & references

1. [Clinician use of language models diverges from how the models are evaluated](https://arxiv.org/abs/2610.11069) – arXiv, 2026-10-09

Last updated: 2026-10-09
