# Qwen3-VL 8B tested on 137 messy documents using a laptop

> A Reddit user compared a small vision model running locally with three hosted models on receipts, forms, invoices and contracts. In the user’s test, Qwen did well on tax forms but struggled with Indian date formats and long contracts.

Oossa · 2026-09-28 · https://oossa.com/en/qwen3-vl-8b-tested-on-137-messy-documents-using-a-laptop

A Reddit user tested Qwen3-VL 8B Instruct on an M5 laptop against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra, using 137 documents. The user reports that documents were fully correct in 59% of Qwen’s results, compared with 89% for Opus and 57% for GPT-5.6 Terra. Qwen got 21 of 32 IRS forms fully right, while GPT-5.6 Terra got seven. But on 10 synthetic Indian bank statements, Qwen read dates in day-month-year format as month-day-year, despite getting the amounts and balances right.

## The facts

- The test covered 137 receipts, invoices, IRS forms, bank statements and contracts.
- Qwen ran on an M5 laptop with 24 GB of memory and took about 30 seconds per document, according to the post.

## Why it matters

The results suggest that a small model running on a laptop can handle some document tasks well, but performance can vary sharply with formats and document types.

## Sources & references

1. [Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]](https://www.reddit.com/r/MachineLearning/comments/1wsbqni/qwen3vl_8b_on_a_laptop_vs_opus_55_sonnet_5_gpt56/) – Reddit r/MachineLearning, 2026-09-28

Last updated: 2026-09-28
