Note · 1 min read
Qwen3-VL 8B tested on 137 messy documents using a laptop
A Reddit user compared a small vision model running locally with three hosted models on receipts, forms, invoices and contracts. In the user’s test, Qwen did well on tax forms but struggled with Indian date formats and long contracts.
Oossa · About Oossa
A Reddit user tested Qwen3-VL 8B Instruct on an M5 laptop against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra, using 137 documents. The user reports that documents were fully correct in 59% of Qwen’s results, compared with 89% for Opus and 57% for GPT-5.6 Terra. Qwen got 21 of 32 IRS forms fully right, while GPT-5.6 Terra got seven. But on 10 synthetic Indian bank statements, Qwen read dates in day-month-year format as month-day-year, despite getting the amounts and balances right.
Why it matters
The results suggest that a small model running on a laptop can handle some document tasks well, but performance can vary sharply with formats and document types.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R] ↗ | Reddit r/MachineLearning | Sep 28, 2026 | I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on: - receipts: C |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.