Qwen3-VL 8B on a MacBook vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats
I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on: - receipts: CORD (Indonesia) and SROIE (Malaysia), 30 each - 20 scanned 1980s-90s invoices , answer keys human-verified - 32 real IRS forms, 4 damage levels (generated this…
Read the full story at r/computervision ↗
Timeline · 1 report
- 2026-09-28 11:17 · r/computervision
Qwen3-VL 8B on a MacBook vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats
More stories
- Opus 5.5 — r/ClaudeAI
- Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents
- If you had to choose only one, which would you pick? — r/GeminiAI
- Can't use Gemini with a VPN? — r/GeminiAI
- I asked Claude Code to make it's own version of that guy's "time" video by Astra 5.6 from yesterday. — r/ChatGPT
- I was curious — r/OpenAI
- Why is Google Antigravity so underrated? — r/AI_Agents
- How is Opus 5.5 cheaper AND better than Fable 5.1? Genuinely trying to understand the "how" — r/ArtificialInteligence
Get the daily brief of stories like this at 6:30 every morning →