AINewsnow

Qwen3-VL 8B on a MacBook vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats

I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on: - receipts: CORD (Indonesia) and SROIE (Malaysia), 30 each - 20 scanned 1980s-90s invoices , answer keys human-verified - 32 real IRS forms, 4 damage levels (generated this…

Read the full story at r/computervision ↗

Timeline · 1 report

  1. 2026-09-28 11:17 · r/computervision
    Qwen3-VL 8B on a MacBook vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats

More stories

  1. Opus 5.5 — r/ClaudeAI
  2. Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents
  3. If you had to choose only one, which would you pick? — r/GeminiAI
  4. Can't use Gemini with a VPN? — r/GeminiAI
  5. I asked Claude Code to make it's own version of that guy's "time" video by Astra 5.6 from yesterday. — r/ChatGPT
  6. I was curious — r/OpenAI
  7. Why is Google Antigravity so underrated? — r/AI_Agents
  8. How is Opus 5.5 cheaper AND better than Fable 5.1? Genuinely trying to understand the "how" — r/ArtificialInteligence

Get the daily brief of stories like this at 6:30 every morning →