AINewsnow

Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]

I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on: - receipts: CORD (Indonesia) and SROIE (Malaysia), 30 each - 20 scanned 1980s-90s invoices , answer keys human-verified - 32 real IRS forms, 4 damage levels (generated this…

Read the full story at r/MachineLearning ↗

Timeline · 2 reports

  1. 2026-09-28 11:17 · r/computervision
    Qwen3-VL 8B on a MacBook vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats
  2. 2026-09-28 11:11 · r/MachineLearning
    Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]

More stories

  1. Opus 5.5 — r/ClaudeAI
  2. Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents
  3. Meta Muse is the worst AI agent for privacy – and I’ve tried them all — ZDNET AI
  4. If you had to choose only one, which would you pick? — r/GeminiAI
  5. Can't use Gemini with a VPN? — r/GeminiAI
  6. I asked Claude Code to make it's own version of that guy's "time" video by Astra 5.6 from yesterday. — r/ChatGPT
  7. I was curious — r/OpenAI
  8. Free to the first 100: a Windows app that runs one prompt past three models in assigned roles and keeps the disagreement — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →