An 8B model given structured context matched a 14B given prose on cross-document temporal reasoning — and with plain retrieval, both scored zero
This story is from 2026-08-30. It is preserved in the archive; the latest stories are on the live feed.
I tested whether structure in the context window can substitute for parameters. Qwen3, five sizes, 0.6B to 14B, so size varies and architecture doesn't. The task: 38 questions asking whether event A precedes event B, where A and B are narrated in different documents in a five-document corpus (260,2…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-08-30 17:51 · r/learnmachinelearning
An 8B model given structured context matched a 14B given prose on cross-document temporal reasoning — and with plain retrieval, both scored zero