Pack Four Docs Without the Block Mask and Doc 4 Spends 83.78% of Its Attention on Documents It Never Saw
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
Sequence packing concatenates short documents into one row so you stop paying for padding. On the length distribution here that lifts utilisation from 27.98% to 99.81%. But attention inside a row is global . Without a block-diagonal mask, the fourth document in the row spends 83.78% of its layer-1…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-17 12:33 · DEV Community — Machine Learning
Pack Four Docs Without the Block Mask and Doc 4 Spends 83.78% of Its Attention on Documents It Never Saw