One Extra Letter, Four Times the Tokens
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
Tokenization isn't preprocessing. It's the meter on your bill and the unit your retrieval is measured in. BPE started as a compression algorithm. OpenAI adopted it for GPT, and GPT-2 pushed the base unit to raw bytes — roughly 50,257 tokens — which is why the same design handles code, emoji, and ev…
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-09-02 11:19 · r/deeplearning
One Extra Letter, Four Times the Tokens