Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.01999v1 Announce Type: new Abstract: We study a variant of the Thompson Sampling (TS) algorithm, called $\alpha$-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We fo…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-03 04:00 · arXiv stat.ML
Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling