Teaching a 27B Model to Write Trading Alphas: 101 Formulas, 12 Rewards and One Unseen Year
Qwen3.8-27B is trained with multi-reward RL to write one-line trading formulas in the language of WorldQuant’s 101 Formulaic Alphas . A deterministic verifier turns each formula into a daily dollar-neutral long-short book on 46 US stocks and scores it on 12 reward channels. On 50 prompts and a year…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-10 19:37 · DEV Community — Machine Learning
Teaching a 27B Model to Write Trading Alphas: 101 Formulas, 12 Rewards and One Unseen Year