Your intent classifier is 12 points worse in Portuguese: benchmarking Laya, Strands Decider and Qwen3 embeddings
A conversational agent first has to decide what the customer wants. I measured that step for customer-service messages in Brazilian Portuguese (PT-BR), using two small "decider" models and two embedding models. Then I translated the whole dataset into English and ran everything again. TL;DR Decider…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-08 08:23 · DEV Community — Machine Learning
Your intent classifier is 12 points worse in Portuguese: benchmarking Laya, Strands Decider and Qwen3 embeddings