I Distilled a 568M Multilingual Model Into a 37M Japanese-English Encoder — Here's What Survived
Cross-lingual retrieval distillation, compression, and the honest cost table. Scope note: English-to-Japanese retrieval only. EN-JA eval is synthetic (opus-100 pairs), not a standard benchmark. All models compared on the same fixed subsample. The problem Tokyo companies with global customers need E…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-10 14:30 · DEV Community — Machine Learning
I Distilled a 568M Multilingual Model Into a 37M Japanese-English Encoder — Here's What Survived