Help diagnosing 100% accuracy (Data Leakage) on DeBERTa & 0% (Label Flip) on a Portuguese DeBERTa model
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
I’m classifying scientific abstracts written in Portuguese into two temporal categories: "Old" vs "Recent". I tested several models, but two of them are giving me massive red flags: DeBERTa (base): Getting exactly 100% accuracy on the test set. Albertina (a Portuguese DeBERTa-based model): Getting…
Read the full story at r/LanguageTechnology ↗
Timeline · 2 reports
- 2026-09-08 00:25 · r/learnmachinelearning
Help diagnosing 100% accuracy (Data Leakage) on DeBERTa & 0% (Label Flip) on a Portuguese DeBERTa model - 2026-09-08 00:24 · r/LanguageTechnology
Help diagnosing 100% accuracy (Data Leakage) on DeBERTa & 0% (Label Flip) on a Portuguese DeBERTa model