I’ve frozen UrduEval v0.3.0 — now I’m testing whether the metrics actually agree with humans
I’ve been working on UrduEval , an open-source evaluation framework for Urdu and Roman Urdu LLMs. After several iterations, I’ve reached a point where I’ve deliberately stopped changing the evaluator. The reason is simple: I don’t want to keep tweaking the metric until I get the result I expect. Wh…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-25 07:08 · r/LocalLLM
I’ve frozen UrduEval v0.3.0 — now I’m testing whether the metrics actually agree with humans