Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.
I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-24 10:31 · r/LocalLLM
Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. - 2026-09-24 10:22 · r/LocalLLaMA
Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.