Which model, which harness? I have data for you.
I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL It measure capabilities (a % of sucess on the various tasks) and…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-05 10:43 · r/LocalLLaMA
Which model, which harness? I have data for you.