Dense200 scores for seven models, including one that writes its boxes as plain text
A lot of the counting posts here end up fighting crowded scenes, so these numbers caught my eye. It's from a paper on a model that writes detection boxes out as plain text, coordinates as tokens, with no box head. It's called SenseNova-Vision, and its starting point was Bagel (also in the chart). T…
Read the full story at r/computervision ↗
Timeline · 1 report
- 2026-10-07 16:42 · r/computervision
Dense200 scores for seven models, including one that writes its boxes as plain text