What Does the Encoder Actually Decide? A Controlled Comparison of Vision Backbones on Joint Tree Segmentation and Stereo Depth
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.13232v1 Announce Type: new Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a vision backbone chosen by reputation rather than measurement. Holding dataset, decoders, losses, sc…
Read the full story at arXiv cs.CV ↗