Modernising the Compressed-Domain Video Captioner: A Controlled Study of SigLIP2 and GPT-2 Substitutions
arXiv:2609.31700v1 Announce Type: new Abstract: Compressed-domain video captioning avoids full video decoding by operating directly on I-frames, motion vectors and residuals, trading a small amount of accuracy for a large gain in inference speed. CoCap established this pipeline using a CLIP vision…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.CV
Modernising the Compressed-Domain Video Captioner: A Controlled Study of SigLIP2 and GPT-2 Substitutions