Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.05766v1 Announce Type: new Abstract: The dominant approach to building domain-specific pretraining corpora is to filter large web archives such as CommonCrawl. This works well for popular domains but breaks down for specialized ones, where relevant content is sparse and often beyond the…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-09 04:00 · arXiv cs.LG
Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora