SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.03047v1 Announce Type: new Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common benchmarks do not systematically test their bibliographic work. They need to know which methods work for their tasks and what those methods require to…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv cs.CL
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking