meituan-longcat/LoHoSearch
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling 📃 Paper • 🏆 Benchmark • 📦 Training Data Abstract Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predominantly human-authored, annotators lack a global perspective on entity statistics and cannot systematically maximize search space size and structural… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/LoHoSearch.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face