BraintrustDataDev/livenewsbench-search-arms
LiveNewsBench Search Arms Paired measurements of four language models answering the same 1,329 news questions under four retrieval conditions. Every question was run in every condition, so each row pairs with 13 others on task_key. The release answers one question: how much of an agent's answer quality comes from the model, and how much from the search system wrapped around it. This is a derivative evaluation-results dataset, not the original LiveNewsBench benchmark. The… See the full description on the dataset page: https://huggingface.co/datasets/BraintrustDataDev/livenewsbench-search-arms.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face