CoolFace
Datasetpublic

mast-benchmark/100k-corpus-2026

MAST 100K Corpus 2026 This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers. This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes489downloads
2 commits on main
8c0eab12mo ago

Update README for MAST 100K corpus

nthakur
3f0e3812mo ago

Duplicate from Tevatron/browsecomp-plus-corpus

nthakur, s42chen