CoolFace
Datasetpublic

SyedNazmusSakib/mirage

MIRAGE — Do Web Agents Investigate Before They Decide? Misleading Investigation Reveals Agent Gaps in Evidence. MIRAGE is a benchmark for investigative competence in autonomous web agents — the ability to recognise when visible information is insufficient, seek hidden context, and integrate discovered evidence into a final decision. The benchmark spans three structurally distinct moderation and policy enforcement domains, each engineered around a two-layer information… See the full description on the dataset page: https://huggingface.co/datasets/SyedNazmusSakib/mirage.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes28downloads
2 commits on main
77b5aa85mo ago

Initial release: MIRAGE benchmark with 3 configs (wikipedia, shopping, reddit)

SyedNazmusSakib
67ac3a95mo ago

initial commit

SyedNazmusSakib