datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.XBRLBenchMIRAGE
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
Ziqian Liu and Stephan Alaniz
Abstract: Instruction-guided image editing has seen remarkable progress with models like FLUX.2 and Qwen-Image-Edit, yet they still struggle with complex scenarios involving multiple similar instances, each requiring individual edits. We observe that state-of-the-art models suffer from severe over-editing and spatial misalignment when faced with multiple identical instances and composite… See the full description on the dataset page: https://huggingface.co/datasets/ziqiangoodgood/MIRAGE.mirage-engine-ledger
Mirage Engine Campaign Ledger
A hash-chained, append-only research journal: 29,707 JSONL records in which each
entry carries a SHA-256 entry_hash over its own body and a prev_hash linking
it to its predecessor. The chain is independently verifiable from the file alone.
Author: Christopher Betances — catqualia.com
License: CC BY 4.0 (see LICENSE)
Language: English (record text); structured JSON in meta
Records: 29,707
Time span: 2026-08-16 03:51:13 UTC → 2026-08-19 05:27:40 UTC… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/mirage-engine-ledger.mirage
MIRAGE — Do Web Agents Investigate Before They Decide?
Misleading Investigation Reveals Agent Gaps in Evidence.
MIRAGE is a benchmark for investigative competence in autonomous web
agents — the ability to recognise when visible information is insufficient,
seek hidden context, and integrate discovered evidence into a final
decision. The benchmark spans three structurally distinct moderation and
policy enforcement domains, each engineered around a two-layer information… See the full description on the dataset page: https://huggingface.co/datasets/SyedNazmusSakib/mirage.Youlln__ECE-MIRAGE-1-15B-details
Dataset Card for Evaluation run of Youlln/ECE-MIRAGE-1-15B
Dataset automatically created during the evaluation run of model Youlln/ECE-MIRAGE-1-15B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__ECE-MIRAGE-1-15B-details.MIRAGEmirage-campaigns
Mirage Engine campaign corpus
301 campaign runs of a differential-testing engine, each measuring how often a system
satisfies its declared specification while its real state diverges from what the
specification means.
The name misleads, so start here: this is not an attack corpus. It is a measurement corpus.
The engine's target is a specific and under-instrumented failure class — spec-compliant
divergence: a component that stays green on every declared check while its actual… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/mirage-campaigns.Youlln__ECE-MIRAGE-1-12B-details
Dataset Card for Evaluation run of Youlln/ECE-MIRAGE-1-12B
Dataset automatically created during the evaluation run of model Youlln/ECE-MIRAGE-1-12B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__ECE-MIRAGE-1-12B-details.test_datasetmirage_urtranslated_mirage
