datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MIRAGE-CanaryDocs
MIRAGE CanaryDocs
MIRAGE CanaryDocs is an English synthetic enterprise-document dataset for structured privacy-unit,
canary, and ordered multi-chunk evaluation. It is the companion dataset for the EMNLP 2026 paper
When Metadata Remembers: Ordered Provenance Enables Document-Level Embedding Inversion.
Project documentation and schemas are also available in the
MIRAGE GitHub repository.
Dataset summary
The dataset contains complete synthetic documents, ordered token… See the full description on the dataset page: https://huggingface.co/datasets/LevenKoko/MIRAGE-CanaryDocs.Youlln__ECE-MIRAGE-1-15B-details
Dataset Card for Evaluation run of Youlln/ECE-MIRAGE-1-15B
Dataset automatically created during the evaluation run of model Youlln/ECE-MIRAGE-1-15B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__ECE-MIRAGE-1-15B-details.MIRAGEmirage-campaigns
Mirage Engine campaign corpus
301 campaign runs of a differential-testing engine, each measuring how often a system
satisfies its declared specification while its real state diverges from what the
specification means.
The name misleads, so start here: this is not an attack corpus. It is a measurement corpus.
The engine's target is a specific and under-instrumented failure class — spec-compliant
divergence: a component that stays green on every declared check while its actual… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/mirage-campaigns.Youlln__ECE-MIRAGE-1-12B-details
Dataset Card for Evaluation run of Youlln/ECE-MIRAGE-1-12B
Dataset automatically created during the evaluation run of model Youlln/ECE-MIRAGE-1-12B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Youlln__ECE-MIRAGE-1-12B-details.
