datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.data-govt-nz-mirror
data.govt.nz — Mirror Catalogue (hourly snapshot)
Mirror publication of the New Zealand open government data catalogue (data.govt.nz).
Each row of catalog.csv is a dataset record as harvested from the data.govt.nz
CKAN instance (national agencies and local councils).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-govt-nz-mirror.data-gov-au-mirror
data.gov.au — Mirror Catalogue (hourly snapshot)
Mirror publication of the Australian open government data catalogue (data.gov.au).
Each row of catalog.csv is a dataset record as harvested from the data.gov.au CKAN
instance (federal, state and territory agencies).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in the… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-gov-au-mirror.Mirror-Prompt-Injection-Dataset
Mirror Prompt Injection Dataset
A ~5,000-pair prompt injection detection dataset built using the Mirror design pattern, as described in:
The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detectionhttps://arxiv.org/abs/2603.11875
Key results from the paper
The paper demonstrates that a sparse character n-gram linear SVM trained on 5,000 Mirror-curated samples achieves 95.97% recall and 92.07% F1 on a holdout set, with sub-millisecond… See the full description on the dataset page: https://huggingface.co/datasets/watchdogsrox/Mirror-Prompt-Injection-Dataset.theogonos-mirror-test
Theogonos Mirror Test
A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position.
Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness.
Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.ibm-aml-mirrormirror
MIRROR Dataset
MIRROR is a synthetic vision–language dataset for multimodal cognitive reframing under client resistance.
Paper: 🪞 MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance
The dataset includes:
Client profile metadata (CACTUS idx, CelebA idx)
Dialogue written in a screenplay format, including stage directions that describe facial expressions
⚠️ Images themselves are not included to comply with the CelebA license.
However, we provide the full image… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reframing/mirror.mirror-cncf-question-and-answer-dataset-for-llm-training
CNCF QA Dataset for LLM Tuning
Description
This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-cncf-question-and-answer-dataset-for-llm-training.mirror-iac-eval
IaC-Eval dataset (v1.1)
IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities.
This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now).
| Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper |
2. Usage instructions
Option 1: Running the… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-iac-eval.Dual_mirror
