datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/DannHiroaki/China-Building-Footprints-CMAB-Mirror.China-Building-Footprints-CMAB-Mirror
Origin Data
@misc{Zhang2025CMAB,
author = {Zhang, Yecheng and Zhao, Huimin and Long, Ying},
title = {{CMAB-The World's First National-Scale Multi-Attribute Building Dataset}},
year = {2025},
month = apr,
publisher = {figshare},
doi = {10.6084/m9.figshare.27992417},
url = {https://doi.org/10.6084/m9.figshare.27992417},
howpublished = {dataset}
}
Paper
@article{Zhang2025SciData,
author = {Zhang, Y. and… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/China-Building-Footprints-CMAB-Mirror.mirandese_g2pMiraData
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
Xuan Ju1*, Yiming Gao1*, Zhaoyang Zhang1*#, Ziyang Yuan1, Xintao Wang1, Ailing Zeng, Yu Xiong, Qiang Xu, Ying Shan1
1ARC Lab, Tencent PCG 2The Chinese University of Hong Kong *Equal Contribution #Project Lead
Introduction
Video datasets play a crucial role in video generation such as Sora.
However, existing text-video datasets often fall short when it comes to handling long video… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/MiraData.NSL-KDD
NSL-KDD
The data set is a data set that converts the arff File provided by the link into CSV and results.
The data set is personally stored by converting data to float64.
If you want to obtain additional original files, they are organized in the Original Directory in the repo.
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
duration
151165
non-null
int64
1
protocol_type
151165
non-null
object
2
service
151165
non-null… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/NSL-KDD.muchomusic
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional… See the full description on the dataset page: https://huggingface.co/datasets/mulab-mir/muchomusic.data-govt-nz-mirror
data.govt.nz — Mirror Catalogue (hourly snapshot)
Mirror publication of the New Zealand open government data catalogue (data.govt.nz).
Each row of catalog.csv is a dataset record as harvested from the data.govt.nz
CKAN instance (national agencies and local councils).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-govt-nz-mirror.UNSW-NB15
UNSW-NB15
This data is provided through the Train, Test CSV file provided by UNSW-NB15.
link
Labels
The label of the data set is as follows.
#
Column
Non-Null
Count
Dtype
0
id
82332
non-null
int64
1
dur
82332
non-null
float64
2
proto
82332
non-null
object
3
service
82332
non-null
object
4
state
82332
non-null
object
5
spkts
82332
non-null
int64
6
dpkts
82332
non-null
int64
7
sbytes
82332
non-null
int64
8
dbytes
82332
non-null
int64
9
rate… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-NB15.ClinicalTrial-gov_QAfunctioncaliingmawi-https-flows-2025data-gov-au-mirror
data.gov.au — Mirror Catalogue (hourly snapshot)
Mirror publication of the Australian open government data catalogue (data.gov.au).
Each row of catalog.csv is a dataset record as harvested from the data.gov.au CKAN
instance (federal, state and territory agencies).
License declaration
This mirror catalogue is published under the Creative Commons Attribution 4.0
International (CC BY 4.0) licence. Individual dataset records reference their own
source licence in the… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-gov-au-mirror.telco-customer-churn
Dataset Card for Telco Customer Churn
This dataset contains information about customers of a fictional telecommunications company, including demographic information, services subscribed to, location details, and churn behavior. This merged dataset combines the information from the original Telco Customer Churn dataset with additional details.
Dataset Details
Dataset Description
This merged Telco Customer Churn dataset provides a comprehensive view of customer… See the full description on the dataset page: https://huggingface.co/datasets/mirunavasile/telco-customer-churn.miRBench
miRBench, curated
miRBench, as published in quality-curated genomic benchmarks - one format, fixed row order, a permanent ID on every row. 3 datasets, 6 files, 2,849,872 rows, one gzipped CSV per split, all Homo sapiens.
Getting the data
Two packages are the way in: genomic-benchmarks-data for people, genomic-benchmarks-data4agents for agents, the same functions either way. They resolve the URL, check the checksum, and carry each dataset's QC results, which this… See the full description on the dataset page: https://huggingface.co/datasets/genomic-benchmarks/miRBench.Mirror-Prompt-Injection-Dataset
Mirror Prompt Injection Dataset
A ~5,000-pair prompt injection detection dataset built using the Mirror design pattern, as described in:
The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detectionhttps://arxiv.org/abs/2603.11875
Key results from the paper
The paper demonstrates that a sparse character n-gram linear SVM trained on 5,000 Mirror-curated samples achieves 95.97% recall and 92.07% F1 on a holdout set, with sub-millisecond… See the full description on the dataset page: https://huggingface.co/datasets/watchdogsrox/Mirror-Prompt-Injection-Dataset.Science_Articlesmovies_reviews_summariesDataset consists of imdb urdu reviews with summaries.
ur_news_sumek100-mir-demo-assetsUNSW-IoT
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Mireu-Lab/UNSW-IoT.function-calling-dataset-Hindi-englishmiriad-1kibm-aml-mirrormir-hf2820-intake
Model Intake Queue
Third-party model submissions awaiting compliance review.
See intake_queue.csv for the pending requests.
FLeW-datamir-hf2816-intake
Model Intake Queue
Third-party model submissions awaiting compliance review.
See intake_queue.csv for the pending requests.
mir-hf2200-intakeMira_DPOscientific_papers_esLongcu
