datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.RealworldQA
RealWorldQA
RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve.
The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.vlmsareblindArXiv - Website
X-AIGD
X-AIGD
X-AIGD is a fine-grained benchmark designed for eXplainable AI-Generated image Detection. It provides pixel-level human annotations of perceptual artifacts in AI-generated images, spanning low-level distortions, high-level semantics, and cognitive-level counterfactuals, aiming to advance robust and explainable AI-generated image detection methods.
For more details, please refer to our paper: Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable… See the full description on the dataset page: https://huggingface.co/datasets/Coxy7/X-AIGD.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-Behavior-Dataset.xAI_Aurora_t2i_human_preferences
Rapidata Aurora Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Aurora across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/xAI_Aurora_t2i_human_preferences.xai-studies
xai-studies
xai = explainable AI. Offline mirror of lyffseba/xai.
GitHub
lyffseba/xai
portal
spaces/lyffseba/xai
use
python3 studies/run.py test
No pip. No network.
catalog
catalog/models.csv
id
name
status
total
active
ctx
experts
inkling
Inkling
weights_public
975B
41B
1048576
6/256+2 shared
kimi-k3
Kimi K3
api_live_weights_pending
2.8T
1048576
16/896
laguna-s-2.1
Laguna S 2.1
weights_public
118B
8B
1048576… See the full description on the dataset page: https://huggingface.co/datasets/lyffseba/xai-studies.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.X-Atlas-Pisces
X-Atlas/Pisces
X-Atlas/Pisces is a Perturb-seq atlas containing 25.6 million perturbed single-cell transcriptomes across 16 biologically diverse contexts, including widely used cell lines (HCT116, HEK293T, HepG2),
induced pluripotent stem cells (iPSCs), resting and CD3/CD28 activated Jurkat T lymphoma cells, and multi-lineage differentiating iPSCs.
(Coming Soon) The following data will be uploaded to this dataset:
All cells from X-Atlas/Orion (cells with a valid dual-guide pair… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Pisces.OpenImages-Inpaintedpriya-sft
Priya SFT + RAG dataset
Synthetic training and retrieval data for the Priya persona — a fictional Senior Customer Success Manager at a fictional B2B SaaS company. Used to train kader-xai/priya-qwen2.5-7b-lora and kader-xai/priya-qwen2.5-7b-gguf.
Fully synthetic. No real person, customer, or company. Generated as the seed corpus for Project Recall, an experiment in employee-continuity AI.
📝 Blog post: Employee Recall — Capturing a Departing Employee's Writing Style and… See the full description on the dataset page: https://huggingface.co/datasets/kader-xai/priya-sft.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.truevislies-resultsaym-xai-datasetFor citing:
@INPROCEEDINGS{11206864,
author={Erdoğanyılmaz, Cihan and Naç, Ali Yasir},
booktitle={2025 10th International Conference on Computer Science and Engineering (UBMK)},
title={Predicting Norm Control Decisions of the {Turkish} {Constitutional} {Court} Using {Explainable} {AI} Techniques},
year={2025},
pages={657-662},
abstract={The application of Natural Language Processing (NLP) to Legal Judgment Prediction (LJP) has gained significant momentum, yet most research in the… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/aym-xai-dataset.taxia-korean-tax-lawsxai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.care-xai
CARE-XAI: Culturally-Aware, Evidence-Grounded Explainable AI for Health
17,803 total rows · 14,254 train / 1,797 validation / 1,752 test · 5 sources · 3 labels
CARE-XAI is a unified health-claim verification dataset consolidating five public health NLP benchmarks into a single schema, augmented with GRADE evidence quality labels, cultural relevance flags, and Gold/Silver explanation annotations.
Sources
Dataset
Rows
%
Explanation
License
PUBHEALTH
9,804
55.1%… See the full description on the dataset page: https://huggingface.co/datasets/Prabhjotschugh/care-xai.credit-default-taiwan
Default of Credit Card Clients (Taiwan) — xaitalk example-data mirror
Mirror of the UCI Default of Credit Card Clients dataset, hosted as a reliable runtime fallback for xaitalk's TreeSHAP example. 30,000 clients x 23 features (credit limit, age, repayment history PAY_*, bill/payment amounts), binary target = default next month.
Source: UCI Machine Learning Repository (https://archive.ics.uci.edu/dataset/350/default+of+credit+card+clients). Credit to the original creator… See the full description on the dataset page: https://huggingface.co/datasets/xaitalk/credit-default-taiwan.B-XAICXAI
Dataset
For the present study, we used data from the GxE competition advocated by the G2F project in 2022 (https://www.maizegxeprediction2022.org/), including genetic markers (G2F-G) for maize inbred lines, phenotypic measurements (G2F-P) collected throughout each growing season, metadata (G2F-M) for each field trial, environmental covariate (EC) data, and environmental (G2F-E) data. G2F-E data were mainly climatic and soil variables captured during crop development in each… See the full description on the dataset page: https://huggingface.co/datasets/AIBreeding/XAI.X_AILabxai-personas-resultspiezo-embedding-benchmark
Piezometric Embedding Benchmark Dataset
Daily groundwater level time series from ~4200 French monitoring stations,
with ERA5 climate covariates and hydrogeological labels.
Notebooks
Notebook
Description
01_data_exploration.ipynb
Dataset overview, label distributions, geographic maps, time series examples
02_benchmark_analysis.ipynb
Encoder comparison, whitening effect, uni vs multi, ranking
Dataset Description
This dataset supports the… See the full description on the dataset page: https://huggingface.co/datasets/xairon/piezo-embedding-benchmark.Cabin-multi-modal-recognition-dataset-of-human
XAILab-CyberSpark/Cabin-multi-modal-recognition-dataset-of-human
由XAI Lab 汽车智能座舱数据集生成引擎生成的座舱垂域任务数据集,面向智能座舱舱内识人场景,本次开源demo数据集包含719张图像及标签,涵盖不同性别、年龄、情绪(表情)、行为动作、衣物等的标签,可用于舱内多模态感知模型SFT训练。
The cockpit vertical domain task dataset generated by the XAI Lab vehicle intelligent cockpit dataset generation engine is oriented to the intelligent cockpit human recognition scene. The Open Source demo dataset contains 719 images and tags, labels covering different genders, ages, emotions… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-multi-modal-recognition-dataset-of-human.X-AIGD-demoThis is a tiny demo subset of X-AIGD for debugging purposes only. Please visit the dataset card of X-AIGD for descriptions.
XAIGID-RewardBench
XAIGID-RewardBench
The test split contains 3,988 benchmark triplets with an image, two policy-model responses, and up to three human-response slots. It contains 4,412 populated human annotations. The human_response split contains the 336-image Human-vs-Policy subset, including 36 GPT-5.5 comparisons, in the same image-and-response schema.
See the project repository and paper.
Xaixai__grok-3_eval_5ed6
xai__grok-3 Evaluation Results
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME25
LiveCodeBenchv5
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
HLE
AIME24
Accuracy
50.0
28.3
90.5
85.0
33.2
88.5
66.5
32.7
29.8
7.3
59.3
AIME25
Average Accuracy: 50.0% ± 1.8%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
50.0%
15
30
2
53.3%
16
30
3
56.7%
17
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/xai__grok-3_eval_5ed6.soliaudit-dasp-sequence-all-xaiXAI_Malware_Prediction
