datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.RealworldQA
RealWorldQA
RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve.
The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.vlmsareblindArXiv - Website
X-AIGD
X-AIGD
X-AIGD is a fine-grained benchmark designed for eXplainable AI-Generated image Detection. It provides pixel-level human annotations of perceptual artifacts in AI-generated images, spanning low-level distortions, high-level semantics, and cognitive-level counterfactuals, aiming to advance robust and explainable AI-generated image detection methods.
For more details, please refer to our paper: Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable… See the full description on the dataset page: https://huggingface.co/datasets/Coxy7/X-AIGD.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-Behavior-Dataset.xAI_Aurora_t2i_human_preferences
Rapidata Aurora Preference
This T2I dataset contains over 400k human responses from over 86k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Aurora across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/xAI_Aurora_t2i_human_preferences.xai-studies
xai-studies
xai = explainable AI. Offline mirror of lyffseba/xai.
GitHub
lyffseba/xai
portal
spaces/lyffseba/xai
use
python3 studies/run.py test
No pip. No network.
catalog
catalog/models.csv
id
name
status
total
active
ctx
experts
inkling
Inkling
weights_public
975B
41B
1048576
6/256+2 shared
kimi-k3
Kimi K3
api_live_weights_pending
2.8T
1048576
16/896
laguna-s-2.1
Laguna S 2.1
weights_public
118B
8B
1048576… See the full description on the dataset page: https://huggingface.co/datasets/lyffseba/xai-studies.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.OpenImages-Inpaintedpriya-sft
Priya SFT + RAG dataset
Synthetic training and retrieval data for the Priya persona — a fictional Senior Customer Success Manager at a fictional B2B SaaS company. Used to train kader-xai/priya-qwen2.5-7b-lora and kader-xai/priya-qwen2.5-7b-gguf.
Fully synthetic. No real person, customer, or company. Generated as the seed corpus for Project Recall, an experiment in employee-continuity AI.
📝 Blog post: Employee Recall — Capturing a Departing Employee's Writing Style and… See the full description on the dataset page: https://huggingface.co/datasets/kader-xai/priya-sft.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.truevislies-resultsaym-xai-datasetFor citing:
@INPROCEEDINGS{11206864,
author={Erdoğanyılmaz, Cihan and Naç, Ali Yasir},
booktitle={2025 10th International Conference on Computer Science and Engineering (UBMK)},
title={Predicting Norm Control Decisions of the {Turkish} {Constitutional} {Court} Using {Explainable} {AI} Techniques},
year={2025},
pages={657-662},
abstract={The application of Natural Language Processing (NLP) to Legal Judgment Prediction (LJP) has gained significant momentum, yet most research in the… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/aym-xai-dataset.xai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.care-xai
CARE-XAI: Culturally-Aware, Evidence-Grounded Explainable AI for Health
17,803 total rows · 14,254 train / 1,797 validation / 1,752 test · 5 sources · 3 labels
CARE-XAI is a unified health-claim verification dataset consolidating five public health NLP benchmarks into a single schema, augmented with GRADE evidence quality labels, cultural relevance flags, and Gold/Silver explanation annotations.
Sources
Dataset
Rows
%
Explanation
License
PUBHEALTH
9,804
55.1%… See the full description on the dataset page: https://huggingface.co/datasets/Prabhjotschugh/care-xai.B-XAICpiezo-embedding-benchmark
Piezometric Embedding Benchmark Dataset
Daily groundwater level time series from ~4200 French monitoring stations,
with ERA5 climate covariates and hydrogeological labels.
Notebooks
Notebook
Description
01_data_exploration.ipynb
Dataset overview, label distributions, geographic maps, time series examples
02_benchmark_analysis.ipynb
Encoder comparison, whitening effect, uni vs multi, ranking
Dataset Description
This dataset supports the… See the full description on the dataset page: https://huggingface.co/datasets/xairon/piezo-embedding-benchmark.X-AIGD-demoThis is a tiny demo subset of X-AIGD for debugging purposes only. Please visit the dataset card of X-AIGD for descriptions.
XAIGID-RewardBench
XAIGID-RewardBench
The test split contains 3,988 benchmark triplets with an image, two policy-model responses, and up to three human-response slots. It contains 4,412 populated human annotations. The human_response split contains the 336-image Human-vs-Policy subset, including 36 GPT-5.5 comparisons, in the same image-and-response schema.
See the project repository and paper.
Xaixai__grok-3_eval_5ed6
xai__grok-3 Evaluation Results
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME25
LiveCodeBenchv5
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
HLE
AIME24
Accuracy
50.0
28.3
90.5
85.0
33.2
88.5
66.5
32.7
29.8
7.3
59.3
AIME25
Average Accuracy: 50.0% ± 1.8%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
50.0%
15
30
2
53.3%
16
30
3
56.7%
17
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/xai__grok-3_eval_5ed6.xai-attack-detection-imagenette
XAI Attack Detection: Imagenette targeted BIM/PGD on ViT-B/16
Private research dataset of paired clean and targeted adversarial Imagenette images. It is
built to study how adversarial attacks change a Vision Transformer's explanation maps and to
support later work on attack detection. Each row is one source image with its clean and its
attacked version.
Summary
Pairs
12,420 (train 8,690 · validation 1,860 · test 1,870)
Source images
Imagenette v2… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-imagenette.XAI_Malware_PredictionexplainDepression-social-media-xaiCleaned, balanced, and clinically annotated social media dataset
for explainable depression detection research.
Combined from real Twitter and Reddit posts, engineered with
DSM-5-aligned clinical lexicon features, and used to train a
DistilBERT model achieving 96.17% accuracy.
━━━━━━━━━━━━━━━━━━━━━━━━━━━
DATASET STATS
Total Rows → 40,770
Class Balance → 50% Depressed / 50% Not Depressed
Feature Columns → 12
Sources → Twitter + Reddit
━━━━━━━━━━━━━━━━━━━━━━━━━━━… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/explainDepression-social-media-xai.philosophai-xai-grok-3
Dataset Card for "philosophai-xai-grok-3"
More Information needed
new_conv_xai_augmented
Dataset Description
This dataset is derived from new_conv_xai. new_conv_xai had the integration of all the images and other metadata together with each of the 30 conversations from the train and test json files.
new_conv_xai_augmented takes it a step further and converts each of those conversations into subsets of conversation histories, creating inputs and outputs that we can finetune the model on.
Each of these input-output pairs are associated with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/mxforml/new_conv_xai_augmented.tradehax-xai-grok-image-capabilities
TradeHax xAI/Grok Image Capabilities
Capability-alignment prompts and responses for xAI/Grok-inspired visual generation behavior.
Owner: Antired
Repo: tradehax-xai-grok-image-capabilities
Synced by: scripts/sync-hf-assets.js
tradehax-xai-grok-trading-visual-prompts
TradeHax xAI/Grok Trading Visual Prompts
Curated trading-scene image prompts and negative prompts tuned for xAI/Grok-inspired visual style.
Owner: Antired
Repo: tradehax-xai-grok-trading-visual-prompts
Synced by: scripts/sync-hf-assets.js
UFB_xainew_conv_xai
