datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.truevislies-resultsaym-xai-datasetFor citing:
@INPROCEEDINGS{11206864,
author={Erdoğanyılmaz, Cihan and Naç, Ali Yasir},
booktitle={2025 10th International Conference on Computer Science and Engineering (UBMK)},
title={Predicting Norm Control Decisions of the {Turkish} {Constitutional} {Court} Using {Explainable} {AI} Techniques},
year={2025},
pages={657-662},
abstract={The application of Natural Language Processing (NLP) to Legal Judgment Prediction (LJP) has gained significant momentum, yet most research in the… See the full description on the dataset page: https://huggingface.co/datasets/LawChatAI/aym-xai-dataset.xai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.credit-default-taiwan
Default of Credit Card Clients (Taiwan) — xaitalk example-data mirror
Mirror of the UCI Default of Credit Card Clients dataset, hosted as a reliable runtime fallback for xaitalk's TreeSHAP example. 30,000 clients x 23 features (credit limit, age, repayment history PAY_*, bill/payment amounts), binary target = default next month.
Source: UCI Machine Learning Repository (https://archive.ics.uci.edu/dataset/350/default+of+credit+card+clients). Credit to the original creator… See the full description on the dataset page: https://huggingface.co/datasets/xaitalk/credit-default-taiwan.B-XAICpiezo-embedding-benchmark
Piezometric Embedding Benchmark Dataset
Daily groundwater level time series from ~4200 French monitoring stations,
with ERA5 climate covariates and hydrogeological labels.
Notebooks
Notebook
Description
01_data_exploration.ipynb
Dataset overview, label distributions, geographic maps, time series examples
02_benchmark_analysis.ipynb
Encoder comparison, whitening effect, uni vs multi, ranking
Dataset Description
This dataset supports the… See the full description on the dataset page: https://huggingface.co/datasets/xairon/piezo-embedding-benchmark.xai__grok-3_eval_5ed6
xai__grok-3 Evaluation Results
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME25
LiveCodeBenchv5
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
HLE
AIME24
Accuracy
50.0
28.3
90.5
85.0
33.2
88.5
66.5
32.7
29.8
7.3
59.3
AIME25
Average Accuracy: 50.0% ± 1.8%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
50.0%
15
30
2
53.3%
16
30
3
56.7%
17
30
4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/xai__grok-3_eval_5ed6.XAI_Malware_Predictionheart-disease-cleveland
Heart Disease (Cleveland) — xaitalk example-data mirror
Mirror of the UCI Heart Disease (Cleveland) dataset, hosted as a reliable runtime fallback for xaitalk's tabular MLP example. 303 patients x 13 features, binary target (disease presence).
Source: UCI Machine Learning Repository (https://archive.ics.uci.edu/dataset/45/heart+disease). All credit to the original creators (Hungarian Inst. of Cardiology / Cleveland Clinic et al.). Redistributed unmodified for example… See the full description on the dataset page: https://huggingface.co/datasets/xaitalk/heart-disease-cleveland.philosophai-xai-grok-3
Dataset Card for "philosophai-xai-grok-3"
More Information needed
explainDepression-social-media-xaiCleaned, balanced, and clinically annotated social media dataset
for explainable depression detection research.
Combined from real Twitter and Reddit posts, engineered with
DSM-5-aligned clinical lexicon features, and used to train a
DistilBERT model achieving 96.17% accuracy.
━━━━━━━━━━━━━━━━━━━━━━━━━━━
DATASET STATS
Total Rows → 40,770
Class Balance → 50% Depressed / 50% Not Depressed
Feature Columns → 12
Sources → Twitter + Reddit
━━━━━━━━━━━━━━━━━━━━━━━━━━━… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/explainDepression-social-media-xai.tradehax-xai-grok-trading-visual-prompts
TradeHax xAI/Grok Trading Visual Prompts
Curated trading-scene image prompts and negative prompts tuned for xAI/Grok-inspired visual style.
Owner: Antired
Repo: tradehax-xai-grok-trading-visual-prompts
Synced by: scripts/sync-hf-assets.js
xai_gab_multip_robertaxai_hd_newClinicalNotes_labeledxai_hd_multipRAB-Cred
RAB-Cred
RAB-Cred is a text classification dataset, where the task is to identify the presence and sentiment of credibility assessments in Danish asylum decision texts. The three classes are:
No credibility assessment: ABSENT
Positive credibility assessment: POSITIVE
Negative credibility assessment: NEGATIVE
The RAB-Cred dataset features high-quality, gold-standard expert annotations and valuable metadata such as annotator confidence and asylum case outcome. Decisions texts were… See the full description on the dataset page: https://huggingface.co/datasets/XAI-CRED/RAB-Cred.xai_epic_baseline_maj_robxai_hdxai_rob_baseline_llm
