datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.ptb-xl-processedHR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.DAAD-XAbout Dataset
The DAADX Dataset is derived from DAAD dataset (https://cvit.iiit.ac.in/research/projects/cvit-projects/daad#dataset), which contains all the captured videos for the Driver Intention Prediction task. We are introducing the first video based explanations dataset for driver intention prediction task. This will be help in further the research interms of making an explainable Autonomous Driving or ADAS System.
DAAD-X contains explanations for each maneuver instance, these… See the full description on the dataset page: https://huggingface.co/datasets/Skyrmion/DAAD-X.Objaverse-XL-Rigged-Animated
Objaverse-XL Rigged & Animated Subset
Every asset here carries both a skeleton and at least one animation clip, selected from
Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and
span characters as well as articulated rigid objects.
Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and
motion on it. This subset isolates that fraction: every file was checked to contain at least one
skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.Frappe_x1
Frappe_x1
Dataset description:
The Frappe dataset contains a context-aware app usage log, which comprises 96203 entries by 957 users for 4082 apps used in various contexts. It has 10 feature fields including user_id, item_id, daytime, weekday, isweekend, homework, cost, weather, country, city. The target value indicates whether the user has used the app under the context. Following the AFN work, we randomly split the data into 7:2:1 as the training set, validation set, and test set… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Frappe_x1.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.DNA_Gen
Citation
Please cite our work using the bibtex below:
BibTeX:
@article{su2025language,
title={Language Models for Controllable DNA Sequence Design},
author={Su, Xingyu and Li, Xiner and Lin, Yuchao and Xie, Ziqian and Zhi, Degui and Ji, Shuiwang},
journal={arXiv preprint arXiv:2507.19523},
year={2025}
}
xauusd
Cleaned XAUUSD Dataset
Dataset Description
This dataset contains cleaned and preprocessed minute-level historical price data for the XAU/USD (Gold vs. US Dollar) pair. The data spans from November 1, 2011, to January 3, 2024, and includes the following columns:
open: The opening price of the minute.
high: The highest price during the minute.
low: The lowest price during the minute.
close: The closing price of the minute.
tickvol: The number of price changes (ticks)… See the full description on the dataset page: https://huggingface.co/datasets/Pcitycrypto/xauusd.linkedin-job-postingsxlwic_wn
Multilingual Word-in-Context (WordNet)
Refer to the documentation and paper for more information.
Criteo_x1
Criteo_x1
Dataset description:
The Criteo dataset is a widely-used benchmark dataset for CTR prediction, which contains about one week of click-through data for display advertising. It has 13 numerical feature fields and 26 categorical feature fields. Following the AFN work, we randomly split the data into 7:2:1* as the training set, validation set, and test set, respectively.
The dataset statistics are summarized as follows:
Dataset Split
Total
#Train
#Validation
#Test… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Criteo_x1.3D-RAD
[ 🎯 NeurIPS 2025 ] 3D-RAD 🩻: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks
📢 News
What's New in This Update 🚀
2025.10.23: 🔥 Updated the latest version of the paper!
2025.09.19: 🔥 Paper accepted to NeurIPS 2025! 🎯
2025.05.16: 🔥 Set up the repository and committed the dataset!
🔍 Overview
💡 In this repository, we present the dataset for "3D-RAD: A… See the full description on the dataset page: https://huggingface.co/datasets/Tang-xiaoxiao/3D-RAD.MicrobenchiPinYou_x1
iPinYou_x1
Dataset description:
The iPinYou Global Real-Time Bidding Algorithm Competition is organized by iPinYou from April 1st, 2013 to December 31st, 2013.The competition has been divided into three seasons. For each season, a training dataset is released to the competition participants, the testing dataset is reserved by iPinYou. The complete testing dataset is randomly divided into two parts: one part is the leaderboard testing dataset to score and rank the participating… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/iPinYou_x1.rednote-xiaohongshu-notes
RedNote (Xiaohongshu) Notes with Engagement and Save Rates
10,427 posts from RedNote (小红书 / Xiaohongshu), across 8 content verticals, with full Chinese post text, four separate engagement metrics, and province-level geography.
Most social datasets give you text and a like count. RedNote separates saves from likes, and that single distinction turns out to measure something the like count cannot.
The finding this dataset exists for
A post can be useful or it can be… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/rednote-xiaohongshu-notes.x402-price-index
x402 Price Index — a census of the machine-payable web
Most directories of x402 services list what merchants claim. This dataset records
what their endpoints actually answer when probed: the HTTP 402 challenge they
return, the price inside it, the chain and asset they want, and whether they respond
at all.
It is built by Animica from an independent prober that walks
every x402 endpoint it can discover and reads each merchant's own accepts[] block.
No merchant self-reporting is… See the full description on the dataset page: https://huggingface.co/datasets/animicaorg/x402-price-index.Synthetic-UAV-Flight-Trajectories
UAV Trajectory Dataset
Summary
This dataset comprises over 5000 random UAV (Unmanned Aerial Vehicle) trajectories collected over 20 hours of flight time. It is intended for training AI models such as trajectory prediction applications. The dataset is generated through an automated pipeline for the creation and preprocessing of UAV synthetic trajectories, making it ready for direct AI model training.
Data Description
The dataset features parameterized… See the full description on the dataset page: https://huggingface.co/datasets/xianguang/Synthetic-UAV-Flight-Trajectories.model-xray-gallery
VG1 status · 22 September 2026 (Europe/Istanbul)
The 7B class opened on 21 September 2026. Any registered account may scan the repositories on the eligible list of the scope page — exact Apache-2.0 revisions, listed there with their status — within the free allowance of 20 browser scans and five distinct source models per calendar month. 7B-class repositories run as single-model quantization simulations. A comparison of two 7B-class checkpoints is currently accepted by the… See the full description on the dataset page: https://huggingface.co/datasets/tetracta/model-xray-gallery.TaobaoAd_x1
TaobaoAd_x1
Dataset description:
Taobao is a dataset provided by Alibaba, which contains 8 days of ad click-through data (26 million records) that are randomly sampled from 1140000 users. By default, the first 7 days (i.e., 20170506-20170512) of samples are used as training samples, and the last day's samples (i.e., 20170513) are used as test samples. Meanwhile, the dataset also covers the shopping behavior of all users in the recent 22 days, including totally seven hundred million… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/TaobaoAd_x1.Tickmill-XAUUSD-TicksCriteo_x4
Criteo_x4
Dataset description:
The Criteo dataset is a widely-used benchmark dataset for CTR prediction, which contains about one week of click-through data for display advertising. It has 13 numerical feature fields and 26 categorical feature fields. Following the setting with the AutoInt work, we randomly split the data into 8:1:1 as the training set, validation set, and test set, respectively.
The dataset statistics are summarized as follows:
Dataset Split
Total
#Train… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Criteo_x4.Avazu_x1
Avazu_x1
Dataset description:
This dataset contains about 10 days of labeled click-through data on mobile advertisements. It has 22 feature fields including user features and advertisement attributes. The preprocessed data are randomly split into 7:1:2* as the training set, validation set, and test set, respectively.
The dataset statistics are summarized as follows:
Dataset
Total
#Train
#Validation
#Test
Avazu_x1
40,428,967
28,300,276
4,042,897
8,085,794
Source:… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Avazu_x1.xMIL-HeatmapsHeatmap data for the multiple instance learning models presented in:
Jamshidi Idaji et al. "Beyond attention heatmaps: How to get better explanations for multiple instance learning models in histopathology". Medical Image Analysis (2026).
Link: https://www.sciencedirect.com/science/article/pii/S1361841526002173
Code: https://github.com/bifold-pathomics/xMIL
@article{
jamshidi26beyond,
title = {Beyond attention heatmaps: How to get better explanations for multiple instance learning models… See the full description on the dataset page: https://huggingface.co/datasets/bifold-pathomics/xMIL-Heatmaps.prompt-injection-attack-datasetCriteo_x2
Criteo_x2
Dataset description:
This dataset employs the Criteo 1TB Click Logs for display advertising, which contains one month of click-through data with billions of data samples. Following the same setting with the AutoGroup work, we select "data 6-12" as the training set while using "day-13" for testing. To reduce label imbalance, we perform negative sub-sampling to keep the positive ratio roughly at 50%. It has 13 numerical feature fields and 26 categorical feature fields. In… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/Criteo_x2.XBRL_analysis
XBRL Extraction Dataset
The is the official dataset introduced in the paper FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
clinical-trials-xml-2018-2024
