datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HR-VILAGE-3K3M
HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression
This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.jepa-qwen3-32b-pure-baselines-2026-05-25
JEPA-Align: Qwen3-32B Safety Defense Matrix
The complete 11-condition Qwen3-32B experiment for Predictive Representation
Alignment (PRA), the paired-view objective introduced in Predictive
Representation Alignment Improves Generalization in LLM Safety.
PRA aligns adversarially rewritten prompts with clean prompts expressing the
same intent. This release contains trained adapters, attack traces, benign
capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.PersonaMem-v3
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell,
Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
A collaboration between:
Meta Recommendation Systems
University of Pennsylvania
MIT
Third release in the PersonaMem series:
PersonaMem-v1: [COLM… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/PersonaMem-v3.BoilingBench-CV
BoilingBench-CV Dataset
Version: v0.1.0
Maintainer: NED3 Laboratory, University of Arkansas
License: CC BY 4.0
DOI: 10.5281/zenodo.22264378
Mirror of the Zenodo deposit of 3 September 2026, published here because most
users of these data work in the Hugging Face ecosystem. The file set was
verified identical to the deposit at upload time: 7,147 files, 4.20 GB
uncompressed.
Authors
Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University),
Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.kernelbench-v3-runs
KernelBench-v3 — Agent Runs
2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py.
Companion datasets:
Infatoshi/kernelbench-v3-problems — 60 problem definitions
Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.aopsargument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.hsk-3.0-dataset
HSK Vocabulary List
Dataset Summary
This dataset contains 5,456 HSK vocabulary entries in a simple CSV format for use on Hugging Face.
Files
hsk.csv: UTF-8 CSV file with five columns
Data Fields
id: integer identifier
hsk_level: HSK level from 1 to 6
chinese: Chinese vocabulary item
pinyin: pinyin with tone marks
english: English gloss, with multiple translations separated by ;
Dataset Statistics
Total rows: 5,456
HSK 1: 500
HSK 2:… See the full description on the dataset page: https://huggingface.co/datasets/Tiagodfs/hsk-3.0-dataset.bitaudit_verification_dataset_v2web3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models.
arc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.daigt_v3TCGA_OncoTree_pt2
TCGA_OncoTree_pt2
1. Tổng quan
[CẦN ĐIỀN THỦ CÔNG: mục đích, ngữ cảnh tạo dataset]
Tổng số bản ghi (cộng tất cả manifest phát hiện được): 23984
Số manifest phát hiện được trong bộ nhớ: 3 (df, labels_df, progress)
Repo HuggingFace chính: ento3686/TCGA_OncoTree_pt2
⚠️ Dataset được lưu trên 2 repo/tài khoản HuggingFace khác nhau:
ento3686/TCGA_OncoTree_pt2 (biến: REPO_ID_2, UPLOAD_REPO_ID, CENTRAL_PROGRESS_REPO_ID, _repo_id_var)
tuna2004/TCGA_OncoTree (biến:… See the full description on the dataset page: https://huggingface.co/datasets/ento3686/TCGA_OncoTree_pt2.everyday-manipulation-3d-raw
Everyday Manipulation 3D (raw RGB-D)
1,513 clips · 10.28 hours · 279 GiB · 4 participants · 10 manipulation tasks · 42 recording sittings
Chest-mounted iPhone Pro capture of everyday two-handed manipulation by
CaryX AI. Clips were recorded with
Record3D, an iOS app that captures the
iPhone's LiDAR RGB-D stream. Each clip is the app's .r3d recording with the
audio track removed; the sensor streams are unmodified: synchronised RGB,
metric LiDAR depth, per-frame ARKit 6-DoF camera… See the full description on the dataset page: https://huggingface.co/datasets/CaryxAI/everyday-manipulation-3d-raw.C3
C3: Cross-View Cross-Modality Correspondence Dataset
Dataset for C3Po: Cross-View Cross-Modality Correspondence with Pointmap Prediction
arXiv | Project Website | GitHub
C3 contains 90K paired floor plans and photos from the Internet across 597 scenes with 153M pixel-level correspondences and 85K camera poses.
Image Pairs
image_pairs/ is split into train/, val, and test/, each with a image_pairs.csv.
image_pairs.csv: Each row represents a plan-photo pair… See the full description on the dataset page: https://huggingface.co/datasets/kwhuang/C3.GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.Jeong-2022-InternationalBCICompetition2020Review-track3
BCI Competition 2020 Track 3: imagined speech classification
This is an unofficial mirror of Track 3 only from the 2020 International BCI
Competition, described by Jeong et al. (2022). It is not affiliated with or
endorsed by the authors, their institutions, or OSF.
Source and attribution
Original data: 2020 International BCI Competition, OSF project pq7vb,
folder Track#3 Imagined speech classification.
Paper: Jeong et al., 2020 International brain–computer… See the full description on the dataset page: https://huggingface.co/datasets/raei/Jeong-2022-InternationalBCICompetition2020Review-track3.P31652b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.maze-30x30-hard-1kPRCA-Net-dataset
Ray-Traced Cross-Frequency Radio Map Dataset
A large ray-traced radio-map (path-loss) dataset for zero-shot
cross-frequency generalization research, generated with
Sionna RT over real urban geometry from
OpenStreetMap.
150 urban scenes across 15 cities, 256×256 rasters
8 transmitters per scene across three deployment strata (street,
rooftop, mast)
6 carrier frequencies: 1.8, 3.5, 7, 28 GHz (training) + 10, 60 GHz
(held out, for interpolation / extrapolation studies)
7,200… See the full description on the dataset page: https://huggingface.co/datasets/SHussain37/PRCA-Net-dataset.new-title-chineseprogram_generation_v3Koala-36M-v1bank-additional-fullin1k_clip_qwen25vl_3b_224res_64tokens_new_ptHUI360
HUI360
HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation (IEEE FG 2026)
Open-access skeleton annotations for HUI360, a large-scale 360° egocentric dataset for human-robot interaction anticipation in the wild. This repository provides the annotations as tabular CSV files (one row per detection), ready for training and evaluation with HUI360-Baselines.
Related resources
Resource
Link
Project… See the full description on the dataset page: https://huggingface.co/datasets/rlorlou/HUI360.agiqa-3kDataset in paper [IEEE TCSVT2023] [Agiqa-3k: An open database for ai-generated image quality assessment](https://huggingface.co/papers/2306.04717)
Code: https://github.com/lcysyzxdxc/AGIQA-3k-Database
@ARTICLE{10262331,
author={Li, Chunyi and Zhang, Zicheng and Wu, Haoning and Sun, Wei and Min, Xiongkuo and Liu, Xiaohong and Zhai, Guangtao and Lin, Weisi},
journal={IEEE Transactions on Circuits and Systems for Video Technology},
title={AGIQA-3K: An Open Database for AI-Generated Image… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/agiqa-3k.in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_pt
