datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.TDC_caco2_wang
TDC Caco-2 Wang
Caco-2 Wang dataset [1], part of TDC [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rate at which drug passes through Caco-2 cells that serve as in vitro simulation of human intestinal tissue.
This dataset is a part of "absorption" subset of ADME tasks.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
910
Recommended splitscaffold
Recommended metric
MAE… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/TDC_caco2_wang.plutchik-storieswandb_logsXBRL_analysis
XBRL Extraction Dataset
The is the official dataset introduced in the paper FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
UVE-Subjective-Benchmark
Dataset Card for UVE-Subjective-Benchmark
🌊 Dataset Summary
The evaluation of Underwater Video Enhancement (UVE) remains a significant challenge due to the complex visual degradations inherent to aquatic environments and the temporal instability frequently introduced by frame-wise processing. Existing objective metrics (e.g., UIQM, UCIQE) often correlate poorly with human subjective judgement, particularly regarding dynamic artifacts like flickering and color… See the full description on the dataset page: https://huggingface.co/datasets/eddy-Wang/UVE-Subjective-Benchmark.reiss-storiesgapa
GAPA: Gender Associations of Physical Attributes
The GAPA dataset contains 14,706 human ratings of how strongly English physical descriptions (n=316, e.g., "a defined jawline", "a soft voice", "broad shoulders") are associated with a woman, a man, or a non-binary person.
Paper Figure 1 — an excerpt of the most gender-distinctive attributes in each ranking pattern, with their per-gender association ratings.
How to use
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/alisa-yingjia-wan/gapa.plutchik-nine-choiceslyu-wang-balius-singh-2019-ampc
Ultra-large docking data: AmpC 96M compounds
These data are from John J. Irwin, Bryan L. Roth, and Brian K. Shoichet's labs. They published it as:
[!NOTE]Lyu J, Wang S, Balius TE, Singh I, Levit A, Moroz YS, O'Meara MJ, Che T, Algaa E, Tolmachova K, Tolmachev AA, Shoichet BK, Roth BL, Irwin JJ.
Ultra-large library docking for discovering new chemotypes. Nature. 2019 Feb;566(7743):224-229. doi: 10.1038/s41586-019-0917-9.
Epub 2019 Feb 6. PMID: 30728502; PMCID: PMC6383769.… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/lyu-wang-balius-singh-2019-ampc.maslow-stories
Dataset Card for [Dataset Name]
annotations_creators:
expert-generated
language_creators:
expert-generated
languages:
english
licenses:
unknown
multilinguality:
monolingual
pretty_name: maslow-stories
size_categories:
unknown
source_datasets: []
task_categories:
question-answering
task_ids:
multiple-choice-qa
discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.wanderplan-dataset
🚀 The Three-Notebook Ecosystem
This dataset serves as the foundational backbone for the WanderPlan AI travel utility platform, developed and executed across three distinct, pipeline-connected Google Colab notebooks:
Stage
Notebook Purpose
Key Operations
Colab Link
1
Data Generation & Engineering
LLM Synthesis, Data Cleaning, Rule-Based QA Patching, Dynamic VIP Injection, Geo-Tagging
Open in Colab
2
Exploratory Data Analysis (EDA)
Cost Distribution Analysis… See the full description on the dataset page: https://huggingface.co/datasets/nirm2002/wanderplan-dataset.EgoVid-5M
Data annotation code
https://github.com/JeffWang987/EgoVid
MapSatisfyBench-MockData-ToolsThe datasets associated with MapSatisfyBench include the benchmark data file "MapSatisfyBench_Benchmark.csv" and the mock data (other .csv files) used by the sandbox tools during simulation execution.
job-educational-parser-dataset-08-0-0805
Job Educational Parser Dataset
招聘领域的岗位与学历要求数据集。
输入:岗位描述 -> 输出:学历要求
Splits
train: 19w_0701.csv (约 19 万条)
test: 2w_0716.csv (约 2 万条)
validation: 4w_0708.csv (约 4 万条)
每条数据至少包含字段:
user: 职位描述
assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低
由 @wangzihaogithub 创建。
wanderlust-ai-dataset
✈️ Wanderlust AI — Travel Recommendation Dataset
A synthetic dataset of AI-generated travel destinations and points of interest, built for a two-layer travel recommendation and itinerary-generation application.
📓 Project Notebook
The full project notebook (data generation, cleaning, EDA, recommendation system, generation, and application code) is available here:
Wanderlust_AI_Final_Project.ipynb
📋 Project Overview
Wanderlust AI matches a user's… See the full description on the dataset page: https://huggingface.co/datasets/Omerinbar/wanderlust-ai-dataset.Verilog_dataMedQA-Calc
Description
MedQA-Calc is a medical calculator dataset used to benchmark LLMs ability to recommend clinical calculators. Each instance in the dataset consists of a truncated patient note, a question asking to recommend a specific clinical calculator, answer options (including "None of the above"), and a final answer value. Our dataset covers 35 different calculators. This dataset contains a training dataset of about 5,000 instances and a testing dataset of 1,009 instances.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas-Wan/MedQA-Calc.thai-mm-dict
🗂️ Dataset Card: Thai-Myanmar Dictionary (2025)
📝 Dataset Summary
The Thai-Myanmar Dictionary (2025) is a high-quality bilingual lexical dataset created by Htet Myet Lynn.It provides direct word-to-word and phrase mappings between Thai and Myanmar (Burmese), supporting both linguistic use and machine learning applications.
The dataset is released under the MIT License, allowing free usage, modification, redistribution, and integration into both academic and commercial… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-mm-dict.glaiveai-reflection-v1Mirror from glaiveai/reflection-v1.
discord-phishing-scam
Discord Scam / Clean Messages Dataset
A small but carefully-curated dataset for binary text-classification:
“Is this Discord message trying to scam / spam users?”
It is intended as a starting point for fine-tuning lightweight BERT-style models that moderate real-time chat servers.
1 Origin & Collection
Source servers – private Discord communities (11 k members in total) run by the author.
Period – 2024-01-01 → 2025-06-01.
Extraction – Discord.py script iterated… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam.chinesespringboot_vullyu-wang-balius-singh-2019-d4
Ultra-large docking data: D4 receptor 115M compounds
These data are from John J. Irwin, Bryan L. Roth, and Brian K. Shoichet's labs. They published it as:
[!NOTE]Lyu J, Wang S, Balius TE, Singh I, Levit A, Moroz YS, O'Meara MJ, Che T, Algaa E, Tolmachova K, Tolmachev AA, Shoichet BK, Roth BL, Irwin JJ.
Ultra-large library docking for discovering new chemotypes. Nature. 2019 Feb;566(7743):224-229. doi: 10.1038/s41586-019-0917-9.
Epub 2019 Feb 6. PMID: 30728502; PMCID: PMC6383769.… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/lyu-wang-balius-singh-2019-d4.reiss-twenty-choicesthai-w2p
Thai W2P
Thai Word-to-Phoneme (W2P) converter.
GitHub: https://github.com/wannaphong/thai_w2p
