datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p2p-full-data
Open Pixel2Play (P2P) Full Dataset
Paper | GitHub | Project Page | Toy Dataset
The p2p-full-data dataset contains 8300+ hours of high-quality human annotated data, spanning across more than 40 popular 3D video games. All gameplay is recorded at 20 FPS by experienced players. Each frame is annotated with keyboard and mouse actions, and text instructionsare provided when available.
If you found the dataset helpful, please consider upvoting the paper so it can reach more people!… See the full description on the dataset page: https://huggingface.co/datasets/elefantai/p2p-full-data.synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kVLAC-Cut-FullData
VLAC-Cut-FullData
VLAC-Cut-FullData is the full-data release for VLAC-Cut. It provides the complete raw-data archive set, benchmark-style JSON files, and a lightweight frame-extraction workflow for reproducing evaluation on the released benchmark protocol.
Contents
benchmark_style_all/
train/video_progress_benchmark_file.json
test_expert_seen/video_progress_benchmark_file.json
test_expert_unseen/video_progress_benchmark_file.json… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-FullData.financial-analyst-data-full
financial-analyst-data-full
A-share historical price + valuation data packaged for financial-analyst —
the 14-agent single-stock deep-dive research workstation.
Published: 2026-05-24
Preset: full — 全 A 股完整包 (含历史退市股). 量化研究员 / 重度用户. lite 全 + TDX 历年财报原始 zip (用户跑 import_tdx_financial.py 解) + F10 原始文本 (公司大事/龙虎榜/主力追踪/最新提示 .txt).
Size: ~14.1 GB
What's included
5450 stocks daily OHLCV + 7 valuation fields (PE/PB/PS/DV/MV/CIRC_MV/turnover_rate)
Date range (daily): 1990-12-19… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-full.beamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
bge-full-datadataset_esc50_full_v1polymarket-full-market-dataset
Polymarket Full Market Dataset
A complete, point-in-time archive of Polymarket: every event and market served by
Polymarket's public Gamma API — from the platform's launch in 2020 through the snapshot
date — with final resolution outcomes, full metadata (questions, rules, tags, timing,
order-book flags), and dense daily OHLC price candles for every outcome token,
assembled from the CLOB price-history API.
One snapshot = one self-contained JSON file. One row = one market, with… See the full description on the dataset page: https://huggingface.co/datasets/od2961/polymarket-full-market-dataset.dataset_gise_full_v1tadabur-lora-data-fulllibrispeech-full-dataset-modeldataset_urbansound8k_full_v1full-dataset-of-combined-malware-samplesfull_tmdb_movies_datasetdata_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.stackcube_data_scaling_new_fullgpt2small_full_training_datapd-voice-full-multimodal-dataset
Parkinson Voice — Full Multimodal Dataset
Complete Parkinson’s vs healthy voice package for classification and explainable Mel reasoning research (EDGE).
Not Mel-only: raw audio, 10 visual modalities, feature CSVs, plus Gemma reasoning traces for Mel.
Contents
Path
Description
audio/
Waveform clips (healthy / parkinsons), 1134 files
images/mel/
Mel spectrograms
images/spectrogram/
Linear spectrograms
images/mfcc/
MFCC maps
images/delta_mfcc/… See the full description on the dataset page: https://huggingface.co/datasets/mdimamhosen/pd-voice-full-multimodal-dataset.vla0-context-trace-full-demo-datasetyt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
pubmed-ai-full-datar2n2_shapenet_dataset_fullONS-Capacity-Factor-Dataset-Full
Dataset Card for ONS-Capacity-Factor-Dataset-Full
Dataset Summary
The ONS-Capacity-Factor-Dataset-Full provides hourly data for wind and solar power plants in Brazil. These values are published by the Operador Nacional do Sistema Elétrico (ONS) — the Brazilian National Electric System Operator — which is responsible for coordinating and controlling electricity generation and transmission in the National Interconnected System (SIN).
The dataset includes data from 2009 to… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/ONS-Capacity-Factor-Dataset-Full.lucidprots_full_data
Dataset Card for "lucidprots_full_data"
More Information needed
Full-Duplex-Bench-DataBenglaAI_Full_Datasetanchorworld-dataset-full
AnchorWorld Dataset
Paper: AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
Project page: https://yuli0103.github.io/AnchorWorld/
Overview
The AnchorWorld Dataset is the processed training dataset used by AnchorWorld, an egocentric world simulation framework controlled by 3D human motion and customizable through spatially grounded anchor views.
The dataset is derived from two public multi-view human-activity datasets:… See the full description on the dataset page: https://huggingface.co/datasets/lyabc/anchorworld-dataset-full.soda-vec-data-full_pmc_title_abstract
SODA-VEC Clean Dataset
This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles.
Dataset Overview
Total examples: 26,573,900
Training set: 26,473,900 examples (99.6%)
Validation set: 50,000 examples (0.2%)
Test set: 50,000 examples (0.2%)
Quality Filtering Applied
This dataset has been processed with the following quality filters:
Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.RoboTwin_Dataset_fullgaze_dataset_full
gaze_dataset_full — StreamGaze_v2 + EgoGazeVQA + HD-EPIC
A single repository containing three complementary benchmarks for
evaluating multimodal LLMs on gaze-grounded egocentric video question
answering:
Subfolder
Source
Held-out (val/test)
Questions
StreamGaze_v2/
egoexolearn, holoassist, egtea
egtea
8 MCQ tasks (4-opt) — gaze-conditioned past/present/future
EgoGazeVQA/
ego4d, egoexo, egtea
egtea
causal / spatial / temporal (5-opt)
HD-EPIC/
P01–P09
P09… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/gaze_dataset_full.
