datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pxr-structure-pose-pool
PXR Structure Challenge — Full Multi-Model Pose Pool (184 ligands)
Every protein–ligand pose generated during the OpenADMET PXR (pregnane X receptor / NR1I2)
structure-prediction challenge, released openly with per-pose labels so the community can
reuse the compute already spent — and, we hope, crack the problem this data makes visible.
What's here
poses/<model>/<SID>.pdb — one best pose per (model, ligand). Protein chain A + ligand
(resname LIG). 15 models, up… See the full description on the dataset page: https://huggingface.co/datasets/xX-its-amit-Xx/pxr-structure-pose-pool.PoseX
PoseX: AI Defeats Physics Approaches on Protein-Ligand Cross Docking
Paper | GitHub | Leaderboard
PoseX is a comprehensive benchmark dataset designed to evaluate molecular docking algorithms for predicting protein-ligand binding poses. It includes curated datasets for both self-docking and cross-docking scenarios.
The file posex_set.zip contains the processed dataset for docking evaluation, while posex_cif.zip contains the raw CIF files from RCSB PDB. For information about creating… See the full description on the dataset page: https://huggingface.co/datasets/CataAI/PoseX.posebusters_benchmarkThe PoseBusters Benchmark dataset accompanying the PoseBench manuscript and benchmarking suite.
PoseStitch-ISLPoseidon
Poseidon: Global Earthquake Dataset (1990-2020)
This is the official dataset for the paper POSEIDON: Physics-Optimized Seismic Energy Inference and Detection Operating Network.
Overview
Poseidon is a largest opensource global earthquake dataset containing 2.8+ million seismic events spanning 30 years (1990-2020). Named after the Greek god of earthquakes, this dataset is designed for machine learning applications including earthquake prediction, seismic hazard analysis… See the full description on the dataset page: https://huggingface.co/datasets/BorisKriuk/Poseidon.SupraDB-PoseFeat
SupraDB-PoseFeat
What it is
SupraBench/SupraDB-PoseFeat is a SupraBench feature dataset generated from the
SupraEngineering compute pipeline. Pipeline position: Phase 1 docked-subset GLIDE pose feature computation by compute_all_features; Phase 2 highest-Boltzmann pose collapse.
The table is designed to join with the Phase 0 SupraDB-GEOM identity table and
the other feature datasets through inchikey.
Schema
Column
Dtype
Units
Meaning… See the full description on the dataset page: https://huggingface.co/datasets/SupraBench/SupraDB-PoseFeat.humanoid-pose-state-dataset-lite
Humanoid Pose State Dataset Lite
Lightweight synthetic dataset for humanoid robot pose classification.
Pose Classes
neutral
walking_pose
running_pose
sitting_pose
lifting_pose
waving_pose
Structure
dataset/
├── train/
├── validation/
Each split contains pose-labeled image folders.
Total Samples
Train: 600
Validation: 150
Image Format
RGB, 224x224
License
MIT
eval2_180_split_inspection_posebasedposeidon_train_datahumanoid_pose_sequencesexample_pose_data
