datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airfrans
AirfRANS
Dataset Description
AirfRANS is a high-fidelity two-dimensional airfoil CFD dataset introduced in a paper from the NeurIPS 2022 Datasets and Benchmarks Track. The data is intended for surrogate modeling of the incompressible steady-state Reynolds-averaged Navier-Stokes (RANS) equations and contains 1,000 subsonic airfoil simulation cases.
The cases cover NACA four-digit and five-digit airfoils, with Reynolds numbers ranging from 2 million to 6 million and… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/airfrans.ShapeNetCar
ShapeNetCar
Dataset Description
The ShapeNetCar dataset comes from the paper Learning Three-dimensional Flow for Interactive Aerodynamic Design by Umetani and Bickel, published in ACM Transactions on Graphics (SIGGRAPH 2018). Based on three-dimensional car geometries from ShapeNet, the dataset uses CFD simulations to obtain velocity fields around the vehicles, surface pressure, and drag coefficients. It supports research on rapidly predicting aerodynamic physical… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/ShapeNetCar.AlphaFold3_dataset
AlphaFold3 Dataset
Dataset Description
AlphaFold3_dataset is the companion dataset for the OneScience-Group/AlphaFold3/ model. It includes inference input examples, public search databases, MMseqs databases, sharded Jackhmmer databases, and alignment-related directories. The dataset ID is fixed as OneScience-Group/AlphaFold3_dataset.
Supported Tasks
This dataset supports the AlphaFold3 biomolecular structure prediction workflow. It provides both… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/AlphaFold3_dataset.ERA5
ERA5
Dataset Description
ERA5 is the fifth-generation global atmospheric reanalysis dataset produced by the European Centre for Medium-Range Weather Forecasts (ECMWF). By combining global observations with physical models, it provides consistent and comprehensive global estimates of hourly, high-resolution (approximately 31 km) climate variables covering the atmosphere, land, and ocean from 1940 to the present.
Supported Tasks
This standardized data… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/ERA5.NEP
NEP
Dataset Description
The NEP dataset provides example data for training materials potentials with OneScience-Group/NEP. It includes several materials systems, such as AuAg, Cu, HfO2, and LiSiC, in pwmat/movement, pwmlff/npy, and extxyz formats. The dataset contains atomic coordinates, cell information, energies, and forces for different configurations of these systems. It is intended as a standard benchmark for learning and validating machine-learning interatomic… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/NEP.Kolmogorov_flow_2d
Kolmogorov Flow 2D
Dataset Description
This dataset contains time series of single-channel vorticity fields obtained from numerical simulations of two-dimensional Kolmogorov Flow. It can be used for turbulence time-series forecasting, neural operator training, partial differential equation surrogate modeling, and long-horizon autoregressive forecasting.
The data describes scalar vorticity fields over a two-dimensional periodic domain. The dataset contains 120… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/Kolmogorov_flow_2d.mp20
MP-20
Dataset Description
MP-20 is derived from the Materials Project database (Jain et al., 2013) and contains approximately 45,000 common inorganic materials. It covers most experimentally known materials with no more than 20 atoms in the unit cell. The data includes elemental compositions, CIF structures, space groups, formation energies per atom, DFT band gaps, bulk moduli, and magnetic densities.
Source paper: A. Jain et al., Commentary: The Materials Project:… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/mp20.cfd_benchmark
CFD Benchmark
Dataset Description
CFD Benchmark is a comprehensive benchmark dataset for training and evaluating neural PDE solvers. It contains six standard tasks on regular grids, structured grids, and irregular geometries. The dataset was used for unified evaluation in the ICML 2024 paper Transolver, and some of its data originates from the FNO and Geo-FNO works.
The dataset contains six subsets: airfoil, darcy, elasticity, ns, pipe, and plas.
Paper: Transolver:… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/cfd_benchmark.beno
BENO
Dataset Description
The BENO dataset originates from the ICLR 2024 paper BENO: Boundary-Embedded Neural Operators for Elliptic PDEs and is designed for solving elliptic partial differential equations under complex boundary conditions. The data contains random boundary geometries with four, three, two, one, or no corners, all standardized to a 32 x 32 grid resolution.
Paper: BENO: Boundary-Embedded Neural Operators for Elliptic PDEs
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/beno.pdenneval
PDENNEval
Dataset Description
PDENNEval is a comprehensive dataset for evaluating neural-network-based PDE solving methods, introduced in an IJCAI 2024 paper. It covers function learning and operator learning tasks and includes 15 types of PDE problems across multiple scientific domains, including fluids, materials, finance, and electromagnetics.
The dataset consists of 10 PDEBench data files and 6 self-generated data files, totaling approximately 286.9 GB. It can… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/pdenneval.MatPL
MatPL
Dataset Description
The MatPL dataset is an example dataset for training material potentials with OneScience-Group/NEP. It contains multiple material systems, including AuAg, Cu, HfO2, and LiSiC, and covers the pwmat/movement, pwmlff/npy, and extxyz data formats. The data include atomic coordinates, simulation cell information, energies, and forces for different material systems in different configurations, and serve as a standard benchmark dataset for… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/MatPL.deepcfd
DeepCFD
Dataset Overview
The DeepCFD dataset is a publicly available two-dimensional steady-state laminar-flow computational fluid dynamics dataset created by a research team affiliated with the German Research Center for Artificial Intelligence (DFKI). The dataset was generated with OpenFOAM's simpleFoam solver and contains approximately 1,000 samples of flow around randomly generated obstacles in a channel. It is suitable for CFD solver acceleration, evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/deepcfd.cfdbench
CFDBench
Dataset Description
CFDBench is a large-scale benchmark dataset for machine learning methods in computational fluid dynamics, designed to evaluate the generalization capabilities of neural operators under unseen boundary conditions, fluid properties, and geometries.
The dataset contains four classic CFD problems: lid-driven cavity flow (cavity), laminar pipe flow (tube), step dam-break flow (dam), and flow around a cylinder (cylinder). For each problem… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/cfdbench.cylinder_flow
Cylinder Flow
Dataset Description
The Cylinder Flow dataset is sourced from DeepMind's MeshGraphNets benchmark and describes flow around a cylinder (vortex shedding) on a two-dimensional unstructured triangular mesh. The data was generated using COMSOL simulations. Each trajectory contains 600 time steps with a time-step size of dt=0.01 and records the fluid velocity and pressure fields.
Paper: Learning Mesh-Based Simulation with Graph Networks
Supported… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/cylinder_flow.alphagenome_dataset
AlphaGenome Dataset
Dataset Description
AlphaGenome Dataset is a curated package of reference genomes, annotation files, and TFRecord data for use with the OneScience AlphaGenome model.
The reference resources include human GRCh38.p13, mouse GRCm38.p6, FAI indexes, Gencode annotations, splice-site annotations, and polyA annotations. The TFRecord data contains 12 bundles from the HOMO_SAPIENS VALID split, with 25 gzip-compressed shards in each bundle.… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/alphagenome_dataset.CMEMS
CMEMS
Dataset Description
CMEMS is an HDF5 gridded dataset for global ocean forecasting tasks. By integrating satellite and in situ observations with numerical ocean models, it provides real-time analyses, forecasts, and historical reconstructions for a variety of variables.
Supported Tasks
This standardized data repository contains 7 annual files covering 1993-1999. Each file provides ocean variable fields, global means, and global standard… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/CMEMS.oxides
Oxides
Dataset Description
Oxides is derived from the oxide polymorph study published by Mehta, Salvador, and Kitchin in 2015. FAIR Chemistry provides the JSON data from the paper's supporting information with its fine-tuning tutorial. This repository extracts PBE/EOS/calculations from that data and standardizes it as ASE SQLite databases.
The dataset covers five oxides—IrO2, RuO2, SnO2, TiO2, and VO2—across 30 oxide/polymorph groups. Each record contains a… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/oxides.State_datasets
State Dataset
Dataset Description
State_dataset is a collection of datasets used for State single-cell expression modeling and perturbation prediction tasks. It comprises four data categories: Parse, Tahoe, Replogle-Nadig, and SE-167M-Human. The primary data is in AnnData/H5AD format, accompanied by gene embeddings (PyTorch .pt), dataset split configurations (TOML), and upstream license files.
Supported Tasks
This repository corresponds to the… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/State_datasets.ani1x
ANI-1x
Dataset Description
The ANI-1x dataset is a molecular-configuration dataset for training and evaluation. The data package contains training, validation, and test HDF5 shards, along with a statistics JSON file and auxiliary extxyz files. Its fields are consistent with the energy and force fields in the MACE ANI-1x configuration, making it a standard benchmark dataset for learning and validating machine-learning interatomic potentials with MACE.
The element set… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/ani1x.proteinmpnn
ProteinMPNN Dataset
Dataset Description
The ProteinMPNN training sample dataset contains PDB-derived structural data for protein sequence design tasks. The data directory is pdb_2021aug02_sample, which contains a structure index table, validation and test cluster split files, and .pt structure tensor files organized by PDB and chain.
Supported Tasks
This resource is intended for validating training data loading, debugging the training workflow, and… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/proteinmpnn.eagle
EAGLE
Dataset Description
The EAGLE dataset is sourced from an ICLR 2023 paper and is designed for predicting two-dimensional unsteady turbulent flows on unstructured dynamic meshes. The data describes the velocity and pressure fields produced by interactions between a moving flow source and various ground geometries. It contains three types of geometric configurations: Cre, Spl, and Tri.
Paper: EAGLE: Large-scale Learning of Turbulent Fluid Dynamics with Mesh… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/eagle.lagrangian
Lagrangian
Dataset Overview
The Lagrangian dataset is sourced from the DeepMind team's ICML 2020 paper Learning to Simulate Complex Physics with Graph Networks. It consists of the two-dimensional Water particle-dynamics data from the paper's Graph Network-based Simulator (GNS) benchmark. The data represents particles as graph nodes and describes fluid evolution over time through particle-position sequences and particle types. It can be used for Lagrangian… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/lagrangian.evo2_dataset
Evo2 Dataset
Dataset Description
Evo2 Dataset is an Evo2 mini genome dataset adapted for OneScience/evo2/. It contains FASTA, compressed FASTA, and merged FASTA files for human chr20, chr21, and chr22, as well as train, validation, and test .bin/.idx splits preprocessed with the Byte-Level tokenizer.
Supported Tasks
This dataset is not the complete OpenGenome2 dataset and is not intended to reproduce full-scale pretraining. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/evo2_dataset.nanotube
nanotube
Dataset Description
The nanotube dataset is an extxyz dataset for carbon nanotube training and evaluation examples. It contains atomic coordinates, energies, and force data for carbon nanotube systems in different configurations and is a standard benchmark dataset for developing and validating machine-learned interatomic potential methods.
The training file nanotube_large.xyz contains 4,000 frames, while the test file nanotube_test.xyz contains 1,032… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/nanotube.oc20
OC20
Dataset Description
The OC20 dataset is an extxyz subset of the Open Catalyst 2020 (OC20) S2EF (Structure to Energy and Forces) dataset. It contains atomic coordinates, unit-cell information, energies, and forces for adsorption configurations on catalytic material surfaces. Its energy and force fields are consistent with those in the UMA OC20 configuration, making it a standard benchmark dataset for developing and validating machine-learned interatomic… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/oc20.DMC
DMC
Dataset Description
DMC is a solvent XTB extxyz dataset for introductory training examples. It contains atomic coordinates, energies, and forces for carbonate molecular systems (VC, EC, PC, DMC, EMC, and DEC) in different configurations, and serves as a standard introductory benchmark dataset for learning and validating machine-learning interatomic potential methods.
The training file solvent_xtb_train_200.xyz contains 203 frames (200 training configurations + 3… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/DMC.water
water
Dataset Description
The water dataset is an extxyz dataset for potential energy surface training and evaluation examples involving water systems. It contains atomic coordinates, unit-cell information, energies, and forces for water molecular systems in different configurations and is a standard benchmark dataset for developing and validating machine-learned interatomic potential methods such as MACE.
The original complete file dataset_1593.xyz contains 1,593… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/water.DeePMD
