CoolFace
Datasetpublic

colabfit/JARVIS_Open_Catalyst_100K

Cite this dataset Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. JARVIS Open Catalyst 100K. ColabFit, 2023. https://doi.org/10.60732/ae1c7e2f This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange:… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_Open_Catalyst_100K.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
0likes38downloads
Dataset Card

<details><summary>Cite this dataset </summary>Chanussot, L., Das, A., Goyal, S., Lavril, T., Shuaibi, M., Riviere, M., Tran, K., Heras-Domingo, J., Ho, C., Hu, W., Palizhati, A., Sriram, A., Wood, B., Yoon, J., Parikh, D., Zitnick, C. L., and Ulissi, Z. JARVIS Open Catalyst 100K. ColabFit, 2023. https://doi.org/10.60732/ae1c7e2f</details>

This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:

https://materials.colabfit.org/id/DSarmtbsouma250

Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.

https://materials.colabfit.org <br><hr>

Dataset Name

JARVIS Open Catalyst 100K

Description

The JARVISOpenCatalyst_100K dataset is part of the joint automated repository for various integrated simulations (JARVIS) DFT database. This subset contains configurations from the 100K training, rest validation and test dataset from the Open Catalyst Project (OCP). JARVIS is a set of tools and datasets built to meet current materials design challenges.

Dataset authors

Lowik Chanussot, Abhishek Das, Siddharth Goyal, Thibaut Lavril, Muhammed Shuaibi, Morgane Riviere, Kevin Tran, Javier Heras-Domingo, Caleb Ho, Weihua Hu, Aini Palizhati, Anuroop Sriram, Brandon Wood, Junwoong Yoon, Devi Parikh, C. Lawrence Zitnick, Zachary Ulissi

Publication

https://doi.org/10.1021/acscatal.0c04525

Original data link

https://figshare.com/ndownloader/files/40902845

License

CC-BY-4.0

Number of unique molecular configurations

124929

Number of atoms

9719646

Elements included

Ag, Al, As, Au, B, Bi, C, Ca, Cd, Cl, Co, Cr, Cs, Cu, Fe, Ga, Ge, H, Hf, Hg, In, Ir, K, Mn, Mo, N, Na, Nb, Ni, O, Os, P, Pb, Pd, Pt, Rb, Re, Rh, Ru, S, Sb, Sc, Se, Si, Sn, Sr, Ta, Tc, Te, Ti, Tl, V, W, Y, Zn, Zr

Properties included

energy <br> <hr>

Usage

  • —ds.parquet : Aggregated dataset information.
  • —co/ directory: Configuration rows each include a structure, calculated properties, and metadata.
  • —cs/ directory : Configuration sets are subsets of configurations grouped by some common characteristic. If cs/ does not exist, no configurations sets have been defined for this dataset.
  • —cs_co_map/ directory : The mapping of configurations to configuration sets (if defined). <br>
ColabFit Exchange documentation includes descriptions of content and example code for parsing parquet files: