datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/Bilsteen/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC
Dataset Summary
This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan 2026), processed into a clean, machine-learning-ready Parquet format.
Unlike raw XML/JSON dumps or standard CSVs, this dataset provides a unified, tabular structure that joins multiple chemical identifiers and descriptors into a single sharded resource:
SMILES: Raw and RDKit-Canonicalized.
SELFIES: Pre-computed 100% robust… See the full description on the dataset page: https://huggingface.co/datasets/th-laurel/PubChem-124M-SMILES-SELFIES-InChI-IUPAC.details_aisquared__dlite-v2-124m
Dataset Card for Evaluation run of aisquared/dlite-v2-124m
Dataset Summary
Dataset automatically created during the evaluation run of model aisquared/dlite-v2-124m on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_aisquared__dlite-v2-124m.details_aisquared__dlite-v1-124m
Dataset Card for Evaluation run of aisquared/dlite-v1-124m
Dataset Summary
Dataset automatically created during the evaluation run of model aisquared/dlite-v1-124m on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_aisquared__dlite-v1-124m.details_xzuyn__GPT-2-SlimOrcaDeduped-airoboros-3.1-MetaMathQA-SFT-124M
Dataset Card for Evaluation run of xzuyn/GPT-2-SlimOrcaDeduped-airoboros-3.1-MetaMathQA-SFT-124M
Dataset automatically created during the evaluation run of model xzuyn/GPT-2-SlimOrcaDeduped-airoboros-3.1-MetaMathQA-SFT-124M on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_xzuyn__GPT-2-SlimOrcaDeduped-airoboros-3.1-MetaMathQA-SFT-124M.details_nicholasKluge__Aira-124M
Dataset Card for Evaluation run of nicholasKluge/Aira-124M
Dataset Summary
Dataset automatically created during the evaluation run of model nicholasKluge/Aira-124M on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nicholasKluge__Aira-124M.details_nicholasKluge__Aira-Instruct-124M
Dataset Card for Evaluation run of nicholasKluge/Aira-Instruct-124M
Dataset Summary
Dataset automatically created during the evaluation run of model nicholasKluge/Aira-Instruct-124M on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nicholasKluge__Aira-Instruct-124M.details_nicholasKluge__Aira-2-124M
Dataset Card for Evaluation run of nicholasKluge/Aira-2-124M
Dataset Summary
Dataset automatically created during the evaluation run of model nicholasKluge/Aira-2-124M on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nicholasKluge__Aira-2-124M.gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
details_MBZUAI__LaMini-GPT-124M
Dataset Card for Evaluation run of MBZUAI/LaMini-GPT-124M
Dataset Summary
Dataset automatically created during the evaluation run of model MBZUAI/LaMini-GPT-124M on the Open LLM Leaderboard.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_MBZUAI__LaMini-GPT-124M.gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
gpt2-124M-qlora-chat-support
Dataset Card for "gpt2-124M-qlora-chat-support"
More Information needed
