datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iSign
iSign: A Benchmark for Indian Sign Language Processing
The iSign dataset serves as a benchmark for Indian Sign Language Processing. The dataset comprises of NLP-specific tasks (including SignVideo2Text, SignPose2Text, Text2Pose, Word Prediction, and Sign Semantics). The dataset is free for research use but not for commercial purposes.
Quick Links
Website: The landing page for iSign
arXiv Paper: Detailed information about the iSign Benchmark.
Dataset on Hugging Face:… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/iSign.MA_Query_Expansion_MLT26openadmet-expansionrx-challenge-data
OpenADMET-ExpansionRx Challenge FULL dataset
This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026.
Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases.
While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.ExploreToM
Data sample for ExploreToM: Program-guided adversarial data generation for theory of mind reasoning
ExploreToM is the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation.
Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs.
Our A* search procedure aims to find… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ExploreToM.platonic-all-experimentsopenadmet-expansionrx-challenge-train-data
OpenADMET-ExpansionRx Challenge training dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.openadmet-expansionrx-challenge-test-data-blinded
OpenADMET-ExpansionRx Challenge blinded test dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-test-data-blinded.rl_llm_experiment_p6hle-flowbench-experiments-20260829
HLE FlowBench experiment archive
Private migration snapshot of the local HLE text-only 100-question research
program through 2026-09-01. It preserves the formal and smoke runs, per-question
Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state,
scores, monitoring, experiment controllers, reports, analyses, the paused-run
migration package, source Git bundles, and HLE-related host orchestration
sessions.
The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.ExpansionRx_OpenADMET_RLM_CLint
ExpansionRx-OpenADMET RLM CLint
RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules.
Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint.
Characteristic
Description
Tasks
1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.language-decoded-experiments
Language Decoded — Experiment Tracking
Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper.
Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code
⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.ExpansionRx_OpenADMET_KSOL
ExpansionRx-OpenADMET KSOL
KSOL dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict KSOL of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
7298
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_KSOL.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
transcript_isoform_expression_prediction
Multi-modal transcript isoform expression dataset
We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.ExpansionRx_OpenADMET_MGMB
ExpansionRx-OpenADMET MGMB
MGMB dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict MGMB of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
431
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_MGMB.ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB
ExpansionRx-OpenADMET Caco-2 Permeability Papp A>B
Caco-2 Permeability Papp A>B dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Papp A>B of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3773
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Papp_AB.Math-Expanded
Massive Step-by-Step Mathematics Instruction Dataset
Dataset Description
This is a 60GB, highly knowledge-dense dataset designed to teach Large Language Models (LLMs) rigorous mathematical reasoning.
Unlike standard math datasets that only provide the final answer, this dataset emphasizes Chain-of-Thought (CoT) reasoning. Every single row contains a detailed, step-by-step breakdown of how to arrive at the solution, making it ideal for supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Math-Expanded.rl_llm_experiment_p9CISLR
CISLR: Corpus for Indian Sign Language Recognition
This repository contains the Indian Sign Language Dataset proposed in the following paper
Paper: CISLR: Corpus for Indian Sign Language Recognition https://preview.aclanthology.org/emnlp-22-ingestion/2022.emnlp-main.707/
Authors: Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
Abstract: Indian Sign Language, though used by a diverse community, still lacks well-annotated… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CISLR.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux
ExpansionRx-OpenADMET Caco-2 Permeability Efflux
Caco-2 Permeability Efflux dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict Caco-2 Permeability Efflux of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
3777
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1:… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_Caco-2_Permeability_Efflux.global-top-Index-exploring-trends-in-stock-Market
Global Top Index: Exploring Trends in Stock Markets
About the Dataset
The Global Top Index dataset offers a detailed view of daily trading activities from several of the world's leading stock market indices. This dataset is ideal for conducting comprehensive analyses to uncover insights and predictive trends in the international stock markets.
Dataset Contents
The dataset encompasses the following key data points for each trading session across multiple dates… See the full description on the dataset page: https://huggingface.co/datasets/pettah/global-top-Index-exploring-trends-in-stock-Market.EXP-Bench
EXP-Bench dataset
❗ (Jan 19, 2026) We are currently working on an update that will improve the robustness and solvability of our dataset tasks. Please stay tuned! ❗
EXP-Bench is a novel benchmark designed to systematically evaluate AI agents on complete research experiments sourced from influential AI publications. Given a research question and incomplete starter code, EXP-Bench challenges AI agents to formulate hypotheses, design and implement experimental procedures, execute them… See the full description on the dataset page: https://huggingface.co/datasets/Just-Curieous/EXP-Bench.Gene_Expression_PredictionExpansionRx_OpenADMET_LogD
ExpansionRx-OpenADMET LogD
LogD dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through
scikit-fingerprints library.
The task is to predict LogD of molecules.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
7309
Recommended split
time
Recommended metric
MAE
References
[1]
OpenADMET team
"Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_LogD.ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
vn-provinces-life-expectancy
Vietnam provinces life expectancy at birth
Life expectancy at birth (years). Coverage 2018-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-life-expectancy.openadmet-expansionrx-challenge-teaser
OpenADMET-ExpansionRx Challenge TEASER dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-teaser.Math-Expanded
Massive Step-by-Step Mathematics Instruction Dataset
Dataset Description
This is a 60GB, highly knowledge-dense dataset designed to teach Large Language Models (LLMs) rigorous mathematical reasoning.
Unlike standard math datasets that only provide the final answer, this dataset emphasizes Chain-of-Thought (CoT) reasoning. Every single row contains a detailed, step-by-step breakdown of how to arrive at the solution, making it ideal for supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Math-Expanded.power-plant-climate-exposure-index
Global Power-Plant Climate Exposure Screening Index (CESI)
Per-plant outdoor-environment severity for the world's power fleet: a 0-100 index, a CES1-CESX screening class and the dominant stressor, derived from each plant's own coordinates (WorldClim normals, Koppen-Geiger class, distance to coast). The full model ships as severity_model.py.
Canonical record: doi.org/10.5281/zenodo.22172589 · Publisher: Inzonex
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Inzinion/power-plant-climate-exposure-index.
