datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.pxr-challenge-train-test
PXR Challenge Train/Test Dataset
A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge.
Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction
Challenge Space: openadmet/pxr-challenge
Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.openadmet-expansionrx-challenge-train-data
OpenADMET-ExpansionRx Challenge training dataset
This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.FARM_training_test
FARM Aerial Radio Map (ARM) Dataset
Paper:
FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking (https://arxiv.org/abs/2604.17362)
Overview
This repository releases the constructed ARM datasets based on ARM-Omni for FARM training, in-domain evaluation (D1-D10), and zero-shot evaluation (P1, F1, and A1). The dataset coverage is summarized below:
Dataset
Frequencies (GHz)
Max Rx Height (m)
Beamwidths
Map Grid Size
Volume
D1
2.1… See the full description on the dataset page: https://huggingface.co/datasets/jliang097/FARM_training_test.STARCOP_allbands_Train1
STARCOP dataset
STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham
Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP
Task:
Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.FrontierCO-TrainReverse-hybrid-train-no-persona-meanReverse-hybrid-correct-train-no-persona-meanDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tripclick-training
TripClick Baselines with Improved Training Data
Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury
https://arxiv.org/abs/2201.00365
tl;dr We create strong re-ranking and dense retrieval baselines (BERTCAT, BERTDOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the… See the full description on the dataset page: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training.Original-hybrid-correct-train-no-persona-meanOriginal-hybrid-train-no-persona-meanSC-train-valid-test_SDG-Descriptionsnli-label:
(0) entailment
(2) contradiction
vn-provinces-trained-labor-rate
Vietnam provinces trained labor rate (age 15+)
Share of the labour force aged 15+ that has received training (percent). Coverage 2008-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (1071 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-trained-labor-rate.isaac-gr00t-ikea-training-report
Isaac GR00T IKEA training report
Portable export of the W&B run unitree_g1_ikea_batch32_20260730. The run stopped after a clean host
shutdown; the last W&B metric is step 31,750 and the
last complete checkpoint is 31,000.
Summary
First logged training loss: 1.4667
Last logged training loss: 0.1256
Lowest 1,000-step rolling loss: 0.1249 at step 31,750
Mean GPU compute utilization: 51.8%
Mean allocated GPU memory: 29.0%
Provisional checkpoint choice: 30,000… See the full description on the dataset page: https://huggingface.co/datasets/ICRA-Competitions/isaac-gr00t-ikea-training-report.ga4-dataset-from-BQ-trainingset-2monthstrain_names_imbalanced
WA Voter Names — unbalanced train split
Training split for binary name classification, built from the Washington State voter
registration database (VRDB) extract dated 2026-09-01. Natural class prevalence.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.
Related repo
Contents
Kymera-Solutions/train_names_balanced
same positives, negatives downsampled 1:1
test split
not yet uploaded — required for evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_imbalanced.Train_routerArcNemotron-3-Nano-RL-Training-Blend-prompt-only
Nemotron-3-Nano-RL-Training-Blend-prompt-only
Prompt-only extraction from nvidia/Nemotron-3-Nano-RL-Training-Blend.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-3-Nano-RL-Training-Blend-prompt-only.bjj-kimura-lesson001-trial
BJJ Kimura from Side Control — Research Trial (sampled across the action arc)
Tier: Research / Evaluation (free, CC BY-NC-SA 4.0)
Source: RTK Motion Intelligence Platform · api.rtkmotion.io
Full commercial dataset: rtk-training/bjj-kimura-lesson001 (gated)
A temporally-sampled trial subset of a full 4D motion-capture session
of a Brazilian Jiu-Jitsu Kimura submission from side control, demonstrated
by a Former IBJJF World Champion (anonymized) with a training partner.… See the full description on the dataset page: https://huggingface.co/datasets/rtk-training/bjj-kimura-lesson001-trial.cyp-challenge-train-test
CYP Challenge Train/Test Dataset
A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge.
Blog post: Announcing OpenADMET’s CYP inhibition blind challenge
Challenge Space: OpenADMET CYP Inhibition Blind Challenge
Challenge period: August 17, 2026 - November 3, 2026
Produced by: OpenADMET
CHANGELOG
Updated… See the full description on the dataset page: https://huggingface.co/datasets/ks121/cyp-challenge-train-test.videophy_train_publicWe have uploaded the videos at: https://huggingface.co/videophysics/videophy-train-videos/tree/main
For more details, please visit the project github: https://github.com/Hritikbansal/videophy
pku-llama3.1-8b-answers-features-trainremote-ai-evaluation-training-market-snapshot
Dataset Description
This is an aggregate August 22, 2026 research snapshot from Specialist AI Work, an independent PatchMedia tracker of reviewed remote AI evaluation, AI training, data annotation-adjacent, and expert-review opportunities.
The live Specialist AI Work inventory has advanced since this snapshot. The counts in this repository describe the immutable August 22 research object; they are not a claim about today's inventory.
Reporting date: 2026-08-22
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/patchmedia-org/remote-ai-evaluation-training-market-snapshot.train_names_balanced
WA Voter Names — balanced train split
1:1 downsampled training split for binary name classification, built from the
Washington State voter registration database (VRDB) extract dated 2026-09-01.
Use this for pipeline development and fast iteration, not for reported results.
Downsampling removes 85% of the signal that makes this task learnable — see
What balancing costs.
Restricted data — see Access and legal restrictions.
This repository is not intended to be public.… See the full description on the dataset page: https://huggingface.co/datasets/Kymera-Solutions/train_names_balanced.IIIT_D1_traingit_good_bench-train
Dataset Summary
GitGoodBench Lite is a subset of 17469 samples for collecting trajectories of AI agents resolving git tasks (see Supported Scenarios) for model training purposes.
We support the programming languages Python, Java and Kotlin and the sample types merge conflict resolution and file-commit chain.
All data in this dataset are collected from 816 unique, open-source GitHub repositories with permissive licenses
that have >= 1000 stars, >= 5 branches, >= 10 contributors and… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains/git_good_bench-train.Train-Datasetseq_level_training_dataMicroG-HAR-train-ready
MicroG-HAR-train-ready dataset.
This dataset is converted from the MicroG-4M dataset.
For more information please:
Refer to our paper
Visit our GitHub
Datasets formatted in this way can be used directly as input data directories for the PySlowFast_for_HAR framework without additional preprocessing, enabling seamless model training and evaluation.
This dataset follows the organizational format of the AVA dataset. The only difference is that the original CSV header… See the full description on the dataset page: https://huggingface.co/datasets/lei-qi-233/MicroG-HAR-train-ready.
