datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dmi-aarhus-predictions
DMI Aarhus Predictions
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
predictions_latest.parquet
Current future + verified prediction store
dmi-collector
frontend_snapshot.json
Primary integration contract for the Vercel frontend
dmi-collector
Compatibility files
File
Status
Notes
predictions.parquet
Legacy
Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.predictionssmart-home-energy-prediction
Smart Home Appliance Energy Prediction
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
15,882
Labeled training data
test
3,853
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
date
object
lights
int64
T1
float64
RH_1
float64
T2
float64
RH_2
float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.Next_Token_Prediction_datasetMIMIC_YOLO_prediction_cxrmeowcat-predictions
MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC
Per-pixel cell-type predictions generated by MeowCat on
H&E whole-slide images from two public cohorts:
Cohort
Tissue
Samples
h5ad payload
TCGA-LUAD
Lung adenocarcinoma
531
~60 GB
CPTAC-CCRCC
Clear-cell renal cell carcinoma
831
~93 GB
File layout
composition.parquet # long format: sample × cell_type → count, fraction
metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.variant-effect-prediction
Updates
[2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons.
Abouts
The human reference genome data is sourced from the NCBI website.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.olist-ecommerce-for-delivery-and-review-prediction
E-Commerce Analytics for Delivery and Review Prediction
This dataset was created for a datathon project. It's a cleaned and feature-engineered version of the public Olist Brazilian E-commerce dataset, specifically prepared to predict shipping delays and customer review scores.
Project Goals
Our project focuses on two key business problems:
Model 1 (Regression): Can we predict how delayed a shipment will be? This helps manage customer expectations proactively.
Model 2… See the full description on the dataset page: https://huggingface.co/datasets/miminmoons/olist-ecommerce-for-delivery-and-review-prediction.ord_predictions
Dataset Details
Dataset Description
The open reaction database is a database of chemical reactions and their conditions
Curated by:
License: CC BY SA 4.0
Dataset Sources
original data source
Citation
BibTeX:
@article{Kearnes_2021,
doi = {10.1021/jacs.1c09820},
url = {https://doi.org/10.1021%2Fjacs.1c09820},
year = 2021,
month = {nov},
publisher = {American Chemical Society ({ACS})},
volume = {143},
number = {45},
pages =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/ord_predictions.fluorescence_prediction
Dataset Card for Fluorescence Prediction Dataset
Dataset Summary
The Fluorescence Prediction task focuses on predicting the fluorescence intensity of green fluorescent protein mutants, a crucial function in biology that allows researchers to infer the presence of proteins within cell lines and living organisms. This regression task utilizes training and evaluation datasets that feature mutants with three or fewer mutations, contrasting the testing dataset, which comprises… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/fluorescence_prediction.suicide_prediction_dataset_phr
Dataset Card for "vibhorag101/suicide_prediction_dataset_phr"
The dataset contains text with binary labels for suicide or non-suicide.
The dataset was cleaned and following steps were applied
Converted to lowercase
Removed numbers and special characters.
Removed URLs, Emojis and accented characters.
Removed any word contractions.
Remove any extra white spaces and any extra spaces after a single space.
Removed any consecutive characters repeated more than 3 times.
Tokenised the… See the full description on the dataset page: https://huggingface.co/datasets/vibhorag101/suicide_prediction_dataset_phr.scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP).
Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability).
ACL-OCL Extended
A dataset for citation count prediction only, based on the ACL-OCL dataset.
Extended with updated citation counts, references and annotated research hypothesis.
OpenReview (Last Update: 1.1.2025)
A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.kalshi-prediction-markets-markets
Kalshi Prediction Markets — Markets Metadata & Quotes
Per-market snapshot data for Kalshi markets spanning Aug 2023–Aug 2025.Includes tickers, lifecycle timestamps, status/result fields, rules, liquidity/open interest, and top-of-book quotes (YES/NO bid/ask + last/previous prices) with integer and dollar-scaled variants.
Rows: ~10,016 markets (single split)Schema stability: stablePrivacy: public market metadata
Dataset Structure
Split
train — one… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-markets.link_predictionkalshi-prediction-markets-betting
Kalshi Prediction Markets — Trades
High-volume trade-level data from Kalshi prediction markets spanning Aug 2023–Aug 2025, suitable for market microstructure, liquidity, and price-impact analysis.
Rows: ~5.08M trades (single split)Schema stability: stablePrivacy: public market data, no PII
Dataset Structure
Split
train — all trades
Features (columns)
name
dtype
description
trade_id
string
Unique trade identifier
ticker
string… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-betting.Driving-Hazard-Prediction-and-Reasoning
Exploring the Potential of Multi-Modal AI for Driving Hazard Prediction
DHPR: Driving Hazard Prediction and Reasoning
Paper
fold_prediction
Dataset Card for Fold Prediction Dataset
Dataset Summary
Fold class prediction is a scientific classification task that assigns protein sequences to one of 1,195 known folds. The primary application of this task lies in the identification of novel remote homologs among proteins of interest, such as emerging antibiotic-resistant genes and industrial enzymes. The study of protein fold holds great significance in fields like proteomics and structural biology, as it… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/fold_prediction.nft_prediction_all_NFTs
Dataset Card for "nft_prediction_all_NFTs"
More Information needed
stability_prediction
Dataset Card for Stability Stability Dataset
Dataset Summary
The Stability Stability task is to predict the concentration of protease at which a protein can retain its folded state. Protease, being integral to numerous biological processes, bears significant relevance and a profound comprehension of protein stability during protease interaction can offer immense value, especially in the creation of novel therapeutics.
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/stability_prediction.Driving-Hazard-Prediction-and-Reasoning
Exploring the Potential of Multi-Modal AI for Driving Hazard Prediction
DHPR: Driving Hazard Prediction and Reasoning
Paper
Github
Forked from original dataset and updated with correct licence and
metadata since original dataset was unmaintained since publication.
Released under BSD-3 Clause Licence.
nft_prediction_all_with_dates
Dataset Card for "nft_prediction_all_with_dates"
More Information needed
predictionsbiomap-research-localization_prediction
localization_prediction
Sourced from biomap-research/localization_prediction and prepared for Hugging Face datasets usage.
Data files
Parquet files are stored under data/ using Hugging Face split naming conventions
(train-*, validation-*, test-*).
Preparation
Preprocess mode: minimal.
Seed: 1957723.
No max sequence length filter was applied.
Renamed source columns: label -> targets, seq -> sequence.
Columns: id, sequence, targets, split.
Validation… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-localization_prediction.article-bias-prediction-media-splits
News Articles with Political Bias Annotations (Media Source Split)
Source
Derived from Baly et al.'s work:
We Can Detect Your Bias: Predicting the Political Ideology of News Articles (Baly et al., EMNLP 2020)
Information
This dataset contains 34,737 news articles manually annotated for political ideology, either "left", "center", or "right".
This version contains media source test/training/validation splits, where the articles in each split are
from different… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/article-bias-prediction-media-splits.tape-proteinnet_contact_prediction
tape-proteinnet-contact-prediction
Sourced from TAPE ProteinNet raw JSON archive: http://s3.amazonaws.com/songlabdata/proteindata/data_pytorch/proteinnet.tar.gz.
Data files
Parquet files are stored under data/ with Hugging Face split names. TAPE valid is published as validation; train_unfiltered is omitted by default and included only when explicitly configured.
Contact definition
Positive contacts are upper-triangle residue pairs (including… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/tape-proteinnet_contact_prediction.biomap-research-contact_prediction_binary
contact_prediction_binary
Sourced from biomap-research/contact_prediction_binary and prepared for Hugging Face datasets usage.
Data files
Parquet files are stored under data/ using Hugging Face split naming conventions
(train-*, validation-*, test-*).
Preparation
Preprocess mode: minimal.
Seed: 1957723.
No max sequence length filter was applied.
Renamed source columns: label -> targets, seq -> sequence.
Columns: id, sequence, targets, split.… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-contact_prediction_binary.fluorescence_prediction
Dataset Card for Fluorescence Prediction Dataset
Dataset Summary
The Fluorescence Prediction task focuses on predicting the fluorescence intensity of green fluorescent protein mutants, a crucial function in biology that allows researchers to infer the presence of proteins within cell lines and living organisms. This regression task utilizes training and evaluation datasets that feature mutants with three or fewer mutations, contrasting the testing dataset, which… See the full description on the dataset page: https://huggingface.co/datasets/chenchaozhao/fluorescence_prediction.task963_librispeech_asr_next_word_prediction
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task963_librispeech_asr_next_word_prediction
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task963_librispeech_asr_next_word_prediction.link_predictionmotive-v2-prediction-rankings
MOTIVE v2 Gene-Compound Prediction Rankings
This Hugging Face repository is the browsable Data Studio mirror of the two Parquet artifacts in Zenodo record 22105202.
Zenodo is the canonical source for citation, provenance, methods, versioning, file integrity, and detailed interpretation.
Exact version DOI: 10.5281/zenodo.22105202
Scientific context and limitations: MOTIVE Issue 12
Paper: MOTIVE: A Drug-Target Interaction Graph For Inductive Link Prediction
Browse… See the full description on the dataset page: https://huggingface.co/datasets/carpenter-singh-lab/motive-v2-prediction-rankings.
