CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes34k downloads3mo agoHugging Face02nexar-ai /nexar_collision_predictiongated Nexar Collision Prediction Dataset This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle. Dataset The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.tabularvideo-classification1K<n<10K19 likes15k downloads1y agoHugging Face03THEORACLEEEE /polymarket-predictions THE ORACLE — Polymarket predictions Live predictions for Polymarket markets, produced by THE ORACLE — an autonomous agent funded by $ORACLE pump.fun creator fees. Each row is a baseline-model forecast over live orderbook signals (momentum, microstructure, liquidity). predictions.json / predictions.csv — 100 markets, refreshed each agent cycle. Columns: question, category, market_prob, oracle_prob, edge, confidence, signal, model, backtest_acc, auc, modelability, volume… See the full description on the dataset page: https://huggingface.co/datasets/THEORACLEEEE/polymarket-predictions.tabularn<1K0 likes11k downloads9m agoHugging Face04scikit-learn /churn-predictionCustomer churn prediction dataset of a fictional telecommunication company made by IBM Sample Datasets. Context Predict behavior to retain customers. You can analyze all relevant customer data and develop focused customer retention programs. Content Each row represents a customer, each column contains customer’s attributes described on the column metadata. The data set includes information about: Customers who left within the last month: the column is called Churn Services that each customer… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/churn-prediction.tabular1K<n<10K20 likes5.7k downloads4y agoHugging Face05Ciroc0 /dmi-aarhus-predictions DMI Aarhus Predictions Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by predictions_latest.parquet Current future + verified prediction store dmi-collector frontend_snapshot.json Primary integration contract for the Vercel frontend dmi-collector Compatibility files File Status Notes predictions.parquet Legacy Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.tabular1K<n<10K6 likes5.3k downloads24m agoHugging Face06victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes5k downloads4y agoHugging Face07orcn /predictionsimage100K<n<1M0 likes1.8k downloads1y agoHugging Face08hoangbang /smart-home-energy-prediction Smart Home Appliance Energy Prediction Dataset Summary A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded. Splits Split Examples Description train 15,882 Labeled training data test 3,853 Public inputs with withheld target labels or annotations Data Fields Field Type date object lights int64 T1 float64 RH_1 float64 T2 float64 RH_2 float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.tabulartabular-regression10K<n<100K0 likes1.3k downloads2mo agoHugging Face09kilian-group /phantom-wiki-v0-5-0-predictions Dataset Card for Dataset Name Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.tabular100K<n<1M0 likes962 downloads2y agoHugging Face10Na0s /Next_Token_Prediction_datasettext1M<n<10M0 likes615 downloads2y agoHugging Face11metaeval /acceptability-prediction@inproceedings{lau-etal-2015-unsupervised, title = "Unsupervised Prediction of Acceptability Judgements", author = "Lau, Jey Han and Clark, Alexander and Lappin, Shalom", booktitle = "Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)", month = jul, year = "2015", address = "Beijing, China", publisher = "Association for… See the full description on the dataset page: https://huggingface.co/datasets/metaeval/acceptability-prediction.tabulartext-classification1K<n<10K1 likes609 downloads4y agoHugging Face12Rookiezz /MIMIC_YOLO_prediction_cxrimage10K<n<100K1 likes556 downloads10mo agoHugging Face13liranmao /meowcat-predictions MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC Per-pixel cell-type predictions generated by MeowCat on H&E whole-slide images from two public cohorts: Cohort Tissue Samples h5ad payload TCGA-LUAD Lung adenocarcinoma 531 ~60 GB CPTAC-CCRCC Clear-cell renal cell carcinoma 831 ~93 GB File layout composition.parquet # long format: sample × cell_type → count, fraction metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.tabularimage-classification10K<n<100K0 likes556 downloads15d agoHugging Face14proteinea /secondary_structure_predictiontext10K<n<100K4 likes547 downloads4y agoHugging Face15haja17 /real-or-fake-fake-jobposting-predictiontabular10K<n<100K0 likes530 downloads7mo agoHugging Face16miminmoons /olist-ecommerce-for-delivery-and-review-prediction E-Commerce Analytics for Delivery and Review Prediction This dataset was created for a datathon project. It's a cleaned and feature-engineered version of the public Olist Brazilian E-commerce dataset, specifically prepared to predict shipping delays and customer review scores. Project Goals Our project focuses on two key business problems: Model 1 (Regression): Can we predict how delayed a shipment will be? This helps manage customer expectations proactively. Model 2… See the full description on the dataset page: https://huggingface.co/datasets/miminmoons/olist-ecommerce-for-delivery-and-review-prediction.tabular100K<n<1M3 likes447 downloads1y agoHugging Face17GenerTeam /variant-effect-prediction Updates [2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons. Abouts The human reference genome data is sourced from the NCBI website. We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis. How to use from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.tabularzero-shot-classification10K<n<100K0 likes442 downloads5mo agoHugging Face18jablonkagroup /ord_predictions Dataset Details Dataset Description The open reaction database is a database of chemical reactions and their conditions Curated by: License: CC BY SA 4.0 Dataset Sources original data source Citation BibTeX: @article{Kearnes_2021, doi = {10.1021/jacs.1c09820}, url = {https://doi.org/10.1021%2Fjacs.1c09820}, year = 2021, month = {nov}, publisher = {American Chemical Society ({ACS})}, volume = {143}, number = {45}, pages =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/ord_predictions.tabular10M<n<100M0 likes407 downloads1y agoHugging Face19proteinglm /fluorescence_prediction Dataset Card for Fluorescence Prediction Dataset Dataset Summary The Fluorescence Prediction task focuses on predicting the fluorescence intensity of green fluorescent protein mutants, a crucial function in biology that allows researchers to infer the presence of proteins within cell lines and living organisms. This regression task utilizes training and evaluation datasets that feature mutants with three or fewer mutations, contrasting the testing dataset, which comprises… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/fluorescence_prediction.texttext-classification10K<n<100K0 likes405 downloads2y agoHugging Face20bryandts /robot-action-prediction-dataset Robotic Action Prediction Dataset Dataset Description This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions. Dataset Structure Data Fields current_frame: Input image (RGB) of the current observation instruction: Textual description of the action to perform future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.imageimage-to-image10K<n<100K0 likes382 downloads1y agoHugging Face21vibhorag101 /suicide_prediction_dataset_phr Dataset Card for "vibhorag101/suicide_prediction_dataset_phr" The dataset contains text with binary labels for suicide or non-suicide. The dataset was cleaned and following steps were applied Converted to lowercase Removed numbers and special characters. Removed URLs, Emojis and accented characters. Removed any word contractions. Remove any extra white spaces and any extra spaces after a single space. Removed any consecutive characters repeated more than 3 times. Tokenised the… See the full description on the dataset page: https://huggingface.co/datasets/vibhorag101/suicide_prediction_dataset_phr.texttext-classification100K<n<1M4 likes332 downloads1y agoHugging Face22thomaswmitch /kalshi-prediction-markets-markets Kalshi Prediction Markets — Markets Metadata & Quotes Per-market snapshot data for Kalshi markets spanning Aug 2023–Aug 2025.Includes tickers, lifecycle timestamps, status/result fields, rules, liquidity/open interest, and top-of-book quotes (YES/NO bid/ask + last/previous prices) with integer and dollar-scaled variants. Rows: ~10,016 markets (single split)Schema stability: stablePrivacy: public market metadata Dataset Structure Split train — one… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-markets.tabularother10K<n<100K0 likes323 downloads1y agoHugging Face23nehalvatss /real-or-fake-fake-jobposting-predictiontabular10K<n<100K0 likes309 downloads3mo agoHugging Face24nhop /scientific-quality-score-predictionDatasets related to the task of Scholarly Document Quality Prediction (SDQP). Each sample is an academic paper for which either the citation count or the review score can be predicted (depending on availability). ACL-OCL Extended A dataset for citation count prediction only, based on the ACL-OCL dataset. Extended with updated citation counts, references and annotated research hypothesis. OpenReview (Last Update: 1.1.2025) A dataset for review score and citation count… See the full description on the dataset page: https://huggingface.co/datasets/nhop/scientific-quality-score-prediction.tabulartext-classification100K<n<1M0 likes299 downloads1y agoHugging Face25CharlesLi /link_predictiontext10K<n<100K5 likes281 downloads1y agoHugging Face26genbio-ai /transcript_isoform_expression_prediction Multi-modal transcript isoform expression dataset We curated the human transcript isoform expression dataset from the GTEx portal following the preprocessing pipeline in Garau-Luis et al. (2024). We downloaded the RNA-seq Transcript TPMs file from the bulk tissue expression in GTEx Analysis V8. The table contains transcript expression collected from 30 non-diseased tissues in nearly 1000 human individuals. We averaged the transcript expression measurements across individuals to… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/transcript_isoform_expression_prediction.tabular100K<n<1M0 likes270 downloads2mo agoHugging Face27thomaswmitch /kalshi-prediction-markets-betting Kalshi Prediction Markets — Trades High-volume trade-level data from Kalshi prediction markets spanning Aug 2023–Aug 2025, suitable for market microstructure, liquidity, and price-impact analysis. Rows: ~5.08M trades (single split)Schema stability: stablePrivacy: public market data, no PII Dataset Structure Split train — all trades Features (columns) name dtype description trade_id string Unique trade identifier ticker string… See the full description on the dataset page: https://huggingface.co/datasets/thomaswmitch/kalshi-prediction-markets-betting.tabularother1M<n<10M4 likes268 downloads1y agoHugging Face28DHPR /Driving-Hazard-Prediction-and-Reasoning Exploring the Potential of Multi-Modal AI for Driving Hazard Prediction DHPR: Driving Hazard Prediction and Reasoning Paper image10K<n<100K11 likes260 downloads2y agoHugging Face29impresso-project /ner-eval-predictionstabular100K<n<1M0 likes256 downloads2mo agoHugging Face30Aurthor /fake_job_post_predictiontabular10K<n<100K2 likes253 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.