CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes33k downloads3mo agoHugging Face02nexar-ai /nexar_collision_prediction Nexar Collision Prediction Dataset This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle. Dataset The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows regular… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.tabularvideo-classification1K<n<10K19 likes15k downloads1y agoHugging Face03THEORACLEEEE /polymarket-predictions THE ORACLE — Polymarket predictions Live predictions for Polymarket markets, produced by THE ORACLE — an autonomous agent funded by $ORACLE pump.fun creator fees. Each row is a baseline-model forecast over live orderbook signals (momentum, microstructure, liquidity). predictions.json / predictions.csv — 100 markets, refreshed each agent cycle. Columns: question, category, market_prob, oracle_prob, edge, confidence, signal, model, backtest_acc, auc, modelability, volume… See the full description on the dataset page: https://huggingface.co/datasets/THEORACLEEEE/polymarket-predictions.tabularn<1K0 likes10k downloads14m agoHugging Face04scikit-learn /churn-predictionCustomer churn prediction dataset of a fictional telecommunication company made by IBM Sample Datasets. Context Predict behavior to retain customers. You can analyze all relevant customer data and develop focused customer retention programs. Content Each row represents a customer, each column contains customer’s attributes described on the column metadata. The data set includes information about: Customers who left within the last month: the column is called Churn Services that each customer… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/churn-prediction.tabular1K<n<10K20 likes5.5k downloads4y agoHugging Face05Ciroc0 /dmi-aarhus-predictions DMI Aarhus Predictions Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by predictions_latest.parquet Current future + verified prediction store dmi-collector frontend_snapshot.json Primary integration contract for the Vercel frontend dmi-collector Compatibility files File Status Notes predictions.parquet Legacy Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.tabular1K<n<10K6 likes5.1k downloads1h agoHugging Face06victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes5k downloads4y agoHugging Face07fassabilf /sea-clip-eval-predictions0 likes2.9k downloads3mo agoHugging Face08onef1shy /Crop-Yield-Prediction-MODIS Crop Yield Prediction MODIS This repository hosts the processed MODIS data used for crop-yield regression in the DFYP project. It was prepared from the MODIS branch of the DFYP project. Repository: https://github.com/onef1shy/DFYP. Paper: https://doi.org/10.1109/TGRS.2026.3684831 Contents datasets/modis/processed_data/<year>/*.npy: preprocessed yearly samples indexed by year, county, and sample id datasets/modis/processed_data/histogram_all_full.npz: histogram data… See the full description on the dataset page: https://huggingface.co/datasets/onef1shy/Crop-Yield-Prediction-MODIS.geospatial0 likes2.3k downloads5mo agoHugging Face09orcn /predictionsimage100K<n<1M0 likes1.7k downloads1y agoHugging Face10pdebench-fno-audit /fno-predictions PDEBench FNO Re-evaluation: Prediction Tensors Test-set prediction arrays from The Unrealized Potential of Fourier Neural Operators: A Systematic Re-evaluation of PDEBench Baselines (NeurIPS 2026 E&D Track submission). File layout For all standard tests (1-27, 29, plus the three supplementary 2D CFD configurations), each .npz file contains: preds: model predictions, shape [N_test, spatial_dims..., T, nc] targets: ground truth, same shape per_sample: per-sample… See the full description on the dataset page: https://huggingface.co/datasets/pdebench-fno-audit/fno-predictions.0 likes1.7k downloads5mo agoHugging Face11aroon-sankoh /western-us-wildfire-prediction Curated Wildfire Detection & Analysis Dataset This dataset consists of 125 fire and 375 control 'scenes', where every scene includes a Sentinel-1 pre image, Sentinel-1 post image, Sentinel-2 pre image, Sentinel-2 post image, ERA5 re-analysis data series, and a json file with metadata on each piece of data. Each fire is paired with three controls that match the fires EPA Level III Eco-region of the fire. 125 fires and 375 control scenes are collected over seven United States… See the full description on the dataset page: https://huggingface.co/datasets/aroon-sankoh/western-us-wildfire-prediction.0 likes1.5k downloads16d agoHugging Face12lwaekfjlk /prediction-market-news0 likes1.3k downloads8mo agoHugging Face13hoangbang /smart-home-energy-prediction Smart Home Appliance Energy Prediction Dataset Summary A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded. Splits Split Examples Description train 15,882 Labeled training data test 3,853 Public inputs with withheld target labels or annotations Data Fields Field Type date object lights int64 T1 float64 RH_1 float64 T2 float64 RH_2 float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/smart-home-energy-prediction.tabulartabular-regression10K<n<100K0 likes1.2k downloads2mo agoHugging Face14kilian-group /phantom-wiki-v0-5-0-predictions Dataset Card for Dataset Name Predictions from https://huggingface.co/datasets/mlcore/phantom-wiki-v050 Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/kilian-group/phantom-wiki-v0-5-0-predictions.tabular100K<n<1M0 likes949 downloads2y agoHugging Face15bernardo-de-almeida /segmentnt_predictionsLabels and predictions of SegmentNT-30kb human model. 0 likes682 downloads2y agoHugging Face16saeedrmd /trajectory-prediction-argoverse21 likes650 downloads8mo agoHugging Face17Na0s /Next_Token_Prediction_datasettext1M<n<10M0 likes619 downloads2y agoHugging Face18metaeval /acceptability-prediction@inproceedings{lau-etal-2015-unsupervised, title = "Unsupervised Prediction of Acceptability Judgements", author = "Lau, Jey Han and Clark, Alexander and Lappin, Shalom", booktitle = "Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)", month = jul, year = "2015", address = "Beijing, China", publisher = "Association for… See the full description on the dataset page: https://huggingface.co/datasets/metaeval/acceptability-prediction.tabulartext-classification1K<n<10K1 likes607 downloads4y agoHugging Face19haja17 /real-or-fake-fake-jobposting-predictiontabular10K<n<100K0 likes582 downloads7mo agoHugging Face20prediction2 /diffroute_exp diffroute_exp Diffroute paper experiment data. 0 likes557 downloads10mo agoHugging Face21liranmao /meowcat-predictions MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC Per-pixel cell-type predictions generated by MeowCat on H&E whole-slide images from two public cohorts: Cohort Tissue Samples h5ad payload TCGA-LUAD Lung adenocarcinoma 531 ~60 GB CPTAC-CCRCC Clear-cell renal cell carcinoma 831 ~93 GB File layout composition.parquet # long format: sample × cell_type → count, fraction metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.tabularimage-classification10K<n<100K0 likes551 downloads14d agoHugging Face22proteinea /secondary_structure_predictiontext10K<n<100K4 likes547 downloads4y agoHugging Face23Rookiezz /MIMIC_YOLO_prediction_cxrimage10K<n<100K1 likes541 downloads10mo agoHugging Face24saeedrmd /trajectory-prediction-nuscenes1 likes540 downloads8mo agoHugging Face25AURORAData /AURORA_predictionAURORA (Architecture Unveiling through RNA Omics and Routine Histology Analysis) trained_models: Pretrained AURORA models for LUAD, KIRC and BRCA. *.pth: model weight; *.json: parameters for the AURORA model; *.csv: supporting information (cell types, gene names and normalizing factors) used by *.json. predictions_112um: Virtual spatial transcriptomics at 112 μm * 112 μm by AURORA of TCGA-LUAD, TCGA-KIRC, TCGA-BRCA and BRCA pre-chemotherapy (https://doi.org/10.1038/s41586-021-04278-5)… See the full description on the dataset page: https://huggingface.co/datasets/AURORAData/AURORA_prediction.0 likes458 downloads4h agoHugging Face26rcds /swiss_judgment_predictionSwiss-Judgment-Prediction is a multilingual, diachronic dataset of 85K Swiss Federal Supreme Court (FSCS) cases annotated with the respective binarized judgment outcome (approval/dismissal), posing a challenging text classification task. We also provide additional metadata, i.e., the publication year, the legal area and the canton of origin per case, to promote robustness and fairness studies on the critical area of legal NLP.text-classification10K<n<100K18 likes452 downloads3y agoHugging Face27miminmoons /olist-ecommerce-for-delivery-and-review-prediction E-Commerce Analytics for Delivery and Review Prediction This dataset was created for a datathon project. It's a cleaned and feature-engineered version of the public Olist Brazilian E-commerce dataset, specifically prepared to predict shipping delays and customer review scores. Project Goals Our project focuses on two key business problems: Model 1 (Regression): Can we predict how delayed a shipment will be? This helps manage customer expectations proactively. Model 2… See the full description on the dataset page: https://huggingface.co/datasets/miminmoons/olist-ecommerce-for-delivery-and-review-prediction.tabular100K<n<1M3 likes442 downloads1y agoHugging Face28GenerTeam /variant-effect-prediction Updates [2025-09-09] We have added ClinVar variant effect prediction results to the repository. The evaluation dataset was sourced from SongLab. The benchmark includes comparisons of GENERator against Evo2, NT, NT-v2, HyenaDNA, GPN-MSA, CADD, phyloP, and phastCons. Abouts The human reference genome data is sourced from the NCBI website. We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis. How to use from datasets… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/variant-effect-prediction.tabularzero-shot-classification10K<n<100K0 likes425 downloads5mo agoHugging Face29pytorch-lifestream /age-group-predictionhttps://ods.ai/competitions/sberbank-sirius-lesson tabulartabular-classification10M<n<100M0 likes419 downloads3y agoHugging Face30proteinglm /fluorescence_prediction Dataset Card for Fluorescence Prediction Dataset Dataset Summary The Fluorescence Prediction task focuses on predicting the fluorescence intensity of green fluorescent protein mutants, a crucial function in biology that allows researchers to infer the presence of proteins within cell lines and living organisms. This regression task utilizes training and evaluation datasets that feature mutants with three or fewer mutations, contrasting the testing dataset, which comprises… See the full description on the dataset page: https://huggingface.co/datasets/proteinglm/fluorescence_prediction.texttext-classification10K<n<100K0 likes404 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.