datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fx-bank-consensus-forecasts
FX Bank Forecast — Cross-Firm FX Consensus & Dispersion (2026)
Aggregated 2026 year-end foreign-exchange forecasts from 18 investment banks,
across 19 G10 and EM currencies — the cross-firm consensus (mean & median),
the range (low/high), and the dispersion (how far apart the banks are).
This is the aggregate view only (mean / median / range / dispersion / firm
count). Per-firm named targets are not included here.
Source: FX Bank Forecast ·
Live, daily-updated version:… See the full description on the dataset page: https://huggingface.co/datasets/fxbankforecast/fx-bank-consensus-forecasts.surya-bench-flare-forecasting
Full-disk Solar Flare Forecasting Dataset
Dataset Summary
This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index.
Two derived binary labels are included for forecasting tasks:
label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.store-sales-time-series-forecasting
taken from this Kaggle competition:
Dataset Description
In this competition, you will predict sales for the thousands of product families sold at Favorita stores located in Ecuador. The training data includes dates, store and product information, whether that item was being promoted, as well as the sales numbers. Additional files include supplementary information that may be useful in building your models.
File Descriptions and Data Field Information… See the full description on the dataset page: https://huggingface.co/datasets/mrcksggcfc/store-sales-time-series-forecasting.forecastability_classificationThis dataset is composed of Claude-labelled fineweb documents.
For each document, Claude is asked if it is 'forecastable' (i.e. would be a reasonable seed for a pastcasting question) and to estimate the date the document was published.
V1 splits were generated by having Claude label ~50K random fineweb documents and v2 splits were augmented with labels on ~30K additional documents that a DebertaV3 classifier finetuned on ratio10_v1 thought were forecastable (Claude thought ~1/3 of these… See the full description on the dataset page: https://huggingface.co/datasets/noanabeshima/forecastability_classification.ipulse-ai-batch5-advisor-forecast-panel
iPulse AI Batch 5 Advisor Forecast Panel
This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters.
The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.multimodal-time-series-forecastinghistorical-us-station-weather-forecasts
Historical US station weather forecasts and observations
December 4, 2025–July 12, 2026 · 7,514,981 forecast records · 882,212 observation records · 14 stations · 19 forecast models
Historical weather data, packaged as one complete SQLite database per station. Each file contains the stored forecast runs, full observation history, daily highs and lows, reported high/low evidence, collection coverage and field descriptions. Weather fields are stored in ordinary columns, without… See the full description on the dataset page: https://huggingface.co/datasets/whodisidk/historical-us-station-weather-forecasts.preference-forecast
HorizonBench Human Longitudinal Preference Dataset
Release Status
Version 1.0.0 is available through gated research access. The structured tables and sanitized participant-authored text pass the documented release checks. Manual review covered every high-priority passage, a fixed random sample of 600 medium-priority passages, and 155 additional medium-priority passages. Text outside the manual sample received the same automated sanitization applied to every row.… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/preference-forecast.weather-forecasting-challenge
Dataset Description
Data Overview
The WiDS Datathon 2023 focuses on a prediction task involving forecasting sub-seasonal temperatures (temperatures over a two-week period, in our case) within the United States. We are using a pre-prepared dataset consisting of weather and climate information for a number of US locations, for a number of start dates for the two-week observation, as well as the forecasted temperature and precipitation from a number of weather… See the full description on the dataset page: https://huggingface.co/datasets/serenia-science/weather-forecasting-challenge.energy-forecasting-filessuperkart-sales-forecast
SuperKart Sales Forecast (Tabular)
This dataset contains product/store-level attributes with the target Product_Store_Sales_Total for supervised learning and forecasting.
Files
data/SuperKart.csv
Schema
Product_Id — Unique identifier of each product (AA… pattern)
Product_Weight — Weight (kg)
Product_Sugar_Content — low sugar / regular / no sugar
Product_Allocated_Area — Ratio of display area allocated to the product
Product_Type — Category (meat, snacks… See the full description on the dataset page: https://huggingface.co/datasets/imambru/superkart-sales-forecast.wood-market-forecast-tournamentsales-forecast-dataset
SuperKart Sales Dataset
This dataset supports a sales prediction pipeline (Product × Store).
Source file: raw/SuperKart.csv
Target: Product_Store_Sales_Total
Expected columns:
Product_Id, Product_Weight, Product_Sugar_Content, Product_Allocated_Area, Product_Type, Product_MRP,
Store_Id, Store_Establishment_Year, Store_Size, Store_Location_City_Type, Store_Type, Product_Store_Sales_Total
forecast-scorecard
AI Energy-Demand Forecast Scorecard
A reproducible audit of how the field forecasts data-centre electricity demand: how the published forecasts disperse, how they get revised, and whether they are transparent enough to reproduce. Primary-sourced, published with the data and a script that regenerates every figure.
AI disclosure: the research is the author's; this text was drafted with AI assistance and reviewed by the author. The model, and the conflict it creates, are named in… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/forecast-scorecard.Sales-Forecast-Predictionsales-forecast-dataset
SuperKart Sales Dataset
This dataset supports a sales prediction pipeline (Product × Store).
Source file: raw/SuperKart.csv
Target: Product_Store_Sales_Total
Expected columns:
Product_Id, Product_Weight, Product_Sugar_Content, Product_Allocated_Area, Product_Type, Product_MRP,
Store_Id, Store_Establishment_Year, Store_Size, Store_Location_City_Type, Store_Type, Product_Store_Sales_Total
power-grid-renewable-output-forecast-coherence-risk-v0.1What this repo is for
Detect grid stress from renewable forecast error.
Focus
forecast vs actual output
ramp variability
backup readiness
storage margin
Why it matters
Forecast error plus no backup creates grid instability fast.
forecastability_classification_oldafrica-synth-energy-oilgas-production-forecasts-nigeria
Africa Synth Energy Oilgas Production Forecasts Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-production-forecasts-nigeria.Weather_Forecast_Dataset
Weather Forecast Dataset
This dataset is ideal for beginners looking to practice machine learning, specifically using classification techniques.
With 2,500 weather observations, it’s a simple yet practical dataset for learning how to predict rainfall based on various weather conditions.
Perfect for use with Python libraries like scikit-learn, this dataset enables experimentation with algorithms such as logistic regression , decision trees , and random forests .
The dataset's… See the full description on the dataset page: https://huggingface.co/datasets/zeeshier/Weather_Forecast_Dataset.SuperKart-Sales-Forecast-Predictorgrocery-sales-forecastingcip-forecasting-Macroeconomic_Indicators_datasetmeta-realworld-forecast-analysis-coherence-test-apms-v0.3
Anthropic Post-Mortem Simulator
RealWorld Forecast Analysis Coherence Test
Meta Cognitive Hygiene Dataset v0.3
Purpose
This dataset tests whether a model can do four things in sequence.
Compute a simple real-world rate vs lab rate
Name plausible failure categories
Detect confounds and internal contradictions in the input
Admit when the data is not sufficient for a definitive diagnosis
This is a coherence audit, not a math quiz.
What… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/meta-realworld-forecast-analysis-coherence-test-apms-v0.3.forecasts-vs-market
NeuPortal Forecast Experiment — AI vs Prediction Market (Bitcoin-Timestamped)
A small but unusual dataset: every row is a football match forecast that was provably made before kickoff — locked, hashed into Bitcoin via OpenTimestamps, and benchmarked against the prediction-market price frozen at the same instant. Nothing here can have been backdated, by anyone, including us.
This is the raw data behind the live public experiment at neuportal.ai/experiment. The cryptographic… See the full description on the dataset page: https://huggingface.co/datasets/neuportal/forecasts-vs-market.Store-Sales-Time-Series-Forecasting-result-0.45607walmart-demand-forecast
Walmart Demand Forecast Dataset
Overview
This dataset contains demand forecasting results generated using historical Walmart sales data.
The goal of this project is to predict future product demand using time series and machine learning techniques.
Models Used
ARIMA
Prophet
LSTM (Deep Learning)
Dataset Description
The dataset includes forecasted weekly sales values based on historical trends.
It is suitable for:
Demand forecasting
Time series… See the full description on the dataset page: https://huggingface.co/datasets/moviebrain01/walmart-demand-forecast.IN5000-MB-TUD-Forecasting
A Framework for Identifying Evolution Patterns of Open-Source Software Projects
IN5000 TU Delft - MSc Computer Science
This repository contains the multivariate time series forecasting datasets related to the following repository:
https://github.com/IN5000-MB-TUD/data-analysis
Contributors
Project developed for the course IN5000 - Master's thesis of the 2023/2024 academic year at TU Delft.
Author:
Mattia Bonfanti
m.bonfanti@student.tudelft.nl
Master's in Computer… See the full description on the dataset page: https://huggingface.co/datasets/MattiaBonfanti-CS/IN5000-MB-TUD-Forecasting.llm-forecast-bench
LLM Forecast Bench — persona vs neutral prompting on a real-money prediction market
A small, fully-reproducible dataset for one question: does wrapping a frontier
LLM in a bespoke "trading persona" prompt change its forecast accuracy versus
running the same model on a neutral prompt — when both stake real money on
on-chain prediction markets and are scored against the live market price?
Every agent here is a real LLM that placed real (USDC) positions on
FlipCoin markets on Base.… See the full description on the dataset page: https://huggingface.co/datasets/Flipcoin/llm-forecast-bench.SupplyChain_Risk_Prediction_Forecast
