datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forecast-news
Forecast News
Deduplicated daily news corpus used by forecast-sim and future-sim.
Snapshot
31,859,020 articles
3,463 daily partitions
Coverage: 2016-08-26 through 2026-08-31
Snapshot published: 2026-09-18
Stored data size: approximately 158.5 GiB
The Parquet files are the canonical complete representation. The repository also
contains daily JSONL files where available and compact headline JSON files used
by article-browsing workflows.
Layout
Files… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news.retail-forecast-optimize-benchmark
Retail Forecast-Optimize Benchmark
Benchmark dataset for predict-then-optimize retail inventory decisions.
Contents
File pattern
Description
{scenario}_s{seed}.json
Full policy comparison per scenario instance
samples/sample_*.json
Retail scenario definitions
eval_results.json
Aggregated leaderboard metrics
chronos-2-live-forecasts/
Live Chronos-2 outputs from HF Jobs (when available)
Scenarios
Baseline Operations
Promotion… See the full description on the dataset page: https://huggingface.co/datasets/aniketraj0224/retail-forecast-optimize-benchmark.forecasting_rawRaw Dataset from "Approaching Human-Level Forecasting with Language Models"
This documentation provides an overview of the raw dataset utilized in our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset originates from forecasting platforms such as Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms engage users in predicting the… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting_raw.forecast-workflow-bench
Forecast Workflow Bench (FWBench)
1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions
with budgeted forecasting tools.
Code ·
Leaderboard
Evaluation protocol
The Language Model evaluation covers
1,251 cases, two hosted and eight local configurations, and 22,518 conversations
including paired local TSFM-removal runs.
Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the
minimum feasible loss with known target demand… See the full description on the dataset page: https://huggingface.co/datasets/Neurogica/forecast-workflow-bench.forecastingDataset from "Approaching Human-Level Forecasting with Language Models"
This document details the curated dataset developed for our research paper, Approaching Human-Level Forecasting with Language Models, authored by Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt.
Data Source and Format
The dataset is compiled from forecasting platforms including Metaculus, Good Judgment Open, INFER, Polymarket, and Manifold. These platforms enable users to predict future events by assigning… See the full description on the dataset page: https://huggingface.co/datasets/YuehHanChen/forecasting.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.pm25-forecasting-data
PM2.5 Forecasting Demo Data
This dataset stores the lightweight artifacts used by the Hugging Face Space demo:
https://huggingface.co/spaces/sumit1703/pm25-forecasting
The files are precomputed outputs from the ANRF AISEHack Phase 2 Theme 2 Pollution Forecasting project. The Space uses these artifacts only for visualization. It does not run live model inference, training, or torch at runtime.
Files
File
Description
demo_preds.npy
Precomputed PM2.5… See the full description on the dataset page: https://huggingface.co/datasets/sumit1703/pm25-forecasting-data.llm-forecast-calibration
LLM Forecast Calibration Study — GLM-5.3 on resolved Manifold Markets questions
Raw generation data for the study "Does sampling K times beat thinking harder?
A controlled study of LLM forecast calibration on resolved binary questions."
Source repo: EzraStone/llm-forecast-calibration.
Data mirrored from GitHub commit 0f12f71a2c2ec8c54cafeb4231fecb87e705e660.
All eight JSONL files match the source data byte for byte. The source repository
remains canonical for analysis code… See the full description on the dataset page: https://huggingface.co/datasets/ezra77/llm-forecast-calibration.retail-forecast-optimize-benchmark
Retail Forecast-Optimize Benchmark
Benchmark dataset for predict-then-optimize retail inventory decisions.
Contents
File pattern
Description
{scenario}_s{seed}.json
Full policy comparison per scenario instance
samples/sample_*.json
Retail scenario definitions
eval_results.json
Aggregated leaderboard metrics
chronos-2-live-forecasts/
Live Chronos-2 outputs from HF Jobs (when available)
Scenarios
Baseline Operations
Promotion… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/retail-forecast-optimize-benchmark.beyond-stuxnet-ii-minnesota-forecast-audit
Beyond Stuxnet II: Minnesota Water Systems Forecast Audit
This repository contains a defensive, evidence-grounded white paper auditing an earlier structural cyber-physical threat model against later reported Minnesota water-system incidents.
The paper does not claim that the earlier work predicted Iran, Minnesota, water utilities, a particular intrusion path, or every reported event detail. Its narrower conclusion is that the earlier public artifact committed to a mechanism… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/beyond-stuxnet-ii-minnesota-forecast-audit.Electricity_Load_Forecasting_using_LLMSbtcusdt_feb_2026_candles_news_forecast
BTCUSDT Multi-Timeframe GARCH Volatility Dataset
Symbol: BTCUSDTPeriod: February 2026 (backtest simulation)Records: ~40 320 (1-minute resolution)Source tool: backtest-kit + garch
Dataset Description
Each row is a 1-minute snapshot of GARCH-predicted volatility (sigma) for BTCUSDT across 8 timeframes, computed during a backtest run. The reliable flag indicates whether the model had enough historical candles to produce a statistically stable estimate.
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/btcusdt_feb_2026_candles_news_forecast.financial-forecast
Financial Forecast Dataset
Vietnamese finance QA dataset focused on company metrics and comparisons.
Columns
id: Unique identifier
question: Formatted question
question_raw: Raw question text
answer_type: Type of answer (e.g., entity)
answer: The numerical answer
time_asked: Date the question was asked
sources: Data sources (e.g., company_quarter)
data_source: Identifier for the data source pipeline (e.g., finance-sql)
reward_model: Reward model config containing:
style:… See the full description on the dataset page: https://huggingface.co/datasets/hung20gg/financial-forecast.Utili-micro-task-forecastingFOReCAst
FOReCAst Dataset
Note: This dataset is temporarily hosted under a separate anonymous account to comply with the double-blind review policy of an ongoing submission. The dataset content is identical to the version hosted under the main project account. Once the review process concludes, the repositories will be consolidated.
FOReCAst (Future Outcome Reasoning and Confidence Assessment) is a benchmark dataset for evaluating language models on reasoning about uncertain future events… See the full description on the dataset page: https://huggingface.co/datasets/DoubleBlindAccount/FOReCAst.Forecastingv1V1 of a custom event forecasting dataset.
Data is not very well filtered or well formatted.
PoliNews-Forecasttraining-forecast-analysisbitcoin-forecaster
