datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edgar-forecast-benchmark
EDGAR-Forecast Benchmark
EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings.
Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs.
Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.inductive-forecasting-data
Inductive Forecasting Study — Anonymous Data Release
This repository is the anonymous data companion to a paper studying behavioral
signatures of inductive reasoning in language-model forecasts. It packages the
frozen inputs, model responses, row-level scores, and aggregate result artifacts
used by the paper's four main experiments, together with synthetic appendix
transfer studies.
The release is organized as Hugging Face dataset configurations so each study can
be loaded… See the full description on the dataset page: https://huggingface.co/datasets/od2961/inductive-forecasting-data.forecastgen-artifacts
Forecast-Generalization: raw evaluation outputs across 38 reasoning models
Complete generation-level outputs, per-seed scores and analysis artifacts from a
study of how well benchmark performance forecasts generalization to held-out
reasoning tasks.
Most released evaluations report only aggregate accuracy. This release keeps the
raw per-problem, per-seed generations, so item-level analyses can be redone
without re-running any inference.
What is here
38 models… See the full description on the dataset page: https://huggingface.co/datasets/dvader13/forecastgen-artifacts.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.gdelt-forecast-freeform
GDELT-Forecast Free-form
924 free-form forecasting questions (named entities, numbers, dates, short narrative answers) generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models on non-binary… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-freeform.Geopol-Forecaster-Predictions
Geopol Forecaster Predictions
Structured predictions extracted from multi-agent geopolitical wargaming simulations run by Geopol Forecaster.
Dataset Structure
File
Description
runs.csv
Simulation run metadata — scenario, model pool, runtime, timestamps
predictions.csv
Discrete testable predictions with probability estimates, time horizons, actor attribution, and denormalized run metadata
assessments.csv
Post-prediction accuracy grades for predictions whose… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Geopol-Forecaster-Predictions.gdelt-forecast-binary
GDELT-Forecast Binary
1,215 yes/no forecasting questions generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models in a strict forecasting posture — the model sees only news that was publicly… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-binary.pg320-forecast-eval
pg-320 — Politics / Geopolitics Forecasting Eval
A 320-question yes/no forecasting benchmark used to evaluate the Anthral Research GRPO LoRA adapters against frontier closed-source baselines.
Each question is bundled with the exact retrieval context that was used during the published evaluation runs — top-N news chunks, all dated strictly before the question creation date — so anyone can reproduce the headline numbers locally without re-running retrieval.
Headline result… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/pg320-forecast-eval.fingpt-forecaster-dow30-sequential_2
FinGPT Forecaster DOW30 — Chain-of-Thought Dataset
Weekly stock price movement predictions for DOW-30 constituents, augmented with
Chain-of-Thought (CoT) reasoning generated via GPT-4.
Built on top of the base dataset:
FinGPT/fingpt-forecaster-dow30-202305-202405
Schema
Field
Description
symbol
DOW-30 ticker (e.g. AXP, MSFT)
period
Forecast week, e.g. 2023-05-14 to 2023-05-21
prompt
Full instruction prompt fed to the model
answer
CoT reasoning +… See the full description on the dataset page: https://huggingface.co/datasets/korra141/fingpt-forecaster-dow30-sequential_2.
