datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edgar-forecast-benchmark
EDGAR-Forecast Benchmark
EDGAR-Forecast is a closed-sandbox benchmark for filing-grounded numerical forecasting from historical SEC filings in EDGAR. The benchmark contains 50 company-level instances and 250 numeric forecast targets from hidden 2026 10-Q filings.
Questions that mention 2025 refer to values disclosed in 2026 Q1 filings; those filings were filed in Q1 2026, so they remain outside the evaluated models' knowledge cutoffs.
Each benchmark directory includes the question… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/edgar-forecast-benchmark.forecastgen-artifacts
Forecast-Generalization: raw evaluation outputs across 38 reasoning models
Complete generation-level outputs, per-seed scores and analysis artifacts from a
study of how well benchmark performance forecasts generalization to held-out
reasoning tasks.
Most released evaluations report only aggregate accuracy. This release keeps the
raw per-problem, per-seed generations, so item-level analyses can be redone
without re-running any inference.
What is here
38 models… See the full description on the dataset page: https://huggingface.co/datasets/dvader13/forecastgen-artifacts.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.llm-forecast-calibration
LLM Forecast Calibration Study — GLM-5.3 on resolved Manifold Markets questions
Raw generation data for the study "Does sampling K times beat thinking harder?
A controlled study of LLM forecast calibration on resolved binary questions."
Source repo: EzraStone/llm-forecast-calibration.
Data mirrored from GitHub commit 0f12f71a2c2ec8c54cafeb4231fecb87e705e660.
All eight JSONL files match the source data byte for byte. The source repository
remains canonical for analysis code… See the full description on the dataset page: https://huggingface.co/datasets/ezra77/llm-forecast-calibration.biomedical-forecasting-lightningrod
Biomedical Forecasting Dataset
A dataset of 1444 binary forecasting questions about biomedical and public health outcomes. Each question is a forward-looking prediction (Yes/No) about a real event, grounded in news and labeled with the actual outcome.
What is in this dataset?
Questions: FDA drug approvals, clinical trial results (Phase 2/3), WHO and CDC declarations, vaccine development, disease outbreaks, gene therapy, and public health policy.
Grounded in real news:… See the full description on the dataset page: https://huggingface.co/datasets/Ainoafv/biomedical-forecasting-lightningrod.gdelt-forecast-freeform
GDELT-Forecast Free-form
924 free-form forecasting questions (named entities, numbers, dates, short narrative answers) generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models on non-binary… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-freeform.financial-forecast
Financial Forecast Dataset
Vietnamese finance QA dataset focused on company metrics and comparisons.
Columns
id: Unique identifier
question: Formatted question
question_raw: Raw question text
answer_type: Type of answer (e.g., entity)
answer: The numerical answer
time_asked: Date the question was asked
sources: Data sources (e.g., company_quarter)
data_source: Identifier for the data source pipeline (e.g., finance-sql)
reward_model: Reward model config containing:
style:… See the full description on the dataset page: https://huggingface.co/datasets/hung20gg/financial-forecast.gdelt-forecast-binary
GDELT-Forecast Binary
1,215 yes/no forecasting questions generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models in a strict forecasting posture — the model sees only news that was publicly… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-binary.pg320-forecast-eval
pg-320 — Politics / Geopolitics Forecasting Eval
A 320-question yes/no forecasting benchmark used to evaluate the Anthral Research GRPO LoRA adapters against frontier closed-source baselines.
Each question is bundled with the exact retrieval context that was used during the published evaluation runs — top-N news chunks, all dated strictly before the question creation date — so anyone can reproduce the headline numbers locally without re-running retrieval.
Headline result… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/pg320-forecast-eval.
