datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.OracleProto
OracleProto: Forecasting Evaluation Set
Chinese doc: [中文文档]
GitHub repo: [MaYiding/OracleProto]
Visit Our Leaderboards: [Website]
View Our Paper: [arXiv]
A SQLite-packaged evaluation set of 80 hand-curated forecasting questions on real-world events, with resolution dates between 2026-03-12 and 2026-04-14, released alongside the GitHub Repo. Both the rows and the byte-stable prompt-reconstruction recipe are packaged in a single file, forecast_eval_set_example.db, which exposes two… See the full description on the dataset page: https://huggingface.co/datasets/MaYiding/OracleProto.quant_eval_golden_oracle_fixtures
quant_eval — Golden oracle fixtures
The locked evaluation fixture set — 1,600 cases across eight agent task families with their deterministic ground truth — plus the crosswalk mapping every run's recorded fixture hash and version label to the published file.
Part of the quant_eval public corpus: a per-case behavioral evaluation of full-weight and quantized large language models across eight agent-relevant task families, with paired statistical testing.
Cite this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/quant_eval_golden_oracle_fixtures.synth-stress-oracle-agent-priv-esc-v0.1What this repo does
This repository implements a synthetic pre-deployment stress oracle for AI agent architectures.
You submit a proposed configuration.
The oracle returns:
cascade probability
risk band
predicted label
sensitivity to small drift
top risk drivers
redesign moves
The goal is to test system stability before deployment.
Core scenario
Example:
You plan to deploy 50 AI agents with shared credentials.
You test the configuration before rollout.
Inputs might look like:
auth_pressure =… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/synth-stress-oracle-agent-priv-esc-v0.1.sqli-err-oracle-falsesqli-oracle-falsecombined_mmlu_llama8b_oracle_difficultycombined_math500_llama70b_oracle_difficultycombined_gpqa_llama8b_oracle_difficultycombined_mmlu_llama70b_oracle_difficultycombined_math500_llama8b_oracle_difficultycombined_gpqa_llama70b_oracle_difficultysqli-oracle-truesqli-err-oracle-trueORACLE-V4
