datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.frames-benchmark-predictions
BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation
Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions)
Executive Summary
This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark.
Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.table-sft-eval-predictions
💾 Raw Predictions for "What Really Matters for Table LLMs?"
This dataset contains the raw model outputs from the experiments in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.Geopol-Forecaster-Predictions
Geopol Forecaster Predictions
Structured predictions extracted from multi-agent geopolitical wargaming simulations run by Geopol Forecaster.
Dataset Structure
File
Description
runs.csv
Simulation run metadata — scenario, model pool, runtime, timestamps
predictions.csv
Discrete testable predictions with probability estimates, time horizons, actor attribution, and denormalized run metadata
assessments.csv
Post-prediction accuracy grades for predictions whose… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Geopol-Forecaster-Predictions.STEM_train_edit_predictions_test
STEM Image Edit Predictions Dataset
This dataset contains AI-generated edit predictions for STEM images based on captions and edit commands.
Dataset Structure
The dataset is organized in batches:
Total batches: 1
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
layout_summary: Concise description of source image layout and key elements
edit_analysis: Specific elements to be edited and changes to be made… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_edit_predictions_test.
