datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
corral-traces
Corral – Evaluation Traces
Full evaluation traces across Corral environments, models, agents, and task granularities
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full evaluation traces collected across all 8 Corral environments.
Each configuration (config) corresponds to a unique combination of model, environment, scope (difficulty… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-traces.corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.corral-oss-trace-logprobs
Corral – OSS-120B Trace Logprobs
Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments.
Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.corral-intervention-traces
Corral – Intervention Traces
Run traces for the intervention ablation study across all evaluated models and Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full message-history traces from the intervention ablation study, covering all evaluated models across all 8 Corral environments.
The intervention ablation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-traces.corral-QAs-topic_reports
Corral – QA Topic Reports
Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.corral-reasoning-annotations
Corral – Reasoning Annotations
LLM epistemic annotations over Corral traces where the annotator judged that the agents do not reason scientifically
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains annotated evaluation traces with LLM-generated epistemic annotations across the Corral benchmark.
The dataset is exposed as three model-specific… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-reasoning-annotations.corral-intervention-reports
Corral – Intervention Ablation Reports
Intervention ablation reports for multiple LLM agents across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the final evaluation reports of the intervention ablation study conducted across multiple LLM agents and all 8 Corral environments.
The intervention ablation examines… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports.corral-oss-intervention-resistorcorral-result-v0
