datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
corral-traces
Corral – Evaluation Traces
Full evaluation traces across Corral environments, models, agents, and task granularities
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full evaluation traces collected across all 8 Corral environments.
Each configuration (config) corresponds to a unique combination of model, environment, scope (difficulty… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-traces.corral-environment-tasks
Corral – Environment Tasks
Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark.
The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.corral-oss-trace-logprobs
Corral – OSS-120B Trace Logprobs
Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments.
Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.corral_lfm_binomial_results
Corral – LFM Binomial IRT Results
Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the fitted parameters of a binomial Item Response Theory (IRT) model estimated from agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results.corral-QAs-reports
Corral – QA Reports
Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.corral-intervention-traces
Corral – Intervention Traces
Run traces for the intervention ablation study across all evaluated models and Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full message-history traces from the intervention ablation study, covering all evaluated models across all 8 Corral environments.
The intervention ablation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-traces.corral-intervention-reports
Corral – Intervention Ablation Reports
Intervention ablation reports for multiple LLM agents across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the final evaluation reports of the intervention ablation study conducted across multiple LLM agents and all 8 Corral environments.
The intervention ablation examines… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports.awesome-prompt-patterns
💬 Awesome Prompt Patterns
Prompt patterns are instructions guiding AI responses for specific tasks and are defined by core contextual statements that enhance the precision and relevancy of an output from an LLM.
View more prompt patterns and techniques on GitHub.
license: cc
