CoolFace
Datasetpublic

jablonkagroup/corral-oss-trace-logprobs

Corral โ€“ OSS-120B Trace Logprobs Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments ๐Ÿ“‹ Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments. Each configuration (config) of this datasetโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes534downloads
Dataset Card

Corral โ€“ OSS-120B Trace Logprobs

<div align="center">

[image]

![Website](https://lamalab-org.github.io/corral/) ![Docs](https://lamalab-org.github.io/corral/docs/) ![GitHub](https://github.com/lamalab-org/corral) ![License: MIT](https://opensource.org/licenses/MIT) ![Paper](https://arxiv.org/abs/2604.18805) ![Dataset](https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs)

Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments

</div>


๐Ÿ“‹ Dataset Summary

This dataset is part of the Corral collection accompanying the paper *AI scientists produce results without reasoning scientifically*. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments.

Each configuration (config) of this dataset corresponds to a unique combination of environment, scope (difficulty level), and granularity (tasks or subtasks). For example, a config such as afm_level_1_subtasks contains the logprob records for the AFM Experiment Execution environment at scope level 1, broken down at the subtask granularity. The full set of configs spans the Cartesian product of the 8 environments, their respective scope levels, and the tasks/subtasks split.

This resource is designed for process-level analysis, interpretability research, and auditing of scientific-agent reasoning โ€” not for general-purpose model pre-training.

๐ŸŽฏ Supported Uses

  • โ€”๐Ÿ” Auditing token-level confidence and uncertainty in scientific-agent completions
  • โ€”๐Ÿ“Š Studying the relationship between model certainty and task success
  • โ€”๐Ÿ” Reproducing and extending the log-probability analyses reported in the paper
  • โ€”๐Ÿ“ Meta-evaluation and calibration studies of frontier LLMs on scientific tasks

๐Ÿงช About Corral

*Corral* is a framework for the science of agents and agents for science. It provides a microservice architecture that decouples agents from environments via a clientโ€“server design (REST API), ensuring flexibility, reproducibility, and robust isolation.

  • โ€”๐ŸŒ Environments define the task space, available tools, and observable feedback โ€” from chemistry labs to HPC clusters.
  • โ€”๐Ÿค– Agents are modular LLM-based entities supporting scaffolds such as ReAct, ToolCalling, LLMPlanner, and Reflection.
  • โ€”๐Ÿ“ Tasks define problems to solve, complete with scoring functions. Tasks can be chained into TaskGroups for complex multi-stage challenges.

Corral currently ships 8 environments, 97 tools, 115 tasks, and 786 subtasks spanning chemistry, physics, and materials science.

๐ŸŒ Environments

EnvironmentDescription๐Ÿ”ง Tools๐Ÿ“ Tasks/scope๐Ÿ”ญ Scopesโฑ๏ธ Avg. trace length
๐Ÿงซ Inorganic Qualitative AnalysisIdentify unknown cations in solution through systematic wet-lab procedures (reagent addition, flame tests, pH measurement, centrifugation, etc.). Observations are computed from thermodynamic data. Three scopes progressively increase the number of candidate ions.1410339.4
โšก Circuit InferenceRecover the topology and component values of a hidden resistor network from pairwise resistance measurements. Tools provide series/parallel calculations, delta-wye transforms, and circuit validation.96115.0
๐Ÿ”ญ Spectroscopic Structure ElucidationDetermine the molecular structure of an unknown compound by requesting and interpreting spectroscopic data (MS, NMR, HSQC, IR) alongside reference databases for chemical shifts and isotope distributions.1620215.1
๐Ÿงฌ Retrosynthetic PlanningDesign multi-step synthetic routes to target molecules under cost, step-count, and commercial-availability constraints, using a template catalogue and functional-group detection tools.158325.5
๐Ÿค– ML-based Property PredictionAssemble a complete ML pipeline to predict formation energies of material polymorphs using data from the Materials Project, covering feature engineering, XGBoost training, and cross-validation.143116.6
๐Ÿ”ฌ AFM Experiment ExecutionAnalyze and interpret atomic force microscopy data for nanoscale surface characterization, including topographical and mechanical property measurements.61426.3
โš›๏ธ Molecular SimulationDesign and execute molecular dynamics simulations with LAMMPS to predict materials properties, covering the full workflow from crystal structure retrieval to force-field queries and log analysis.82โ€“3230.4
๐Ÿ—๏ธ Adsorption Surface ConstructionBuild adsorbateโ€“slab configurations from bulk crystal structures for heterogeneous catalysis studies, integrating Materials Project retrieval, slab generation, and adsorption-site enumeration.153119.6

๐Ÿ—‚๏ธ Dataset Structure

Configs

Each config name encodes {environment}_{scope}_{granularity}, where:

  • โ€”environment is a short identifier for one of the 8 Corral environments (e.g., afm, circuit_inference, spectroscopic, retrosynthesis, ml_property, molecular_simulation, adsorption).
  • โ€”scope is the difficulty level (e.g., level_1, level_2, level_3).
  • โ€”granularity is either tasks or subtasks.

Data Splits

All configs expose a single train split.

Data Instances

Each row corresponds to one token-level log-probability record from a GPT-Oss-120B completion produced during an agent evaluation run.


๐Ÿ—๏ธ Dataset Creation

Curation Rationale

This dataset was created as part of Corral to enable process-level analysis of LLM-based scientific agents, specifically to study how token-level confidence relates to scientific reasoning quality and task outcomes.

Source Data

Records are derived from agent evaluation runs on Corral benchmark tasks, capturing the log-probabilities returned by GPT-Oss-120B for each generated token across all 8 environments and their scope levels.


๐Ÿ”— Relation to Other Corral Artifacts

This dataset is one component of the broader Corral release and is best interpreted together with the matching task definitions, execution traces, reports, aggregate results, and reasoning annotations available in the *Corral* collection.


๐Ÿ“„ Citation

bibtex
@article{rรญos-garcรญa2026ai,
  title   = {AI scientists produce results without reasoning scientifically},
  author  = {Martiรฑo Rรญos-Garcรญa and Nawaf Alampara and Chandan Gupta and Indrajeet Mandal and Sajid Mannan and Ali Asghar Aghajani and N. M. Anoop Krishnan and Kevin Maik Jablonka},
  year    = {2026},
  journal = {arXiv preprint arXiv: 2604.18805}
}

๐Ÿ“œ License

This dataset is released under the MIT License.

Changelog

2026-04-22

  • โ€”Initial release of the dataset card.