CoolFace
14 results

corral

jablonkagroup /corral-traces Corral – Evaluation Traces Full evaluation traces across Corral environments, models, agents, and task granularities 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full evaluation traces collected across all 8 Corral environments. Each configuration (config) corresponds to a unique combination of model, environment, scope (difficulty… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-traces.tabulartext-generation10K<n<100K0 likes1.3k downloads3mo agoHugging Facejablonkagroup /corral-environment-tasks Corral – Environment Tasks Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark. The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.texttext-generationn<1K0 likes1.1k downloads3mo agoHugging Facejablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes540 downloads3mo agoHugging Facejablonkagroup /corral-oss-trace-logprobs Corral – OSS-120B Trace Logprobs Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments. Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.tabulartext-generation100K<n<1M0 likes534 downloads3mo agoHugging Facejablonkagroup /corral_lfm_binomial_results Corral – LFM Binomial IRT Results Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the fitted parameters of a binomial Item Response Theory (IRT) model estimated from agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results.text-generation1K<n<10K0 likes349 downloads5mo agoHugging Facejablonkagroup /corral-QAs-reports Corral – QA Reports Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.texttext-generation1K<n<10K0 likes213 downloads3mo agoHugging Face