CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jablonkagroup /corral-traces Corral – Evaluation Traces Full evaluation traces across Corral environments, models, agents, and task granularities 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full evaluation traces collected across all 8 Corral environments. Each configuration (config) corresponds to a unique combination of model, environment, scope (difficulty… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-traces.tabulartext-generation10K<n<100K0 likes1.3k downloads3mo agoHugging Face02jablonkagroup /corral-environment-tasks Corral – Environment Tasks Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark. The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.texttext-generationn<1K0 likes1.1k downloads3mo agoHugging Face03jablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes540 downloads3mo agoHugging Face04jablonkagroup /corral-oss-trace-logprobs Corral – OSS-120B Trace Logprobs Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments. Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.tabulartext-generation100K<n<1M0 likes534 downloads3mo agoHugging Face05jablonkagroup /corral_lfm_binomial_results Corral – LFM Binomial IRT Results Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the fitted parameters of a binomial Item Response Theory (IRT) model estimated from agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results.text-generation1K<n<10K0 likes349 downloads5mo agoHugging Face06jablonkagroup /corral-QAs-reports Corral – QA Reports Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.texttext-generation1K<n<10K0 likes213 downloads3mo agoHugging Face07jablonkagroup /corral-intervention-traces Corral – Intervention Traces Run traces for the intervention ablation study across all evaluated models and Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the full message-history traces from the intervention ablation study, covering all evaluated models across all 8 Corral environments. The intervention ablation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-traces.tabulartext-generation1K<n<10K0 likes103 downloads3mo agoHugging Face08jablonkagroup /corral-intervention-reports Corral – Intervention Ablation Reports Intervention ablation reports for multiple LLM agents across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the final evaluation reports of the intervention ablation study conducted across multiple LLM agents and all 8 Corral environments. The intervention ablation examines… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports.tabulartext-generationn<1K0 likes34 downloads3mo agoHugging Face09corralm /awesome-prompt-patterns 💬 Awesome Prompt Patterns Prompt patterns are instructions guiding AI responses for specific tasks and are defined by core contextual statements that enhance the precision and relevancy of an output from an LLM. View more prompt patterns and techniques on GitHub. license: cc texttext-generationn<1K1 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.