datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CaT-Bench
Dataset Card for CaT-Bench
CaT-Bench is a benchmark dataset designed to evaluate large language models' (LLMs) understanding of causal and temporal dependencies in natural language plans, specifically in cooking recipes based on the English Recipe Flow Graph Corpus by Yamakata et al. (2020). It consists of questions that test whether one step must necessarily occur before or after another, requiring reasoning about preconditions, effects, and the overall structure of the plan.… See the full description on the dataset page: https://huggingface.co/datasets/vanyacohen/CaT-Bench.catbench-results
CatBench results
Model outputs for CatBench, a small
benchmark that asks a model to draw a cute kitten two ways and looks at what comes
back. Produced by the /catbench command in
blockquant.
Upstream publishes its own results at
Katehuuh.github.io/demos/CatBench/assets.
This dataset holds runs for models that are not in that set. /catbench checks both
and only rents a pod when neither has the model, so the two do not duplicate
each other.
The prompts
Verbatim… See the full description on the dataset page: https://huggingface.co/datasets/Honkware/catbench-results.
