datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mawps-asdiv-a_svamp
Dataset Card for "mawps-asdiv-a_svamp"
More Information needed
cv_svamp_augmented_fold0
Dataset Card for "cv_svamp_augmented_fold0"
More Information needed
X-SVAMP_en_zh_ko_it_es
X-SVAMP
🤗 Paper | 📖 arXiv
Dataset Description
X-SVAMP is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the math reasoning abilities of LLMs. The dataset is translated by GPT-4-turbo from the original English-version SVAMP.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-SVAMP_en_zh_ko_it_es.cv_svamp_augmented_fold1
Dataset Card for "cv_svamp_augmented_fold1"
More Information needed
bad_data_gsm8k_svamp.csvSome bad data discovered in the popular GSM8K and SVAMP LLM benchmarking datasets.
These examples have incorrect answers in the corresponding math problem benchmark dataset, and should not be used to evaluate AI models.
We detected this bad data automatically using Cleanlab's Trustworthy Language Model. TLM's estimated trustworthiness score for each example is also provided.
Example error found in the GSM8K dataset:
Question: After scoring 14 points, Erin now has three times… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/bad_data_gsm8k_svamp.csv.svampcv_svamp_augmented_fold2
Dataset Card for "cv_svamp_augmented_fold2"
More Information needed
cv_svamp_augmented_fold3
Dataset Card for "cv_svamp_augmented_fold3"
More Information needed
cv_svamp_augmented_fold4
Dataset Card for "cv_svamp_augmented_fold4"
More Information needed
Calc-svamp-Taggedcodi-gpt2-svamp-activations
CODI-GPT-2 activations at the : token (SVAMP + counterfactuals)
This dataset contains the residual stream of CODI-GPT-2 captured at
the : token position during answer emission, for SVAMP and a family of
counterfactual SVAMP variants.
What is the : position?
CODI-GPT-2 emits answers in the template The answer is: <number>. After the
latent reasoning loop (6 latent steps) ends with the EOT marker, the model
autoregressively decodes:
decode step
input (= previously… See the full description on the dataset page: https://huggingface.co/datasets/sandrajyluo/codi-gpt2-svamp-activations.
