CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cq01 /mawps-asdiv-a_svamp Dataset Card for "mawps-asdiv-a_svamp" More Information needed tabular1K<n<10K0 likes4k downloads3y agoHugging Face02sethapun /cv_svamp_augmented_fold0 Dataset Card for "cv_svamp_augmented_fold0" More Information needed tabular1K<n<10K0 likes126 downloads4y agoHugging Face03zhihz0535 /X-SVAMP_en_zh_ko_it_es X-SVAMP 🤗 Paper | 📖 arXiv Dataset Description X-SVAMP is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish). It is intended to evaluate the math reasoning abilities of LLMs. The dataset is translated by GPT-4-turbo from the original English-version SVAMP. In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-SVAMP_en_zh_ko_it_es.tabularquestion-answering1K<n<10K2 likes71 downloads3y agoHugging Face04sethapun /cv_svamp_augmented_fold1 Dataset Card for "cv_svamp_augmented_fold1" More Information needed tabular1K<n<10K0 likes45 downloads4y agoHugging Face05Cleanlab /bad_data_gsm8k_svamp.csvSome bad data discovered in the popular GSM8K and SVAMP LLM benchmarking datasets. These examples have incorrect answers in the corresponding math problem benchmark dataset, and should not be used to evaluate AI models. We detected this bad data automatically using Cleanlab's Trustworthy Language Model. TLM's estimated trustworthiness score for each example is also provided. Example error found in the GSM8K dataset: Question: After scoring 14 points, Erin now has three times… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/bad_data_gsm8k_svamp.csv.tabularn<1K3 likes40 downloads2y agoHugging Face06abhaygupta1266 /svamptabular1K<n<10K0 likes40 downloads1y agoHugging Face07sethapun /cv_svamp_augmented_fold2 Dataset Card for "cv_svamp_augmented_fold2" More Information needed tabular1K<n<10K0 likes37 downloads4y agoHugging Face08sethapun /cv_svamp_augmented_fold3 Dataset Card for "cv_svamp_augmented_fold3" More Information needed tabular1K<n<10K0 likes27 downloads4y agoHugging Face09sethapun /cv_svamp_augmented_fold4 Dataset Card for "cv_svamp_augmented_fold4" More Information needed tabular1K<n<10K0 likes25 downloads4y agoHugging Face10vishruthnath /Calc-svamp-Taggedtabularn<1K0 likes24 downloads3y agoHugging Face11sandrajyluo /codi-gpt2-svamp-activations CODI-GPT-2 activations at the : token (SVAMP + counterfactuals) This dataset contains the residual stream of CODI-GPT-2 captured at the : token position during answer emission, for SVAMP and a family of counterfactual SVAMP variants. What is the : position? CODI-GPT-2 emits answers in the template The answer is: <number>. After the latent reasoning loop (6 latent steps) ends with the EOT marker, the model autoregressively decodes: decode step input (= previously… See the full description on the dataset page: https://huggingface.co/datasets/sandrajyluo/codi-gpt2-svamp-activations.tabularn<1K0 likes18 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.