CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clembench-playpen /lm-pragmatics Citation @misc{hu2023finegrainedcomparisonpragmaticlanguage, title={A fine-grained comparison of pragmatic language understanding in humans and language models}, author={Jennifer Hu and Sammy Floyd and Olessia Jouravlev and Evelina Fedorenko and Edward Gibson}, year={2023}, eprint={2212.06801}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2212.06801},} tabularquestion-answering1K<n<10K0 likes261 downloads2y agoHugging Face02clembench-playpen /working-memory0 likes175 downloads2y agoHugging Face03clembench-playpen /natural-plan-calendartext1K<n<10K1 likes79 downloads2y agoHugging Face04clembench-playpen /natural-plan-meetingtext1K<n<10K1 likes63 downloads2y agoHugging Face05clembench-playpen /glue_diagnostics Citation @inproceedings{wang2019glue, title={{GLUE}: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding}, author={Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R.}, note={In the Proceedings of ICLR.}, year={2019}} textquestion-answeringn<1K0 likes57 downloads2y agoHugging Face06clembench-playpen /natural-plan-triptext1K<n<10K0 likes49 downloads2y agoHugging Face07pm-25 /clembench-rlvr-dataset Clembench RLVR Dataset – Full (Wins + Losses) combines SFT-Final (https://huggingface.co/datasets/clembench-playpen/SFT-Final-Dataset) and DPO_dialogue (https://huggingface.co/datasets/clembench-playpen/DPO_dialogue). Info: "id": concat of "game"+"episode" "query": "string in llama-format" "reward": rewards given (1=="chosen" from dpo-dialogue and "success" from sft-final, 0== "rejected" from dpo-dialogue) "origin": marks origin of sample "player": kept for dpo-dialogue samples… See the full description on the dataset page: https://huggingface.co/datasets/pm-25/clembench-rlvr-dataset.text10K<n<100K0 likes23 downloads1y agoHugging Face08clembench-playpen /DPO_dialogue Dataset Details This is the training dataset for the DPO Dialogue strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback. This preference dataset has been obtained from these games' instances using this script, with --preference_depth dialogue. Dataset Description Preference Dataset where chosen vs rejected continuations are the successful vs unsuccessful interactions stored. Language(s) (NLP): English Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_dialogue.text10K<n<100K0 likes21 downloads1y agoHugging Face09clembench-playpen /DPO_1neg_Aborted_FINALtext10K<n<100K0 likes18 downloads2y agoHugging Face10clembench-playpen /DPO_1neg_Aborted_old_LAtext1K<n<10K0 likes17 downloads2y agoHugging Face11clembench-playpen /DPO_1neg_Aborted_best_models_old_LAtext1K<n<10K0 likes17 downloads2y agoHugging Face12LuckyLukke /clembench_REFUEL_0.7tabular1K<n<10K0 likes16 downloads1y agoHugging Face13LuckyLukke /clembench_REFUEL_SFT_3tabular1K<n<10K0 likes14 downloads2y agoHugging Face14LuckyLukke /clembench_REFUEL_SFT_1_iter2_tabular1K<n<10K0 likes14 downloads2y agoHugging Face15clembench-playpen /DPO_turntabular10K<n<100K0 likes14 downloads1y agoHugging Face16clembench-playpen /lmentryLMentry is a benchmark for measuring language model performance on tasks that are trivial to humans. LMentry consists of 25 tasks which humans are generally expected to perform perfectly, e.g. writing a sentence containing a specific word, identifying which words in a list belong to a specific category, choosing which of two words is longer, or identifying which of two words rhymes with a third word.question-answering0 likes13 downloads2y agoHugging Face17clembench-playpen /KTO_Aborted_FINALtabular100K<n<1M0 likes13 downloads2y agoHugging Face18clembench-playpen /DPO_turn_bug Dataset Details This is the training dataset for the DPO Turn strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback. These preference dataset have been obtained from these games' instances using this script, with --preference_depth turn. Given the huge number of chosen vs rejected pairs in the first turn of the conversation, we limit the numbers of chosen and rejected pairs for the first turn to 10k samples (--first_turn_limit True).… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_turn_bug.tabulartoken-classification10K<n<100K0 likes13 downloads1y agoHugging Face19LuckyLukke /clembench_REFUEL_SFT_1tabular1K<n<10K0 likes12 downloads2y agoHugging Face20clembench-playpen /KTO_Aborted_best_models_Ftabular1K<n<10K0 likes12 downloads2y agoHugging Face21clembench-playpen /DPO_2neg_Aborted_FINALtext10K<n<100K0 likes11 downloads2y agoHugging Face22clembench-playpen /warm-up_synthetic-datatabular10K<n<100K0 likes11 downloads1y agoHugging Face23clembench-playpen /DPO_turn_allneg_old_and_newtabular100K<n<1M0 likes11 downloads1y agoHugging Face24clembench-playpen /KTO_Aborted_same_family_model_FINALtabular10K<n<100K0 likes10 downloads2y agoHugging Face25clembench-playpen /DPO_1neg_Aborted_same_family_model_old_LAtext1K<n<10K0 likes9 downloads2y agoHugging Face26clembench-playpen /SFT-Final-Datasettabular1K<n<10K0 likes9 downloads1y agoHugging Face27clembench-playpen /KTO_Aborted_Ftabular10K<n<100K0 likes8 downloads2y agoHugging Face28clembench-playpen /DPO_2neg_Aborted_same_family_model_old_LAtext1K<n<10K0 likes8 downloads2y agoHugging Face29clembench-playpen /DPO_allneg_Aborted_old_and_new_exptabular100K<n<1M0 likes8 downloads1y agoHugging Face30clembench-playpen /DNLI Adam Ek, Bill Noble, Stergios Chatzikyriakidis, Robin Cooper, Simon Dobnik, Eleni Gregoromichelaki, Christine Howes, Staffan Larsson, Vladislav Maraev, Gregory Mills, and Gijs Wijnholds. 2024. I hea- umm think that’s what they say: A Dataset of Inferences from Natural Language Dialogues. In Proceedings of the 28th Workshop on the Semantics and Pragmatics of Dialogue. 0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.