datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-pragmatics
Citation
@misc{hu2023finegrainedcomparisonpragmaticlanguage, title={A fine-grained comparison of pragmatic language understanding in humans and language models}, author={Jennifer Hu and Sammy Floyd and Olessia Jouravlev and Evelina Fedorenko and Edward Gibson}, year={2023}, eprint={2212.06801}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2212.06801},}
working-memorynatural-plan-calendarnatural-plan-meetingglue_diagnostics
Citation
@inproceedings{wang2019glue, title={{GLUE}: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding}, author={Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R.}, note={In the Proceedings of ICLR.}, year={2019}}
natural-plan-tripclembench-rlvr-dataset
Clembench RLVR Dataset – Full (Wins + Losses)
combines SFT-Final (https://huggingface.co/datasets/clembench-playpen/SFT-Final-Dataset) and DPO_dialogue (https://huggingface.co/datasets/clembench-playpen/DPO_dialogue).
Info:
"id": concat of "game"+"episode"
"query": "string in llama-format"
"reward": rewards given (1=="chosen" from dpo-dialogue and "success" from sft-final, 0== "rejected" from dpo-dialogue)
"origin": marks origin of sample
"player": kept for dpo-dialogue samples… See the full description on the dataset page: https://huggingface.co/datasets/pm-25/clembench-rlvr-dataset.DPO_dialogue
Dataset Details
This is the training dataset for the DPO Dialogue strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback.
This preference dataset has been obtained from these games' instances using this script, with --preference_depth dialogue.
Dataset Description
Preference Dataset where chosen vs rejected continuations are the successful vs unsuccessful interactions stored.
Language(s) (NLP): English
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_dialogue.DPO_1neg_Aborted_FINALDPO_1neg_Aborted_old_LADPO_1neg_Aborted_best_models_old_LAclembench_REFUEL_0.7clembench_REFUEL_SFT_3clembench_REFUEL_SFT_1_iter2_DPO_turnlmentryLMentry is a benchmark for measuring language model performance on tasks that are trivial to humans. LMentry consists of 25 tasks which humans are generally expected to perform perfectly, e.g. writing a sentence containing a specific word, identifying which words in a list belong to a specific category, choosing which of two words is longer, or identifying which of two words rhymes with a third word.KTO_Aborted_FINALDPO_turn_bug
Dataset Details
This is the training dataset for the DPO Turn strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback.
These preference dataset have been obtained from these games' instances using this script, with --preference_depth turn.
Given the huge number of chosen vs rejected pairs in the first turn of the conversation, we limit the numbers of chosen and rejected pairs for the first turn to 10k samples (--first_turn_limit True).… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_turn_bug.clembench_REFUEL_SFT_1KTO_Aborted_best_models_FDPO_2neg_Aborted_FINALwarm-up_synthetic-dataDPO_turn_allneg_old_and_newKTO_Aborted_same_family_model_FINALDPO_1neg_Aborted_same_family_model_old_LASFT-Final-DatasetKTO_Aborted_FDPO_2neg_Aborted_same_family_model_old_LADPO_allneg_Aborted_old_and_new_expDNLI
Adam Ek, Bill Noble, Stergios Chatzikyriakidis, Robin Cooper, Simon Dobnik, Eleni Gregoromichelaki, Christine Howes, Staffan Larsson, Vladislav Maraev, Gregory Mills, and Gijs Wijnholds. 2024. I hea- umm think that’s what they say: A Dataset of Inferences from Natural Language Dialogues. In Proceedings of the 28th Workshop on the Semantics and Pragmatics of Dialogue.
