datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-pragmatics
Citation
@misc{hu2023finegrainedcomparisonpragmaticlanguage, title={A fine-grained comparison of pragmatic language understanding in humans and language models}, author={Jennifer Hu and Sammy Floyd and Olessia Jouravlev and Evelina Fedorenko and Edward Gibson}, year={2023}, eprint={2212.06801}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2212.06801},}
clembench_REFUEL_0.7clembench_REFUEL_SFT_3clembench_REFUEL_SFT_1_iter2_DPO_turnKTO_Aborted_FINALDPO_turn_bug
Dataset Details
This is the training dataset for the DPO Turn strategy in PLAYPEN: An Environment for Exploring Learning From Dialogue Game Feedback.
These preference dataset have been obtained from these games' instances using this script, with --preference_depth turn.
Given the huge number of chosen vs rejected pairs in the first turn of the conversation, we limit the numbers of chosen and rejected pairs for the first turn to 10k samples (--first_turn_limit True).… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/DPO_turn_bug.clembench_REFUEL_SFT_1KTO_Aborted_best_models_Fwarm-up_synthetic-dataDPO_turn_allneg_old_and_newKTO_Aborted_same_family_model_FINALSFT-Final-DatasetKTO_Aborted_FDPO_allneg_Aborted_old_and_new_expclembench_REFUEL_base_1KTO_Aborted_same_family_model_FDPO_allneg_Aborted_old_and_new_10Klimitclembench_REFUEL_base_1_DPO_allneg_Aborted_old_and_new_exp_noexpclembench_REFUEL_SFT_1_KTO_FINALclembench_REFUEL_0.7_DPO_turn_solved_oldclembench_REFUEL_SFT_1_iter2turn-level-DPODPO_turn_allneg_old_6mcladderTaken from https://huggingface.co/datasets/causal-nlp/CLadder (the original was broken as of and not usable within the lm_eval_harness)
Citation
@inproceedings{jin2023cladder, author = {Zhijing Jin and Yuen Chen and Felix Leeb and Luigi Gresele and Ojasv Kamal and Zhiheng Lyu and Kevin Blin and Fernando Gonzalez and Max Kleiman-Weiner and Mrinmaya Sachan and Bernhard Sch{"{o}}lkopf}, title = "{CL}adder: {A}ssessing Causal Reasoning in Language Models", year = "2023"… See the full description on the dataset page: https://huggingface.co/datasets/clembench-playpen/cladder.DPO_turn_allneg_old
