datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildchat-1m-tagged
WildChat 1M with tagging
This dataset replicates the category annotation process introduced in the paper named self-taught evaluator.
This dataset additionally contains three categories, i.e., category, complexity, and length, annotated by Mistral-7B-Instruct-v0.3 for WildChat-1M sessions.
I do my best to follow the technical details explained in the above paper while some difference is inevitably made due to the computational constraint.
The difference and notes for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/wildchat-1m-tagged.ro-WildChatThis dataset is a translation of allenai/WildChat using LLMic, a bilingual Romanian-English LLM.
WildChat is a collection of 650K conversations between human users and ChatGPT.
License: ODC-BY
@inproceedings{
zhao2024wildchat,
title={WildChat: 1M Chat{GPT} Interaction Logs in the Wild},
author={Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng},
booktitle={The Twelfth International Conference on Learning Representations},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-WildChat.
