datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toxic-chat
Update
[01/31/2024] We update the OpenAI Moderation API results for ToxicChat (0124) based on their updated moderation model on on Jan 25, 2024.[01/28/2024] We release an official T5-Large model trained on ToxicChat (toxicchat0124). Go and check it for you baseline comparision![01/19/2024] We have a new version of ToxicChat (toxicchat0124)!
Content
This dataset contains toxicity annotations on 10K user prompts collected from the Vicuna online demo.
We utilize a human-AI… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/toxic-chat.lmsys_chatbot_arena_conversationsdatasource: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH
lmsys-kaggle
LMSYS/Kaggle Preference Data
This repository contains 91,667 rows in the default/train split and is viewable through the Hugging Face Dataset Viewer.
Important Provenance Notice
The repository currently does not include a source manifest, processing script, or exact upstream revision. Its name and schema are consistent with LMSYS/Kaggle preference data, but the row count should not be interpreted as proof that it is an unchanged copy of any single official… See the full description on the dataset page: https://huggingface.co/datasets/keithtyser/lmsys-kaggle.lmsys-chat-1m-GEO-2-25
📊 Understanding User Intent in Chatbot Conversations
This dataset, sections the LMSYS-Chat-1M dataset into GEO and Marketing relevant categories to help maketers access real life prompts and chatlogs that are relevant to them.
💡 Why I'm doing this
GEO (generative engine optimization) is a new marketing discipline, similar to SEO, that aims to measure and optimize the responses of LLMs and LLM-powered apps for marketing purposes.
The question of "search volume" in LLMs… See the full description on the dataset page: https://huggingface.co/datasets/GPAWeb/lmsys-chat-1m-GEO-2-25.lmsys-train
