datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roleplay 🎭 Roleplay TTL
Let AI be any characters you want to play with!
Dataset Overview
This dataset trains conversational AI to embody a wide range of original characters, each with a unique persona. It includes fictional characters, complete with their own backgrounds, core traits, relationships, goals, and distinct speaking styles.
Dataset Details
Curated by: Hieu Minh Nguyen
Language(s) (NLP): Primarily English (with potential for multilingual extensions)
License:… See the full description on the dataset page: https://huggingface.co/datasets/hieunguyenminh/roleplay.Cantonese-Prototype-Roleplay-Dataset
Prototype-Dataset-Cantonese
角色扮演訓練集(粵語版)
簡介
本數據集是基於 Moemu/Muice-Dataset 進行製作的粵語 (Cantonese) 衍生版本。旨在模擬二次元角色的語言風格,包含日常生活、情感話題、自我認知增強等約 3,500 條對話。
數據處理說明
語言轉換:將原有的簡體中文對話轉換為繁體粵語口語。
格式保持:嚴格遵循原版的多輪對話 JSONL 格式。
限制聲明
日常導向:本訓練集主要圍繞日常話題展開,對專業性問題(如代碼、高級推理)做了簡化處理,模型可能會產生事實性錯誤。
性格特徵:為了模仿動漫角色的說話風格,訓練集可能含有傲嬌、偏見或不禮貌的回答。如需構建高度安全性的模型,請謹慎使用或進行篩選。
倫理風險:使用本數據集訓練模型所引起的任何法律或倫理風險,由訓練者自行承擔。
許可與鳴謝
原作者: Moemu
粵語版本製作: AhYin
許可證:… See the full description on the dataset page: https://huggingface.co/datasets/AhYin/Cantonese-Prototype-Roleplay-Dataset.roleplayLLM_Chinese
该数据集主要用于对LLM进行角色扮演过程中的微调。
数据集仅适用于可以使用prompt进行微调的大语言模型。
数据集中大部分数据源自于Chinese-Roleplay-SingleTurn数据集,并且引入了部分《原神》与《崩坏 星穹铁道》中角色的对话内容进行构建。
数据集中包含了部分公开的文本内容,使用过程中应遵守开源协议,并且不可用于商用。
项目
roleplay_odiaThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Odia using OdiaGenAI English=>Indic translation app.
roleplay_hindiThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Hindi using OdiaGenAI English=>Indic translation app.
roleplay_englishllm-roleplay
