CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lemonilia /roleplaying-forums-raw Roleplaying forum scrapes (raw) Here are mostly original/raw files for some of the roleplaying forums I scraped in the past (and some newly scraped ones), repacked as HTML strings + some metadata on a one-row-per-thread basis instead of a one-row-per-message basis, which should make them more convenient to handle. Unlike the previously uploaded archive, they shouldn't have issues with spaces between adjacent HTML tags, as that occurred by mistake in an intermediate processing step… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/roleplaying-forums-raw.text100K<n<1M8 likes1.4k downloads2y agoHugging Face02beyoru /Aesir-Character-CoT-roleplay Overview Think with your role. Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character. Continue updating until money run out, I will try to update this dataset in near future Stats 1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content) ~14,349 assistant turns, each with full character-POV reasoning Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.tabulartext-generation1K<n<10K32 likes1.2k downloads5mo agoHugging Face03silk-road /ChatHaruhi-RolePlaying ChatHaruhi Reviving Anime Character in Reality via Large Language Model Chat-Haruhi-Suzumiyais a language model that imitates the tone, personality and storylines of characters like Haruhi Suzumiya, https://github.com/LC1332/Chat-Haruhi-Suzumiya Using this to load character and chat with him/her from ChatHaruhi import ChatHaruhi chatbot = ChatHaruhi( role_from_hf = "silk-road/ChatHaruhi-RolePlaying/haruhi",\ llm = 'openai' ,\… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-RolePlaying.text10K<n<100K16 likes766 downloads3y agoHugging Face04agentlans /combined-roleplay Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama 3.1… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/combined-roleplay.texttext-generation1M<n<10M23 likes685 downloads2y agoHugging Face05shibing624 /roleplay-zh-sharegpt-gpt4-data roleplay 数据集 数据 我们有4个数据集文件: "sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.texttext-generation1K<n<10K73 likes545 downloads2y agoHugging Face06IlyaGusev /gpt_roleplay_realm GPT Role-play Realm Dataset: The AI-generated character compendium This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play. 219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4. 20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.imagetext-generationn<1K105 likes529 downloads2y agoHugging Face07Gryphe /Sonnet3.5-Charcard-Roleplay⚠️ WARNING ⚠️ Many of these simulated character cards are highly NSFW in nature and may potentially describe disturbing scenes. Consider yourself very thoroughly warned! 9736 carefully simulated character card-based roleplay dialogues produced using an unrestrained Sonnet 3.5, now available as a ShareGPT dataset. Enjoy. How this dataset was produced Each card was enriched with a simulated user, which was either male or female with four distinct personalities. An effort was made to… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Sonnet3.5-Charcard-Roleplay.texttext-generation1K<n<10K96 likes498 downloads2y agoHugging Face08MiniMaxAI /role-play-bench Role-play Benchmark A comprehensive benchmark for evaluating Role-play Agents in Chinese and English scenarios. Dataset Summary Role-play Benchmark is designed to evaluate Role-play Agents' ability to deliver immersive role-play experiences through Situated Reenactment. Unlike traditional benchmarks with verifiable answers, Role-play is fundamentally non-verifiable, e.g., there's no single "correct" response when a tsundere character is asked "Do you like me?".… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/role-play-bench.tabulartext-generation1K<n<10K151 likes484 downloads8mo agoHugging Face09dasodefa /role-play-dataimagen<1K1 likes483 downloads4h agoHugging Face10lazyweasel /roleplay-bench RP-Bench: Roleplay Quality Benchmark for LLMs A multi-dimensional evaluation framework for measuring how well LLMs perform in roleplay scenarios — not just writing quality, but character consistency, user agency respect, lorebook integration, temporal reasoning, and genre-specific craft. The LLM-as-judge signals in this benchmark disagree with real users about half the time. We're calibrating against human preferences via a public blind-arena. Help out at arena.l3vi4th4n.ai — each… See the full description on the dataset page: https://huggingface.co/datasets/lazyweasel/roleplay-bench.tabulartext-generation1K<n<10K5 likes481 downloads5mo agoHugging Face11Syntaxdevloperangraeactionrpkaralho /action-roleplay-data Action Roleplay Data Data package for the Action SA-MP Android client. The client connects to 92.119.165.177:5636. The files/ directory contains the extracted game data, cache.zip is the archive consumed by the initial installer, files.json is the file-by-file manifest, and client_config.json contains the public endpoints. Runtime logs were excluded from the distributable package. The APK included here is a debug build for testing and is signed with a debug key. geospatialn<1K0 likes316 downloads20d agoHugging Face12AlekseyKorshuk /gpt-roleplay-realm-chatml Follow me HuggingFace: https://huggingface.co/AlekseyKorshuk GitHub: https://github.com/AlekseyKorshuk Twitter / X: https://x.com/alekseykorshuk text1K<n<10K5 likes279 downloads2y agoHugging Face13zerofata /Roleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters. The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.textn<1K31 likes258 downloads1y agoHugging Face14Johnson8187 /role-play-chinese繁體中文 English Role-Play Chinese Dataset 簡介 這是一個專為角色扮演對話設計的中文數據集,數據由 AI 生成,適用於訓練和評估自然語言處理(NLP)模型,特別是對話生成和角色扮演相關的任務。數據集以 Alpha 格式 儲存,方便進行微調和進一步的模型訓練。數據集包含多種場景和角色設定,能夠幫助模型學習如何在不同的情境下生成符合角色性格和背景的對話。 數據集結構 數據以Alpha格式儲存方便微調,包含以下字段: instruction: 任務指令,描述模型需要完成的任務。 input: 輸入內容,包含場景描述、過去的對話以及當前對話的上下文。 output: 期望的模型輸出,即符合角色設定的回應。 system: 角色設定和背景故事,幫助模型理解角色的性格和行為模式。 範例 { "instruction": "在給定的場景中,請根據角色設定回應對話。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/role-play-chinese.texttext-generation10K<n<100K5 likes247 downloads9d agoHugging Face15ShiniChien /reddit_creepypasta_roleplaytext100K<n<1M0 likes237 downloads7mo agoHugging Face16rickRossie /bluemoon_roleplay_chat_data_300k_messages Dataset Card for "bluemoon_roleplay_chat_data_300k_messages" More Information needed text100K<n<1M101 likes232 downloads3y agoHugging Face17beezza /japanese-young-character-roleplay-vlm Japanese Young Character Roleplay VLM 日本語のVLM向けに作られた、画像接地型ロールプレイSFTデータセットです。幼い雰囲気の 完全な架空キャラクターが、自分を名前で呼びながら、画像について安全な日常会話を 続けます。 既存作品のキャラクター、実在人物、Webから取得した画像は一切含みません。画像・会話・ メタデータは、このリポジトリの決定的な手続き生成器だけで作成しています。 規模 split rows assistant turns train 98,000 294,000 validation 1,000 3,000 test 1,000 3,000 total 100,000 300,000 各行には、256×192 WebP画像が1枚と、systemを含む7メッセージ (user/assistant 3往復)が入っています。48の架空名、12の場面、16の物体、8色、 6タスクを決定的に組み合わせています。… See the full description on the dataset page: https://huggingface.co/datasets/beezza/japanese-young-character-roleplay-vlm.imageimage-text-to-text100K<n<1M0 likes217 downloads2mo agoHugging Face18LooksJuicy /Chinese-Roleplay-CAItext1K<n<10K17 likes207 downloads2y agoHugging Face19LooksJuicy /Chinese-Roleplay-Novel一直以来,中文角色扮演开源数据集更关注超拟人方向或纯角色对话方向,严重缺乏交互游戏方向的开源数据,因此许多模型尤其参数量较小的模型对酒馆类的角色卡支持较差。 为了解决这一困境,本项目抛砖引玉,基于4500条小说文本使用GPT4o构建出约260条酒馆style的数据集,均为多轮对话,每轮对话都包括状态数据,如时间、角色状态、任务进度等。 数据key对应含义如下: world:表示当前故事的世界观,通常可以加入到system prompt中 scence:表示当前故事发生场景,包括时间、地点、环境、任务目标 character:表示当前故事中可能出现的角色和对应简介 field:表示这条数据每轮对话中需要生成的状态信息 conversations:表示这条数据的对话内容,分为问候语、主角(user)和系统(assistant) fields_format:表示状态信息的填充格式prompt,可能是列表、表格、JSON等各种形式 format_list:表示状态信息的填充结果 状态信息的示例如下 **健康状态**: 🌿 良好,身体颤抖 **精神状态**: 🌟 恐惧,极度紧张… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-Novel.textn<1K88 likes189 downloads2y agoHugging Face20CausalLM /Kingfall-Roleplay Kingfall-Roleplay Dataset Summary CausalLM/Kingfall-Roleplay is a preview subset of a larger synthetic corpus generated with Gemini Kingfall. This release contains 10K adapted samples selected for public preview and research use. It is not the full Kingfall-generated corpus, nor is it a release of the original unmodified data. Kingfall refers here to a reported confidential Gemini-family model that briefly became accessible during a limited availability window. Community… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Kingfall-Roleplay.text10K<n<100K19 likes184 downloads4mo agoHugging Face21hieunguyenminh /roleplay 🎭 Roleplay TTL Let AI be any characters you want to play with! Dataset Overview This dataset trains conversational AI to embody a wide range of original characters, each with a unique persona. It includes fictional characters, complete with their own backgrounds, core traits, relationships, goals, and distinct speaking styles. Dataset Details Curated by: Hieu Minh Nguyen Language(s) (NLP): Primarily English (with potential for multilingual extensions) License:… See the full description on the dataset page: https://huggingface.co/datasets/hieunguyenminh/roleplay.texttext-generation1K<n<10K92 likes158 downloads3y agoHugging Face22LooksJuicy /Chinese-Roleplay-SingleTurn请注意,个人模型经过characterEval的reward model进行DPO训练,因此使用本数据集进行SFT的模型在该榜单上会存在bias,导致分数异常偏高,请勿直接使用该榜单进行测试 简介 因已找到更优数据合成方案,为填充中文角色扮演数据集的空白,现开源部分中文角色扮演单轮对话数据集。 使用Refined-Anime-Text作为system prompt,使用小黄鸡随机query作为输入,调用个人角色扮演模型作为输出。 已处理为alpaca数据格式,方便大家处理和训练。经过验证,仅使用该数据集进行Lora微调即可获取一个效果还不错的模型~ chatGPT对比 character question answer_us answer_chatGPT 黑须彼方是(省略……)黑须彼方有着许多有趣的爱好和特点。她是一个有点毒舌的人,但总能犀利地指出问题所在。她有着敏锐的洞察力,擅长看透人心。她经常以此来捉弄加贺正午。她与正午有着相同的口癖,张扬的性格(省略……)她的个性和爱好使她成为一个备受喜爱的角色。… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-SingleTurn.texttext-generation1K<n<10K41 likes153 downloads2y agoHugging Face23dim /roleplay_instruct_v2_final Dataset Card for "roleplay_instruct_v2_final" More Information needed text1K<n<10K4 likes141 downloads3y agoHugging Face24Tarklanse /Traditional_Chinese_roleplay_chat_Dataset Traditional_Chinese_roleplay_chat_Dataset 這個資料集是以繁體中文為主,將各種由ChatGPT生成與極小部分個人撰寫的對話內容整理為alpaca dataset format的格式 以一層一層堆疊的方式,將一則對話紀錄拆成數筆資料(共約1000則對話),在幾次嘗試性的訓練中能夠讓llama2重現原本英文那種很活躍的對話風格,並且能夠維持善於扮演各種角色的能力 目前個人有以這個資料集製作一個lora 2023/09/07 更新 為資料集加入一些中英翻譯的句子,以期AI能以更好的文字去描寫他的動作,並增加了一些與食物有關的對話,希望能降低AI生出奇怪食物名的機率 texttext-generation1K<n<10K42 likes140 downloads3y agoHugging Face25OmniAICreator /Japanese-Roleplay-Dialogues Japanese-Roleplay-Dialogues This is a dialogue corpus collected from Japanese role-playing forum (commonly known as "なりきりチャット(narikiri chat)"). Each record corresponds to a single thread. For the original version, no filtering has been applied. For the filtered version, the following filtering and cleaning conditions have been applied: If the number of unique poster in the posts of each record is 1 or less, delete the entire record. If the length of the posts is 10 or less, delete… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/Japanese-Roleplay-Dialogues.texttext-generation10K<n<100K17 likes136 downloads2y agoHugging Face26laion /emotional-roleplay-finetuning-dataset Artificial Voice Roleplay Dataset 67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character voice-direction captions with generated audio, across German, English, Spanish, and French (German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre, zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost, vampire …) and high-arousal emotional delivery (rage, fear, grief, menace). Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.audiotext-to-speech10K<n<100K4 likes122 downloads2mo agoHugging Face27asd567557275 /zhtw-roleplay-space-grimoire Space Grimoire RP Corpus (Traditional Chinese) Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0. 中文說明在下方 Dataset Summary Source text 283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.tabulartext-generation10K<n<100K1 likes121 downloads11d agoHugging Face28rx1lora /StoryPlay_RolePlay-NPCv2 RolePlay-NPCv2 The newest RP dataset containing some high-quality dataset for Gemma3NPC. We combined pippa, NPC-Dialogue_v2, Sonnet-Roleplay and ReLe_Synthetic_v1_json. WARNING -- Some conversations contain highly NSFW content, use it with caution! texttext-generation10K<n<100K1 likes118 downloads2mo agoHugging Face29Arketov /ru_roleplay_conversationlima, pipa и bluemoon. Переведены на русский, нуждаются в допополнтельной фильтрации. Длина некоторых последовательностей очень большая, а не которых очень маленькая. Есть шанс очень редких дубликатов. texttext-generation10K<n<100K3 likes115 downloads3y agoHugging Face30Exxe /literary-roleplay Dataset Card for Literary Roleplay SFT Dataset Summary An instruction-tuning dataset for training models to roleplay properly, derived from literary sources across five languages and three roleplay-engine logic frameworks. The dataset contains 346 rows spanning English (164), Russian (68), Hindi (38), Sanskrit (38), and Japanese (38), drawn from the works of Gogol, Bulgakov, Perumov, Golovachev, Vedic canon (Upanishads, Mahabharata, Ramayana), classic sci-fi… See the full description on the dataset page: https://huggingface.co/datasets/Exxe/literary-roleplay.texttext-generationn<1K0 likes110 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.