CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4.9k downloads2y agoHugging Face02BEE-spoke-data /Long-Data-Col-rp_pile_pretrain Dataset Card for "Long-Data-Col-rp_pile_pretrain" This dataset is a subset of togethercomputer/Long-Data-Collections, namely the rp_sub.jsonl.zst and pile_sub.jsonl.zst files from the pretrain split. Like the source dataset, we do not attempt to modify/change licenses of underlying data. Refer to the source dataset (and its source datasets) for details. changes as this is supposed to be a "long text dataset", we drop all rows where text contains <= 250 characters.… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/Long-Data-Col-rp_pile_pretrain.texttext-generation10M<n<100M3 likes1.2k downloads9mo agoHugging Face03kevinpro /R-PRM 📘 R-PRM Dataset (SFT + DPO) This dataset is developed for training Reasoning-Driven Process Reward Models (R-PRM), proposed in our ACL 2025 paper. It consists of two stages: SFT (Supervised Fine-Tuning): collected from strong LLMs prompted with limited annotated examples, enabling reasoning-style evaluation. DPO (Direct Preference Optimization): constructed by sampling multiple reasoning trajectories and forming preference pairs without additional labels. These datasets are used… See the full description on the dataset page: https://huggingface.co/datasets/kevinpro/R-PRM.texttext-generation100K<n<1M1 likes352 downloads1y agoHugging Face04AwaleSagar /gpio-llm-rpi5-actions GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a small on-device model should produce: a hardware operation, a clarifying question when the pin or device is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter English model that runs offline on the Pi. Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.tabulartext-generation1M<n<10M0 likes189 downloads5d agoHugging Face05taozi555 /novel-rpgated Novel-RP: Multilingual Novel Role-Playing Dataset A multilingual novel-based role-playing dataset for training and evaluating LLMs on character persona simulation. 📖 Overview Novel-RP is a multilingual role-playing dataset built from web novels and role-playing conversations, specifically designed for training large language models on character role-playing tasks. This dataset contains two main subsets: train: Novel-based role-playing data (ShareGPT format) - from… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-rp.tabulartext-generation10K<n<100K11 likes147 downloads9mo agoHugging Face06ResplendentAI /NSFW_RP_Format_DPOThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action* This dataset contains NSFW content. texttext-generationn<1K82 likes133 downloads3y agoHugging Face07chargoddard /rpguildData scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape. As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/rpguild.texttext-generation100K<n<1M33 likes106 downloads3y agoHugging Face08BEE-spoke-data /rp_books-en Dataset Card for "rp_books-en" Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections The default config: Dataset({ features: ['meta', 'text'], num_rows: 26372 }) token count default GPT-4 tiktoken token count: token_count count 2.637200e+04 mean 1.009725e+05 std 1.161315e+05 min 3.811000e+03 25% 3.752750e+04 50% 7.757950e+04 75% 1.294130e+05 max 8.687685e+06 Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.texttext-generation100K<n<1M1 likes83 downloads9mo agoHugging Face09Aratako /Japanese-RP-Bench-testdata-SFW Japanese-RP-Bench-testdata-SFW 本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。 ベンチマークの詳細については記事を参照してください。 データの概要 本データは以下のようなキーを持ちます。 genre: ロールプレイのジャンル tag: ロールプレイの年齢区分 world_setting: ロールプレイの世界観設定 scene_setting: ロールプレイのシーン設定 user_setting: ロールプレイのユーザー側キャラクター設定 assistant_setting: ロールプレイのアシスタント側キャラクター設定 dialogue_tone: ロールプレイの対話のトーン first_user_input: ロールプレイの最初のユーザー発話 response_format: ロールプレイの応答形式 id: データのid ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.texttext-generationn<1K5 likes73 downloads2y agoHugging Face10TechPowerB /RPRevamped-Small RPRevamped-Small-v1.0 Dataset Description RPRevamped is a synthetic dataset generated by various numbers of models. It is very diverse and is recommended if you are fine-tuning a roleplay model. This is the Small version with Medium and Tiny version currently in work. Github: RPRevamped GitHub Here are the models used in creation of this dataset: DeepSeek-V3-0324 Gemini-2.0-Flash-Thinking-Exp-01-21 DeepSeek-R1 Gemma-3-27B-it Gemma-3-12B-it Qwen2.5-VL-72B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/TechPowerB/RPRevamped-Small.texttext-generation1K<n<10K1 likes68 downloads1y agoHugging Face11chimbiwide /NPC-RP-Post-Thinking AIIDE-POST-THINKING This is the post-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026 Dataset Details This is the training dataset for chimbiwide/Gemma3-4B-post-thinking Corresponding Links Repository: [To be updated] Paper: [To be updated] Demo: [To be updated] Uses Suprevised-Finetuning Dataset Creation For more details, consult our paper. Citation If you… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-Post-Thinking.texttext-generation1K<n<10K0 likes62 downloads29d agoHugging Face12rpotham /agent-misalignment-dataset Agent Misalignment Dataset v0.1.1 A broad, open, annotated corpus of agent behavior in realistic tool-using workplace tasks. 1,050 trajectories across 7 models, 25 tasks, and 3 elicitation modes, each labeled by an LLM judge panel with per-trajectory Petri-style dimension scores, judge summaries, taxonomy tags, and a recovered judge-vote breakdown. This is a v0.1.1 release. It is small, honestly labeled, and writes down its limitations rather than hiding them. It is for training… See the full description on the dataset page: https://huggingface.co/datasets/rpotham/agent-misalignment-dataset.tabulartext-classification1K<n<10K0 likes60 downloads4mo agoHugging Face13rpisano /nemotron-cc-atomic-simplification-gemma4-31b nemotron-cc atomic-statement simplification (Gemma 4 31B-it) 2,000,000 records: source text from nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object statements. Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to <=8192 templated tokens. Fields id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.texttext-generation1M<n<10M0 likes59 downloads16d agoHugging Face14Aratako /Rosebleu-1on1-Dialogues-RP Rosebleu-1on1-Dialogues-RP 2025/05/17 3人での対話のデータを追加&無駄な改行の削除 @matsuxrさんが公開しているRosebleuデータセットを加工したAratako/Rosebleu-1on1-Dialoguesを元に、キャラクターや作品の設定などを付け加えたうえで、ロールプレイ的な文脈になるように加工したデータセットです。 LLMのファインチューニングにおけるロールプレイングタスクの学習を想定しています。 OpenAI APIのようにroleとcontentのペアの形式となっており、tokenizer.apply_chat_template()によって簡単に各モデルのチャットテンプレートのデータセットへと変換可能です。 データセットの詳細 各キャラの設定や各作品の世界観・あらすじなどをWikipediaやニコニコ大百科からまとめ、ロールプレイ向けにシステムメッセージへと埋め込んでいます。 現在、以下の2パターンのデータセットを用意してあります。主に地の文の処理方法が異なります。… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Rosebleu-1on1-Dialogues-RP.texttext-generation1K<n<10K18 likes58 downloads2y agoHugging Face15joujiboi /bluemoon-fandom-1-1-rp-jp-translated bluemoon-fandom-1-1-rp-jp-translated A subset of Squish42/bluemoon-fandom-1-1-rp-cleaned translated to Japanese using command-r-08-2024. Misc. info I used openrouter's api for inference with command-r-08-2024. Doing so is roughly 4x quicker than running the model locally, doesn't use up 95% of my vram, and doesn't make my 3090 as loud as my neighbours. I decided to use command-r-08-2024 because it is completely uncensored for nsfw translation and provides translation… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/bluemoon-fandom-1-1-rp-jp-translated.tabulartext-generationn<1K3 likes53 downloads1y agoHugging Face16marcuscedricridia /Qwill-RP-CreativeWriting-Reasoning Qwill RP CreativeWriting Reasoning Dataset 📝 Dataset Summary Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes: Reasoning, wrapped in <think>...</think> Final Answer, wrapped in <answer>...</answer> The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.tabulartext-generation1K<n<10K8 likes53 downloads1y agoHugging Face17aimeri /rp-reasoning RP Reasoning Traces Multi-turn roleplay conversations augmented with structured reasoning traces, designed to teach models to maintain narrative coherence over long RP sessions. What This Is Each conversation from Gryphe/Sonnet3.5-Charcard-Roleplay has been processed to inject <think> blocks before assistant turns. These blocks contain a "writer's scratchpad" — the kind of running state a skilled collaborative fiction writer would mentally track to keep a scene coherent… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/rp-reasoning.texttext-generation1K<n<10K1 likes50 downloads6mo agoHugging Face18chimbiwide /NPC-RP-Pre-Thinking AIIDE-PRE-THINKING This is the pre-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026 Dataset Details This is the training dataset for chimbiwide/Gemma3-4B-pre-thinking Corresponding Links Repository: [To be updated] Paper: [To be updated] Demo: [To be updated] Uses Suprevised-Finetuning Dataset Creation For more details, consult our paper. Citation If you… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-Pre-Thinking.texttext-generation1K<n<10K0 likes47 downloads29d agoHugging Face19chimbiwide /NPC-RP-No-Thinking AIIDE-NO-THINKING This is the no-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026 Dataset Details This is the training dataset for chimbiwide/Gemma3-4B-no-thinking Corresponding Links Repository: [To be updated] Paper: [To be updated] Demo: [To be updated] Uses Suprevised-Finetuning Dataset Creation For more details, consult our paper. Citation If you find… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-No-Thinking.texttext-generation1K<n<10K0 likes44 downloads29d agoHugging Face20developer-lunark /kaidol-phase2-rp-base-v0.3 KAIDOL Phase 2 RP Base Dataset v0.3 Dataset Description KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality. What's New in v0.3 GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.texttext-generation10K<n<100K0 likes36 downloads9mo agoHugging Face21Fizzarolli /rpguild_processedpreprocessed version of chargoddard/rpguild into a fun little prompt format for finetuning texttext-generation10K<n<100K4 likes28 downloads2y agoHugging Face22laion /exp_rpt_manybugs-v2 exp_rpt_manybugs-v2 A fixed, non-gameable release of DCAgent/exp_rpt_manybugs: 164 C bug-repair tasks (ManyBugs) packaged as Harbor sandbox tasks. Tasks 164 Projects php (102), libtiff (24), python (15), wireshark (7), lighttpd (9), gzip (5), gmp (2) Unique environments 7 snapshots (ubuntu:22.04 base, one -dev set per project) Format tasks.parquet — columns path (task id), task_binary (gzipped task tarball) Each task contains a single buggy .c file from a… See the full description on the dataset page: https://huggingface.co/datasets/laion/exp_rpt_manybugs-v2.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face23aipracticecafe /wataoshi-dialogues-rpこのデータセットは「私の推しは悪役令嬢。」のアニメから少しクリーニングされたセリフです。私はこのアニメの権利を持ってません、データセットの使い方について、責任がない。 Userは大体レイが言ったセリフが、他のキャラも含めてる。Assistantはクレアの答え。 texttext-generationn<1K0 likes24 downloads2y agoHugging Face24CJJones /RPG_Scenes_LLM_SyntheticThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. RPG Scene Generator Dataset Summary A collection of 100,000 procedurally generated fantasy RPG scenes, richly structured and atmosphere-enhanced. Each scene features narrative depth, modular logic, and multi-sensory context, making this dataset ideal for AI storytelling, tabletop roleplaying, video game prototyping… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_Scenes_LLM_Synthetic.texttext-generation10K<n<100K0 likes24 downloads7mo agoHugging Face25aipracticecafe /test_rpFormat to have User, Assistant in order. def merge_roles(data): merged_data = [] current_role = None current_content = [] for entry in data["messages"]: # print(entry) role = entry['role'] if role == "system": role = "user" content = entry['content'] if role == current_role: current_content.append(content) else: ifcurrent_role is not None: merged_data.append({"role":… See the full description on the dataset page: https://huggingface.co/datasets/aipracticecafe/test_rp.texttext-generationn<1K0 likes22 downloads2y agoHugging Face26CJJones /RPG_DM_Simulation_Combat_LLM_Trainingname: RPG_DM_Simulation_Combat_LLM_Training pretty_name: Magician MUD Conversations description: 20 turn-by-turn gameplay conversations from a text-based dungeon crawler RPG (MUD style). Each conversation captures strategic decision-making in fantasy combat, including player status, enemy encounters, resource management, and combat outcomes. Ideal for fine-tuning language models for RPG dialogue generation, tactical decision-making, and game state understanding. Get the full 30K… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_DM_Simulation_Combat_LLM_Training.texttext-generation1K<n<10K0 likes20 downloads7mo agoHugging Face27QuasarResearch /Quasar-RP-DPO-v0.1 Quasar-RP-DPO-v0.1 Dataset Card Overview The Quasar-RP-DPO-v0.1 dataset is a collection of roleplaying preference data designed primarily for training and aligning language models. With a focus on improving creative outputs such as prose, poetry, and overall narrative creativity, this dataset is also well-suited for developing reward models and preference optimization techniques (e.g., Direct Preference Optimization or DPO). Dataset Details Total Rows: 2642… See the full description on the dataset page: https://huggingface.co/datasets/QuasarResearch/Quasar-RP-DPO-v0.1.texttext-generation1K<n<10K4 likes19 downloads2y agoHugging Face28joujiboi /bluemoon-fandom-1-1-rp-jp-translated-v2Reattempt at what I did with bluemoon-fandom-1-1-rp-jp-translated v1. This dataset has 538 conversations and 9606 messages, making this dataset about 15% bigger. I used deepseek-v3.2-exp from translation this time. tabulartext-generationn<1K0 likes18 downloads10mo agoHugging Face29allura-org /SynthRP-RpR-convertedFork of SynthRP by Epiculous with reasoning traces backadded via ArliAI's RpR conversion script. The dataset is in Axolotl's input-output format, since otherwise reasoning won't be properly trained. To be used, you first need to manually add chat template to it. A script for adding ChatML formatting is provided in the repo. Example of usage in Axolotl: datasets: - path: ./synthrp_with_uuids-segments-full.jsonl type: input_output This dataset was provided to Allura by OwenArli.… See the full description on the dataset page: https://huggingface.co/datasets/allura-org/SynthRP-RpR-converted.texttext-generation1K<n<10K4 likes16 downloads1y agoHugging Face30RParslow /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/RParslow/prompts.chat.textquestion-answering1K<n<10K0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.