datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script.
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.Long-Data-Col-rp_pile_pretrain
Dataset Card for "Long-Data-Col-rp_pile_pretrain"
This dataset is a subset of togethercomputer/Long-Data-Collections, namely the rp_sub.jsonl.zst and pile_sub.jsonl.zst files from the pretrain split.
Like the source dataset, we do not attempt to modify/change licenses of underlying data. Refer to the source dataset (and its source datasets) for details.
changes
as this is supposed to be a "long text dataset", we drop all rows where text contains <= 250 characters.… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/Long-Data-Col-rp_pile_pretrain.R-PRM
📘 R-PRM Dataset (SFT + DPO)
This dataset is developed for training Reasoning-Driven Process Reward Models (R-PRM), proposed in our ACL 2025 paper. It consists of two stages:
SFT (Supervised Fine-Tuning): collected from strong LLMs prompted with limited annotated examples, enabling reasoning-style evaluation.
DPO (Direct Preference Optimization): constructed by sampling multiple reasoning trajectories and forming preference pairs without additional labels.
These datasets are used… See the full description on the dataset page: https://huggingface.co/datasets/kevinpro/R-PRM.gpio-llm-rpi5-actions
GPIO-LLM: Raspberry Pi 5 GPIO request-to-action dataset
Requests to a Raspberry Pi 5 in plain English, paired with the structured, validated GPIO action a
small on-device model should produce: a hardware operation, a clarifying question when the pin or device
is unknown, or a refusal when the request is invalid or unsafe. It was built to train a ~20M-parameter
English model that runs offline on the Pi.
Safety. Model output must never drive hardware directly. Every action is… See the full description on the dataset page: https://huggingface.co/datasets/AwaleSagar/gpio-llm-rpi5-actions.novel-rp
Novel-RP: Multilingual Novel Role-Playing Dataset
A multilingual novel-based role-playing dataset for training and evaluating LLMs on character persona simulation.
📖 Overview
Novel-RP is a multilingual role-playing dataset built from web novels and role-playing conversations, specifically designed for training large language models on character role-playing tasks.
This dataset contains two main subsets:
train: Novel-based role-playing data (ShareGPT format) - from… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-rp.NSFW_RP_Format_DPOThis dataset aims to align a model to output the most common roleplaying format: "dialogue" *action*
This dataset contains NSFW content.
rpguildData scraped from roleplayerguild and parsed into prompts with a conversation history and associated character bio. Thanks to an anonymous internet stranger for the original scrape.
As usernames can be associated with multiple character biographies, assignment of characters is a little fuzzy. The char_confidence feature reflects how likely this assignment is to be correct. Not all posts in the conversation history necessarily have an associated character name. The column has_nameless reflects… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/rpguild.rp_books-en
Dataset Card for "rp_books-en"
Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections
The default config:
Dataset({
features: ['meta', 'text'],
num_rows: 26372
})
token count
default
GPT-4 tiktoken token count:
token_count
count 2.637200e+04
mean 1.009725e+05
std 1.161315e+05
min 3.811000e+03
25% 3.752750e+04
50% 7.757950e+04
75% 1.294130e+05
max 8.687685e+06
Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.Japanese-RP-Bench-testdata-SFW
Japanese-RP-Bench-testdata-SFW
本データセットは、LLMの日本語ロールプレイ能力を計測するベンチマークJapanese-RP-Bench用の評価データセットです。
ベンチマークの詳細については記事を参照してください。
データの概要
本データは以下のようなキーを持ちます。
genre: ロールプレイのジャンル
tag: ロールプレイの年齢区分
world_setting: ロールプレイの世界観設定
scene_setting: ロールプレイのシーン設定
user_setting: ロールプレイのユーザー側キャラクター設定
assistant_setting: ロールプレイのアシスタント側キャラクター設定
dialogue_tone: ロールプレイの対話のトーン
first_user_input: ロールプレイの最初のユーザー発話
response_format: ロールプレイの応答形式
id: データのid
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-RP-Bench-testdata-SFW.RPRevamped-Small
RPRevamped-Small-v1.0
Dataset Description
RPRevamped is a synthetic dataset generated by various numbers of models. It is very diverse and is recommended if you are fine-tuning a roleplay model. This is the Small version with Medium and Tiny version currently in work.
Github: RPRevamped GitHub
Here are the models used in creation of this dataset:
DeepSeek-V3-0324
Gemini-2.0-Flash-Thinking-Exp-01-21
DeepSeek-R1
Gemma-3-27B-it
Gemma-3-12B-it
Qwen2.5-VL-72B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/TechPowerB/RPRevamped-Small.NPC-RP-Post-Thinking
AIIDE-POST-THINKING
This is the post-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026
Dataset Details
This is the training dataset for chimbiwide/Gemma3-4B-post-thinking
Corresponding Links
Repository: [To be updated]
Paper: [To be updated]
Demo: [To be updated]
Uses
Suprevised-Finetuning
Dataset Creation
For more details, consult our paper.
Citation
If you… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-Post-Thinking.agent-misalignment-dataset
Agent Misalignment Dataset v0.1.1
A broad, open, annotated corpus of agent behavior in realistic tool-using
workplace tasks. 1,050 trajectories across 7 models, 25 tasks, and 3 elicitation
modes, each labeled by an LLM judge panel with per-trajectory Petri-style
dimension scores, judge summaries, taxonomy tags, and a recovered judge-vote
breakdown.
This is a v0.1.1 release. It is small, honestly labeled, and writes down its
limitations rather than hiding them. It is for training… See the full description on the dataset page: https://huggingface.co/datasets/rpotham/agent-misalignment-dataset.nemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.Rosebleu-1on1-Dialogues-RP
Rosebleu-1on1-Dialogues-RP
2025/05/17 3人での対話のデータを追加&無駄な改行の削除
@matsuxrさんが公開しているRosebleuデータセットを加工したAratako/Rosebleu-1on1-Dialoguesを元に、キャラクターや作品の設定などを付け加えたうえで、ロールプレイ的な文脈になるように加工したデータセットです。
LLMのファインチューニングにおけるロールプレイングタスクの学習を想定しています。
OpenAI APIのようにroleとcontentのペアの形式となっており、tokenizer.apply_chat_template()によって簡単に各モデルのチャットテンプレートのデータセットへと変換可能です。
データセットの詳細
各キャラの設定や各作品の世界観・あらすじなどをWikipediaやニコニコ大百科からまとめ、ロールプレイ向けにシステムメッセージへと埋め込んでいます。
現在、以下の2パターンのデータセットを用意してあります。主に地の文の処理方法が異なります。… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Rosebleu-1on1-Dialogues-RP.bluemoon-fandom-1-1-rp-jp-translated
bluemoon-fandom-1-1-rp-jp-translated
A subset of Squish42/bluemoon-fandom-1-1-rp-cleaned translated to Japanese using command-r-08-2024.
Misc. info
I used openrouter's api for inference with command-r-08-2024. Doing so is roughly 4x quicker than running the model locally, doesn't use up 95% of my vram, and doesn't make my 3090 as loud as my neighbours.
I decided to use command-r-08-2024 because it is completely uncensored for nsfw translation and provides translation… See the full description on the dataset page: https://huggingface.co/datasets/joujiboi/bluemoon-fandom-1-1-rp-jp-translated.Qwill-RP-CreativeWriting-Reasoning
Qwill RP CreativeWriting Reasoning Dataset
📝 Dataset Summary
Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes:
Reasoning, wrapped in <think>...</think>
Final Answer, wrapped in <answer>...</answer>
The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.rp-reasoning
RP Reasoning Traces
Multi-turn roleplay conversations augmented with structured reasoning traces, designed
to teach models to maintain narrative coherence over long RP sessions.
What This Is
Each conversation from Gryphe/Sonnet3.5-Charcard-Roleplay
has been processed to inject <think> blocks before assistant turns. These blocks
contain a "writer's scratchpad" — the kind of running state a skilled collaborative
fiction writer would mentally track to keep a scene coherent… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/rp-reasoning.NPC-RP-Pre-Thinking
AIIDE-PRE-THINKING
This is the pre-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026
Dataset Details
This is the training dataset for chimbiwide/Gemma3-4B-pre-thinking
Corresponding Links
Repository: [To be updated]
Paper: [To be updated]
Demo: [To be updated]
Uses
Suprevised-Finetuning
Dataset Creation
For more details, consult our paper.
Citation
If you… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-Pre-Thinking.NPC-RP-No-Thinking
AIIDE-NO-THINKING
This is the no-thinking dataset for our paper accepted as a poster presentation to AIIDE-2026
Dataset Details
This is the training dataset for chimbiwide/Gemma3-4B-no-thinking
Corresponding Links
Repository: [To be updated]
Paper: [To be updated]
Demo: [To be updated]
Uses
Suprevised-Finetuning
Dataset Creation
For more details, consult our paper.
Citation
If you find… See the full description on the dataset page: https://huggingface.co/datasets/chimbiwide/NPC-RP-No-Thinking.kaidol-phase2-rp-base-v0.3
KAIDOL Phase 2 RP Base Dataset v0.3
Dataset Description
KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality.
What's New in v0.3
GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns
Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.rpguild_processedpreprocessed version of chargoddard/rpguild into a fun little prompt format for finetuning
exp_rpt_manybugs-v2
exp_rpt_manybugs-v2
A fixed, non-gameable release of DCAgent/exp_rpt_manybugs:
164 C bug-repair tasks (ManyBugs) packaged as Harbor sandbox tasks.
Tasks
164
Projects
php (102), libtiff (24), python (15), wireshark (7), lighttpd (9), gzip (5), gmp (2)
Unique environments
7 snapshots (ubuntu:22.04 base, one -dev set per project)
Format
tasks.parquet — columns path (task id), task_binary (gzipped task tarball)
Each task contains a single buggy .c file from a… See the full description on the dataset page: https://huggingface.co/datasets/laion/exp_rpt_manybugs-v2.wataoshi-dialogues-rpこのデータセットは「私の推しは悪役令嬢。」のアニメから少しクリーニングされたセリフです。私はこのアニメの権利を持ってません、データセットの使い方について、責任がない。
Userは大体レイが言ったセリフが、他のキャラも含めてる。Assistantはクレアの答え。
RPG_Scenes_LLM_SyntheticThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
RPG Scene Generator
Dataset Summary
A collection of 100,000 procedurally generated fantasy RPG scenes, richly structured and atmosphere-enhanced. Each scene features narrative depth, modular logic, and multi-sensory context, making this dataset ideal for AI storytelling, tabletop roleplaying, video game prototyping… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_Scenes_LLM_Synthetic.test_rpFormat to have User, Assistant in order.
def merge_roles(data):
merged_data = []
current_role = None
current_content = []
for entry in data["messages"]:
# print(entry)
role = entry['role']
if role == "system":
role = "user"
content = entry['content']
if role == current_role:
current_content.append(content)
else:
ifcurrent_role is not None:
merged_data.append({"role":… See the full description on the dataset page: https://huggingface.co/datasets/aipracticecafe/test_rp.RPG_DM_Simulation_Combat_LLM_Trainingname: RPG_DM_Simulation_Combat_LLM_Training
pretty_name: Magician MUD Conversations
description:
20 turn-by-turn gameplay conversations from a text-based dungeon crawler RPG (MUD style).
Each conversation captures strategic decision-making in fantasy combat, including player status,
enemy encounters, resource management, and combat outcomes. Ideal for fine-tuning language models
for RPG dialogue generation, tactical decision-making, and game state understanding.
Get the full 30K… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/RPG_DM_Simulation_Combat_LLM_Training.Quasar-RP-DPO-v0.1
Quasar-RP-DPO-v0.1 Dataset Card
Overview
The Quasar-RP-DPO-v0.1 dataset is a collection of roleplaying preference data designed primarily for training and aligning language models. With a focus on improving creative outputs such as prose, poetry, and overall narrative creativity, this dataset is also well-suited for developing reward models and preference optimization techniques (e.g., Direct Preference Optimization or DPO).
Dataset Details
Total Rows: 2642… See the full description on the dataset page: https://huggingface.co/datasets/QuasarResearch/Quasar-RP-DPO-v0.1.bluemoon-fandom-1-1-rp-jp-translated-v2Reattempt at what I did with bluemoon-fandom-1-1-rp-jp-translated v1.
This dataset has 538 conversations and 9606 messages, making this dataset about 15% bigger.
I used deepseek-v3.2-exp from translation this time.
SynthRP-RpR-convertedFork of SynthRP by Epiculous with reasoning traces backadded via ArliAI's RpR conversion script.
The dataset is in Axolotl's input-output format, since otherwise reasoning won't be properly trained.
To be used, you first need to manually add chat template to it. A script for adding ChatML formatting is provided in the repo.
Example of usage in Axolotl:
datasets:
- path: ./synthrp_with_uuids-segments-full.jsonl
type: input_output
This dataset was provided to Allura by OwenArli.… See the full description on the dataset page: https://huggingface.co/datasets/allura-org/SynthRP-RpR-converted.prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can be… See the full description on the dataset page: https://huggingface.co/datasets/RParslow/prompts.chat.
