datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zhtw-roleplay-space-grimoire
Space Grimoire RP Corpus (Traditional Chinese)
Speaker-attributed dialogue from the original novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage; 283 chapters, ~1.7M characters), cut into scenes and assembled into ShareGPT-style role-play training data. The novel and this dataset are the work of 睡半夜怎麼三更, who holds the copyright and has no exclusive platform agreement. Data: CC BY 4.0. Code: Apache 2.0.
中文說明在下方
Dataset Summary
Source text
283… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/zhtw-roleplay-space-grimoire.ebm-role-play-datasetMoral-RolePlay
Moral RolePlay
Paper | Code & Project Page
Abstract
Large Language Models (LLMs) are increasingly tasked with creative generation, including the simulation of fictional characters. However, their ability to portray non-prosocial, antagonistic personas remains largely unexamined. We hypothesize that the safety alignment of modern LLMs creates a fundamental conflict with the task of authentically role-playing morally ambiguous or villainous characters. To investigate this… See the full description on the dataset page: https://huggingface.co/datasets/Zihao1/Moral-RolePlay.RolePlay_Collection_random_ShareGPT
RolePlay_Collection_random_ShareGPT
This collection of random datasets for roleplay has been gathered from various sources. It has been processed, cleaned, and grammar-checked. However, it still requires significant work to be fully usable. Additional cleaning is necessary. I hope this proves helpful to someone.
chinese-roleplayfactual-multiagent-roleplay-ft-ru
march228/factual-multiagent-roleplay-ft-ru
Небольшой русскоязычный synthetic finetuning dataset для обучения модели следованию ролевым системным инструкциям личности при сохранении фактической опоры на контекст.
Что это за датасет
Этот набор сделан как instruction / finetuning dataset, а не как benchmark.
В каждой записи есть:
плотный system с персоной и тоном;
context, на который нужно опираться;
пользовательский question;
внутренние thoughts;
финальный answer.… See the full description on the dataset page: https://huggingface.co/datasets/march228/factual-multiagent-roleplay-ft-ru.chai-roleplay-labelednsfw: true means safe, false means nsfw
context_ratio: content between ** / total char
total_token_length: total amount of token from memory,prompt, chat_history, simply added, didn't remove repetitive tokens
vicgalle__Roleplay-Llama-3-8B-details
Dataset Card for Evaluation run of vicgalle/Roleplay-Llama-3-8B
Dataset automatically created during the evaluation run of model vicgalle/Roleplay-Llama-3-8B
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__Roleplay-Llama-3-8B-details.bunnycore__SmolLM2-1.7B-roleplay-lora-details
Dataset Card for Evaluation run of bunnycore/SmolLM2-1.7B-roleplay-lora
Dataset automatically created during the evaluation run of model bunnycore/SmolLM2-1.7B-roleplay-lora
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__SmolLM2-1.7B-roleplay-lora-details.fhai50032__RolePlayLake-7B-details
Dataset Card for Evaluation run of fhai50032/RolePlayLake-7B
Dataset automatically created during the evaluation run of model fhai50032/RolePlayLake-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__RolePlayLake-7B-details.roleplay-test
roleplay-test
Test split for lm-eval-harness.
