datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
anime-understanding-dataset
Anime Understanding Benchmark (WIP)
Evaluate anime knowledge found in existing LLMs. We hope to provide an easy to run evaluation on knowledge understanding in anime/manga. Better understanding in anime/manga knowledge should resulted in task such as waifu role play.
Any suggestion is open in discussion tab.
Currently in the works
[] Eval on popular models such as gpt, hermes, dolphin, llama base model
[] Add more metadata regarding of anime/manga year span
[] Suggestions… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/anime-understanding-dataset.Roleplay-Anime-CharactersA small synthetic (mostly) SFW dataset of (mostly) one on one character RP. Focuses on anime and game characters etc. This dataset tries to leverage the information larger models know about the characters to play them in character better than would normally be possible with generic characters.
The situations are generally absurd so the model is forced to generalize. It focuses on teaching the model how to be proactive, creative, emotional and take existing characters it may know about and… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Roleplay-Anime-Characters.anime-caption-danbooru-2021-sfw-5m-hq
Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq
Dataset Summary
This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated.
Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.Refined-Anime-TextA reupload of CausalLM/Refined-Anime-Text. Disclaimer: it's not a good dataset.
Legal to re-upload because the original license was wtfpl.
nyaa-anime-filenames
nyaa.si anime filenames
A snapshot of anime release filenames scraped from nyaa.si, covering the Anime categories English-translated, Non-English-translated, and Raw. It contains metadata only: filenames, file sizes, and post details. No torrent files, magnet links, or media content are included.
Snapshot date: 2026-09-16. 3,064 posts, 8,148 files after cleaning.
Files
File
Rows
Description
raw.jsonl
9,778
One row per file, exactly as listed on each… See the full description on the dataset page: https://huggingface.co/datasets/valentin-marquez/nyaa-anime-filenames.anime-waifu-personality-chat
Anime Waifu Personality
This dataset contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
Here's a few example:
{
"trait": "tsundere",
"dialogue": "H-Holding hands?! W-Well, I guess if you’re that desperate..."
},
{
"trait": "yandere",
"dialogue": "If I can't… See the full description on the dataset page: https://huggingface.co/datasets/Shxbhxm21/anime-waifu-personality-chat.Anime_novel_datasetsanimelist-datasetA JSON based anime dataset containing the most important meta data as well as cross references to various anime sites such as MAL, ANIDB, ANILIST, KITSU and more...
Credits: https://github.com/manami-project/anime-offline-database
Instruct-AnimeThis dataset is a mix of various open source datasets rewritten in the voice of anime characters. Generated using Gemini 2.5 Pro.
Credits go to the original dataset creators. Data was taken from (that I remember, the below datasets)
shisa-ai/shisa-v2-sharegpt
nvidia/Llama-Nemotron-Post-Training-Dataset (chat)
ConicCat/EzInstruct1-15K-MagpieProMT
Delta-Vector/Hydrus-AM-Thinking-IF-No-Think
anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
Anime-AMA-ProseDataset of anime / game characters and VTubers being asked various questions referencing their wiki page.
Prompts were generated by GLM 4.6
Scenes were generated by GLM 4.6
Response Plan was generated by GLM 4.6
Initial response was generated by GLM 4.6
Rewrite (if enough overused words were found) was done by Gemma 3 27b.
The rewrites are generally lower quality than the response provided by GLM, but they offer some different prose / word choices if preferred. Additionally, some rewrites… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Anime-AMA-Prose.AnimeMangaCharacters-247K
Anime Manga Characters Dataset
This dataset is a metafile containing information about 247,034 anime and manga characters sourced from 2,372 fandom wiki sites. Each entry represents a character along with its associated metadata. The dataset has been deduplicated based on the url field to avoid redundancy, although a single character may still have multiple associated URLs.
Potential Applications
Multimodal Data Creation: Use the URLs to download the respective wiki… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/AnimeMangaCharacters-247K.Instruct-Anime-CreativeWritingA small synthetic dataset of instruct prompts from various fandoms, answered by anime / game characters instructed to respond in a creative writing style.
Data was created by a mix of Claude 3.7, Deepseek-v3 & Gemini Flash.
Dataset has been cleaned for major slop. There's a decent amount of one word (testaments, fractures etc.) slop in here, but generally this one is pretty clean and has a positive impact on writing style for models it's applied to.
Creation Process:
Character is randomly… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Instruct-Anime-CreativeWriting.AnimeSubtitleRefined-Anime-Text
Refined Anime Text for Continual Pre-training of Language Models
This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.anime-waifu-personality-chat
Anime Waifu Personality
This dataset contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
Here's a few example:
{
"trait": "tsundere",
"dialogue": "H-Holding hands?! W-Well, I guess if you’re that desperate..."
},
{
"trait": "yandere",
"dialogue": "If I can't… See the full description on the dataset page: https://huggingface.co/datasets/qwertt2005/anime-waifu-personality-chat.anime-metaphor-engine
CatQualia anime metaphor engine — fiction mechanisms to engineering constructs
53,931 rows · 38,685,099 bytes · JSON Lines, one object per line.
What this is
Fiction mechanisms mapped to engineering isomorphisms with a buildable artifact and an explicit grade. Each row carries canon_rules — the rules the source fiction actually establishes — and an ASSUMED provenance note where the canon needs verification. The status field records whether the artifact was… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/anime-metaphor-engine.anime-unslop-10k~10k samples from CausalLM/Refined-Anime-Text passed through Claude 3.5 Sonnet to appear more human-like.
Summaries-Anime-FandomPagesA bunch of fandom summaries from scraped webpages.
Created as part of the process of making https://huggingface.co/datasets/zerofata/Instruct-Anime-CreativeWriting
None of these have been checked for quality but they were all made by Claude / Deepseek-v3 so quality on average is probably fine.
animepedia
Animepedia
A large-scale, multi-format English text dataset covering 22,099 unique anime titles — designed for language model training, instruction tuning, and encyclopedic knowledge injection. Every title is represented in at least three distinct writing styles, resulting in a structurally complete corpus with zero corrupt records and zero entries below the minimum quality threshold.
Dataset Description
Animepedia is a richly structured collection of… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/animepedia.Manga_ita_animeclickList of mangas (snapshot 28.02.2025) listed on Animeclick. The dataset contains the following details:
{
"url": "https://www.animeclick.it/manga/46610/restart-rui-si-yi-kai",
"titolo_originale": "Gui Ling",
"titolo_inglese": "Zona 0: Restart / Return To Zero",
"titolo_kanji": "归零",
"nazionalita": "Cina",
"casa_editrice": "KuaiKan Manhua",
"storia": "Rui Si",
"disegni": "Rui Si",
"anno": "2020",
"stato_patria": "completato",
"stato_italia": "in corso"… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Manga_ita_animeclick.danbooru-anime-screenshotactsrefined-anime-instruct-en-641k
Dataset Card for refined-anime-instruct-en-641k
Dataset Summary
This is 641,497 instructions for an expert model that knows about the following things:
Anime
Manga
Live Action Shows
Children's Films
Western Comics
Agatha Christie Novels and Adaptations (not sure why this is over-represented)
Video Games
It is derived from Refined-Anime-Text by filtering out all ZH entries. According to their README.md, these outputs are completions derived from GPT3.5 and GPT4.… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/refined-anime-instruct-en-641k.aditya2803_one-piece-anime
ONE PIECE ANIME
Complete dataset of one piece anime in .csv and .json format
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 4,105
Files: 2
Files
ONE PIECE.csv
One Piece json.json
Mirrored from Kaggle
anime-character-prompt-15kanimebench-alphaAlpha version of some form of anime-bench
** This needs more data, if you have anime QA in text format, send it
200 questions, would consider the easy version of anime bench, as its just anime trivia,
anime challenge should be deeper into the lore of several anime, maybe even iceberg chart stuff
dataset rn:
25% one piece QA
25% DBZ QA
50% Anime Quotes, guess the person who said it
not in MCQ format, no random guess baseline, if it knows, it knows
QA made by forcing kindly requesting gemini… See the full description on the dataset page: https://huggingface.co/datasets/VatsaDev/animebench-alpha.Anime-News
このアニメは女の子のかわい
popular_anime_corpus
popular_anime_corpus
有名なアニメのあらすじや主人公の情報をまとめたコーパスです。
1960年代以降の人気アニメに関するテキストをWikipediaから抽出し、テキストを要約して、Alpaca形式でコーパスにまとめています。下記のようなデータとなっています。
{
"train": [
{
"instruction": "主人公は誰ですか?",
"input": "名探偵コナン 絶海の探偵",
"output": "主人公は「江戸川 コナン(えどがわ コナン)」です。\n\n江戸川 コナンは、本来の姿は「東の高校生探偵」として名を馳せている工藤新一だが、黒ずくめの組織に飲まされた毒薬・APTX4869の副作用で小学生の姿になっている。本作の主人公であり、物語の中心人物として活躍します。"},
{
"instruction": "主人公やあらすじを教えてください。",
"input": "ヴァイオレット・エヴァーガーデン",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/kujirahand/popular_anime_corpus.
