datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
narutonarutoshippuden
Bangumi Image Base of Naruto Shippuden
This is the image base of bangumi Naruto Shippuden, we detected 196 characters, 36722 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/narutoshippuden.naruto-blip-captions
Dataset Card for Naruto BLIP captions
Dataset used to train TBD.
The original images were obtained from narutopedia.com and captioned with the pre-trained BLIP model.
For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the accompanying text caption. Only a train split is provided.
Example stable diffusion outputs
"Bill Gates with a hoodie", "John Oliver with Naruto style", "Hello Kitty with Naruto style", "Lebron… See the full description on the dataset page: https://huggingface.co/datasets/lambda/naruto-blip-captions.narutomovies
Bangumi Image Base of Naruto [movies]
This is the image base of bangumi NARUTO [Movies], we detected 37 characters, 3111 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/narutomovies.CharacterRAG
CharacterRAG dataset
Overview
CharacterRAG is a high-quality dataset for RAG-based role-playing agents, which consists of persona documents for 15 distinct fictional characters totaling 976K written characters, and 450 question–answer pairs.
All external information about works featuring characters that could affect persona consistency has been manually removed by human annotators (e.g., character popularity poll).
Dataset Structure
├── anya_forger
├──… See the full description on the dataset page: https://huggingface.co/datasets/naruto-soop/CharacterRAG.naruto-reddit-commentsThis dataset is extracted from Reddit datasets for research purpose
details_FelixChao__NarutoDolphin-10B
Dataset Card for Evaluation run of FelixChao/NarutoDolphin-10B
Dataset automatically created during the evaluation run of model FelixChao/NarutoDolphin-10B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_FelixChao__NarutoDolphin-10B.details_FelixChao__NarutoDolphin-7B
Dataset Card for Evaluation run of FelixChao/NarutoDolphin-7B
Dataset automatically created during the evaluation run of model FelixChao/NarutoDolphin-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_FelixChao__NarutoDolphin-7B.naruto-vision-promptsnaruto-blip-captions
Dataset Card for Naruto BLIP captions
Dataset used to train TBD.
The original images were obtained from narutopedia.com and captioned with the pre-trained BLIP model.
For each row the dataset contains image and text keys. image is a varying size PIL jpeg, and text is the accompanying text caption. Only a train split is provided.
Example stable diffusion outputs
"Bill Gates with a hoodie", "John Oliver with Naruto style", "Hello Kitty with Naruto style", "Lebron… See the full description on the dataset page: https://huggingface.co/datasets/Milabench/naruto-blip-captions.NarutoShinobiLordDadosfatima-fellowship-naruto-blip-captionsnaruto-shippuden-ultimate-ninja-storm-4-road-to-boruto-eu-en-1.05naruto-shippuden-ultimate-ninja-storm-trilogy-us-en-1.01naruto-shippuden-ultimate-ninja-storm-4-us-en-1.11naruto-x-boruto-ultimate-ninja-storm-connections-eu-en-1.50naruto-boruto-wiki
Naruto & Boruto Wiki Dataset
A structured text dataset containing all articles from the Naruto and Boruto Fandom wikis. Extracted via the MediaWiki API with clean, readable markdown-formatted text suitable for LLM pretraining, fine-tuning, or knowledge-augmented tasks.
Dataset Summary
This dataset contains 8,705 articles (7,934 from Naruto Wiki + 771 from Boruto Wiki) totaling 23.7 million characters (5.9M estimated tokens). Each article is stored as clean… See the full description on the dataset page: https://huggingface.co/datasets/TheOneWhoWill/naruto-boruto-wiki.gyoza-naruto-goal-synth-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101",
"total_episodes": 260,
"total_frames": 66924,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:260"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YUGOROU/gyoza-naruto-goal-synth-v2.haruno_sakura_naruto
Dataset of haruno_sakura (NARUTO)
This is the dataset of haruno_sakura (NARUTO), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
TTS_data_train2hyuuga_hinata_naruto
Dataset of hyuuga_hinata (NARUTO)
This is the dataset of hyuuga_hinata (NARUTO), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
naruto_trainnaruto-vision-prompts-dposhizune_naruto
Dataset of shizune (NARUTO)
This is the dataset of shizune (NARUTO), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
anauzumaki_kushina_naruto
Dataset of uzumaki_kushina (NARUTO)
This is the dataset of uzumaki_kushina (NARUTO), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
driver_licensenaruto-full-captionslayoutlmv3hyuuga_hanabi_naruto
Dataset of hyuuga_hanabi (NARUTO)
This is the dataset of hyuuga_hanabi (NARUTO), containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
