datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
naruto-boruto-wiki
Naruto & Boruto Wiki Dataset
A structured text dataset containing all articles from the Naruto and Boruto Fandom wikis. Extracted via the MediaWiki API with clean, readable markdown-formatted text suitable for LLM pretraining, fine-tuning, or knowledge-augmented tasks.
Dataset Summary
This dataset contains 8,705 articles (7,934 from Naruto Wiki + 771 from Boruto Wiki) totaling 23.7 million characters (5.9M estimated tokens). Each article is stored as clean… See the full description on the dataset page: https://huggingface.co/datasets/TheOneWhoWill/naruto-boruto-wiki.naruto-instruction-promptsNaruto_training
