datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ALE-Bench
ALE-Bench
Dataset Description
ALE-Bench is a benchmark for evaluating AI systems on score-based algorithmic programming contests.
This dataset is officially provided by AtCoder Inc..
Please be sure to check the "License" section below.
Please read our blog post and our paper for more details.
Related resources:
Preprint paper (arXiv)
Sakana AI Blog (English)
Sakana AI Blog (Japanese)
GitHub repository
Leaderboard
Usage
Our Python library automatically… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/ALE-Bench.sakugabooru2025
Sakugabooru2025: Curated Animation Clips from Enthusiasts
Sakugabooru.com is a booru-style imageboard dedicated to collecting and sharing noteworthy animation clips, emphasizing Japanese anime but open to creators worldwide. Over the years, it has amassed more than 240,000 animation clips, alongside informative blog posts for anime fans everywhere.
With the growing interest in generative video models and AI animations, the scarcity of proper animation-related video datasets has… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/sakugabooru2025.sakuratrick
Bangumi Image Base of Sakura Trick
This is the image base of bangumi Sakura Trick, we detected 17 characters, 1556 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sakuratrick.SAKE-Twitteraloha_real_agilex_folding_fabricaloha_real_agilex_nesting_dollsakurasounopetnakanojo
Bangumi Image Base of Sakurasou No Pet Na Kanojo
This is the image base of bangumi Sakurasou no Pet na Kanojo, we detected 24 characters, 4107 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/sakurasounopetnakanojo.JA-VG-VQA-500
JA-VG-VQA-500
Dataset Description
JA-VG-VQA-500 is a 500-sample subset of Japanese Visual Genome VQA dataset.
This dataset was used in the evaluation of EvoVLM-JP-v1-7B.
Please refer to our report and blog for more details.
We are grateful to the developers for making the dataset available under Creative Commons Attribution 4.0 License.
Visual Genome
Japanese Visual Genome VQA dataset
Usage
Use the code below to get started with the dataset.
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/JA-VG-VQA-500.aloha_real_agilex_clean_cup_drawer_and_clean_cube_container-20250809aloha_real_agilex_clean_to_drawer-20250805reuse-vs-recraft-substrate
reuse-vs-recraft — substrate package
Everything the bucket-repair arc authors or derives itself: training preference
pairs, the validation set, held-out and drift rows, the conflict-scene image
substrate, the frozen probe set and its items, the repair-slice bundle, and
the PhD panel id lists.
Paired with reuse-vs-recraft-battery-bundle, which carries cached copies of
public benchmark items. Together they mean a fresh evaluation machine downloads
no benchmark dataset at all:… See the full description on the dataset page: https://huggingface.co/datasets/saketh-chervu/reuse-vs-recraft-substrate.sakimi-chan_LoRA
Sakimi-chan LoRA
Who is Sakimi-chan?
Sakimi-chan is a Canadian artist best known for her digital paintings and unique style. She mainly draws fanart of games and popular characters and creates gifs for her fans with voiceovers.
Patreon: https://www.patreon.com/sakimichan
Use Cases
The LoRA is in itself very compatible with the most diverse model. However, it is most effective when used with Kenshi or AbyssOrangeMix2.
The LoRA itself was trained with the token:… See the full description on the dataset page: https://huggingface.co/datasets/Nerfgun3/sakimi-chan_LoRA.aloha_real_agilex_clean_cup_drawer_and_clean_cube_container-20250813agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
reuse-vs-recraft-slice150
The 150 conflict slices — the reuse-vs-recraft repair/damage test set
The standing conflict-slice metric of the reuse-vs-recraft programme
(program-wide fix, 2026-08-16). 150 items: for each of three benchmarks,
25 repair-side items the untrained model (Qwen2.5-VL-7B-Instruct,
greedy, 48 new tokens) answers WRONGLY — it follows the misleading text —
and 25 cost-side items it answers CORRECTLY — it resists. A trained
model is scored on both sides:
fixed — repair-side items it… See the full description on the dataset page: https://huggingface.co/datasets/saketh-chervu/reuse-vs-recraft-slice150.aloha_real_agilex_clean_to_drawer_and_clean_cube_container-20250807JA-VLM-Bench-In-the-Wild
JA-VLM-Bench-In-the-Wild
Dataset Description
JA-VLM-Bench-In-the-Wild is Japanese version of LLaVA-Bench-In-the-Wild.
We carefully collected a diverse set of 42 images with 50 questions in total. (For LLaVA-Bench-In-the-Wild, 24 images with 60 questions)
The images contain Japanese culture and objects in Japan. The Japanese questions and answers were generated with assistance from GPT-4V (gpt-4-vision-preview), OpenAI’s large-scale language-generation model and removed… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/JA-VLM-Bench-In-the-Wild.signalsakha-ocr-synth
Синтетические строки якутского текста для OCR
500 000 изображений строк с точной разметкой. Сделано для обучения
распознавателя якутского (саха) текста: готовые движки для этого языка не
работают, а размеченных строк почти нет.
Зачем вообще синтетика: у tesseract-rus и ABBYY FineReader 10 на якутской
печати доля правильно прочитанных специфических букв ҕ ҥ ө һ ү равна нулю.
Не «низкая» — ноль на 460 тысячах букв, при том что эти буквы составляют около
7% всех букв и встречаются… See the full description on the dataset page: https://huggingface.co/datasets/lab-ii/sakha-ocr-synth.mazes-largeJA-Multi-Image-VQA
JA-Multi-Image-VQA
Dataset Description
JA-Multi-Image-VQA is a dataset for evaluating the question answering capabilities on multiple image inputs.
We carefully collected a diverse set of 39 images with 55 questions in total.
Some images contain Japanese culture and objects in Japan. The Japanese questions and answers were created manually.
Usage
from datasets import load_dataset
dataset = load_dataset("SakanaAI/JA-Multi-Image-VQA", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/JA-Multi-Image-VQA.alaska_dead_trees
Alaskan Dead Tree Dataset
Dataset Details
This dataset consists of high resolution aerial images of forest from the Kenai peninsula area of Alaska taken by NASA Goddard's G-LiHT instrument.
Pixel-wise annotations were created using k-means clustering augmented by a Gray Level Co-occurrence Matrix (GLCM). A value of 1 denotes a dead tree, and 0 denotes everything else. The masks and images were tiled into 256x256 pixel squares and paired together, then any empty masks… See the full description on the dataset page: https://huggingface.co/datasets/saking3/alaska_dead_trees.diy-project-code-based-on-hardware-imageLIBERO-SOG10-LeRobotambiguous-imagesMetaThinker-dataKamonBench
KamonBench
A grammar-based image-to-structure benchmark for evaluating compositional
factor recovery in vision-language models, built around Japanese family crests
(kamon, 家紋).
Each composite crest is paired with:
a formal kamon description language string (KDL, kamon yōgo, 家紋用語),
a segmented Japanese analysis,
an English translation,
a non-linguistic program code over the generator factors.
Because every crest is synthesized from a known triple of generator factors
(container C… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/KamonBench.sakshi-handwritingisac-simo-mortar-seg-demo
isac-simo-mortar-seg-demo
Tiny synthetic masonry set with pixel-exact mortar masks (255 = mortar, 0 = brick), for the wall segmentation task. Masks are exact by construction since the walls are procedurally drawn. Five train / three validation. Teaching use only.
SAKS_JEWELRY
