datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.Athar-EmbeddingsendoslamUltraX-Preview
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
📜 Paper |
💻 Code |
🤖 Models |
📦 UltraData Collection
English |
中文
📚 Introduction
UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/Kangarooz/UltraX-Preview.kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok
Bangumi Image Base of Kanchigai No Atelier Meister: Eiyuu Party No Moto Zatsuyougakari Ga, Jitsu Wa Sentou Igai Ga Sss Rank Datta To Iu Yoku Aru Hanashi
This is the image base of bangumi Kanchigai no Atelier Meister: Eiyuu Party no Moto Zatsuyougakari ga, Jitsu wa Sentou Igai ga SSS Rank Datta to Iu Yoku Aru Hanashi, we detected 55 characters, 5388 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok.vggfaceBotFails
BotFails: A Multimodal Dataset for Robotic Failure Detection
Overview
BotFails is a novel dataset specifically designed to support research on general failure detection in robotic manipulation. Addressing the scarcity of publicly available benchmarks in this domain, BotFails provides multimodal observations — including vision, proprioception, and natural language task instructions — collected across a semantically diverse set of manipulation scenarios.
Data collection was… See the full description on the dataset page: https://huggingface.co/datasets/kantine/BotFails.agi-structural-intelligence-protocols
AGI Structural Intelligence Protocols
Current positioning: SI-Core specifications, evaluation materials, implementation scaffolds, and historical LLM protocol experiments
Status note
The repository name reflects the project's early history. It is not a claim that AGI, machine consciousness, persistent selfhood, or permanent model transformation has been achieved.
This repository now contains two distinct generations of work:
Historical prompt-level experiments that explored… See the full description on the dataset page: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols.surg-vla-datasetPuffin-4M
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
📖 Project Page | 🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Dataset Details
Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence.
To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.Athar-Shamela4
Shamela 4 — Full Islamic Library Corpus
A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text.
Dataset Structure
stage0_raw/
├── _meta/ # Cross-cutting metadata (Parquet + JSONL)
│ ├── extraction_manifest.json # Global extraction record… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Shamela4.kanitakorn-th-sft
Kanitakorn — Thai-focused SFT corpus + tools (beats Typhoon-S-8B on ThaiExam / MATH / HotpotQA)
A Thai-language SFT dataset (4,147 records → 23,715 with Round 2 augmentation) and the training/eval
toolchain we used to fine-tune Qwen3-8B and Qwen3-4B-Instruct-2507 into Thai-benchmark-targeted models
that beat Typhoon-S-8B on multiple benchmarks.
Released models
8B variant: https://huggingface.co/Jnx03/kanitakorn-qwen3-8b-sft-v1
4B small-device variant:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-th-sft.japanese-listening-voicevox-backupDL3DV-Depth-DA3-Aligned
DL3DV-Depth-DA3-Aligned
Per-frame depth annotations for the DL3DV dataset, produced by
Depth-Anything-3 (DA3) and then aligned to each scene's sparse depth
from the original DL3DV reconstruction. We use this refined dataset for 3D world generation and reconstruction in our Puffin-World.
Sample Videos
Each clip is a 1×3 comparison — RGB | Original Depth | Our Aligned Depth —
with depth rendered by vision banana representation. It shows
how the DA3-aligned depth… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/DL3DV-Depth-DA3-Aligned.kangaroo_2025_5_6
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 5-6 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_5_6.kangaroo_2025_7_8
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 7-8 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_7_8.kangaroo_2025_9_10
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 9-10 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_9_10.kangaroo_2025_3_4
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 3-4 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_3_4.kangaroo_2025_1_2
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 1-2 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_1_2.kangaroo_2025_11_12
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 11-12 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_11_12.KANT_Domain_Corpus
KANT RAG Corpus Release
This Hugging Face dataset contains the retrieval materials used by the KANT RAG experiments. The main payload is split into 512 MiB parts so it can be uploaded and downloaded reliably.
The release includes:
rag_release_materials.tar.zst.part-*: split parts of the self-contained RAG materials archive.
SHA256SUMS.parts: checksums for the split archive parts.
SHA256SUMS.original_archives: checksums for the reconstructed archives.
Reconstruct… See the full description on the dataset page: https://huggingface.co/datasets/leo20000306/KANT_Domain_Corpus.knights-and-knaves
📘 knights-and-knaves Dataset [Project Page]
The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs.
🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning.
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl")
Available subset: test, train.
Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.flux_vgg50k_inv28_infer28_uncondIDTrueindian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.kanojookarishimasu3rdseason
Bangumi Image Base of Kanojo, Okarishimasu 3rd Season
This is the image base of bangumi Kanojo, Okarishimasu 3rd Season, we detected 70 characters, 8863 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanojookarishimasu3rdseason.kangaroo_math_mc_questionskanojookarishimasu
Bangumi Image Base of Kanojo, Okarishimasu
This is the image base of bangumi Kanojo, Okarishimasu, we detected 44 characters, 6680 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanojookarishimasu.kangaroo_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
competition (string): Competition or… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025.kantaicollectionkancolle
Bangumi Image Base of Kantai Collection: Kancolle
This is the image base of bangumi Kantai Collection: KanColle, we detected 54 characters, 3057 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kantaicollectionkancolle.
