datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.Athar-EmbeddingsUltraX-Preview
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
📜 Paper |
💻 Code |
🤖 Models |
📦 UltraData Collection
English |
中文
📚 Introduction
UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/Kangarooz/UltraX-Preview.Puffin-4M
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
📖 Project Page | 🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Dataset Details
Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence.
To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.kangaroo_2025_5_6
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 5-6 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_5_6.kangaroo_2025_7_8
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 7-8 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_7_8.kangaroo_2025_9_10
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 9-10 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_9_10.kangaroo_2025_3_4
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 3-4 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_3_4.kangaroo_2025_1_2
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 1-2 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_1_2.kangaroo_2025_11_12
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 11-12 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
Source Data
The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_11_12.knights-and-knaves
📘 knights-and-knaves Dataset [Project Page]
The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs.
🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning.
Loading the dataset
To load the dataset:
from datasets import load_dataset
data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl")
Available subset: test, train.
Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated
Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,217
Total Duration: 42.77 hours
Average Duration: 5.87 seconds
Duration Range: 0.65s - 122.23s
Average Phonemes: 18.5 per sample
Average Kanade Tokens: 151.1 per sample
Global Embedding Dimension: 128
New Columns
This dataset adds the following columns:
duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.indian-case-laws
Indian Case Laws
Open Indian case-law data for AI, search, and legal research.
This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems.
KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.
Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.kangaroo_math_mc_questionskangaroo_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Kangaroo 2025 used for the MathArena Leaderboard.
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
image (image): Problem image.
competition (string): Competition or… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025.kantaicollectionkancolle
Bangumi Image Base of Kantai Collection: Kancolle
This is the image base of bangumi Kantai Collection: KanColle, we detected 54 characters, 3057 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kantaicollectionkancolle.kangaroo_math_benchmark
Känguruh Wettbewerb Dataset
This dataset is a German benchmark from the official Känguru der Mathematik competition materials covering 1998--2025 (Link).
Each instance is a single multiple-choice problem with five options (A - E) from school grade levels 3 - 13. The benchmark targets curriculum-aligned mathematical reasoning in German with explicit visual grounding (diagrams, geometric figures, spatial arrangements) while enabling controlled analyzes by year, grade group… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/kangaroo_math_benchmark.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT… See the full description on the dataset page: https://huggingface.co/datasets/kantor3/CADS-dataset.refcocoPR1-Datasets-Groundingwizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
kanatanoastra
Bangumi Image Base of Kanata No Astra
This is the image base of bangumi Kanata no Astra, we detected 25 characters, 2286 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanatanoastra.wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
Llama-Nemotron-VLM-Dataset-v1-OCR4Sekai2_Real_World
Sekai2 Real World
This repository releases the reproducible URL/timestamp metadata and paired
camera-pose/caption annotations for the perspective-video portion of
Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling.
Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper
The perspective MP4 clips are not redistributed here. Each row in
sekai2_clips.csv provides the source URL and the exact half-open frame range
[start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.unified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
pixmo-point-count-concat_0-20sd2.1_cocohebrew_speech_kan
Dataset Card for Dataset Name
Dataset Summary
Hebrew Dataset for ASR
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
{'audio': {'path': '/root/.cache/huggingface/datasets/downloads/extracted/8ce7402f6482c6053251d7f3000eec88668c994beb48b7ca7352e77ef810a0b6/train/e429593fede945c185897e378a5839f4198.wav',
'array': array([-0.00265503, -0.0018158… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_kan.kangaroo_dataset
German Kangaroo Benchmark
The complete German Mathematical Kangaroo archive from 1998 to 2025 as a
multiple-choice benchmark: 3,886 items from 140 exams in five grade groups
(3--4, 5--6, 7--8, 9--10, 11--13), worth 3, 4, or 5 points each. 1,746 items are
multimodal, with a question diagram, image-based answer options, or both. The
accompanying paper describes the extraction, the evaluation protocol, and the
results.
Files
kangaroo.parquet: the benchmark, 3,886… See the full description on the dataset page: https://huggingface.co/datasets/kangaroo-dataset-german/kangaroo_dataset.
