CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads4mo agoHugging Face02Kandil7 /Athar-Embeddingstabular1M<n<10M0 likes5.8k downloads5mo agoHugging Face03Kangarooz /UltraX-Preview UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing 📜 Paper | 💻 Code | 🤖 Models | 📦 UltraData Collection English | 中文 📚 Introduction UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/Kangarooz/UltraX-Preview.texttext-generation100M<n<1B0 likes5.4k downloads2mo agoHugging Face04KangLiao /Puffin-4M Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation    📖 Project Page  |    🖥️ GitHub    |   🤗 Hugging Face   |    📑 Paper    Dataset Details Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence. To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.imagetext-to-image1B<n<10B35 likes2.3k downloads9mo agoHugging Face05MathArena /kangaroo_2025_5_6 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 5-6 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_5_6.imagen<1K0 likes1.8k downloads4mo agoHugging Face06MathArena /kangaroo_2025_7_8 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 7-8 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_7_8.imagen<1K0 likes1.8k downloads4mo agoHugging Face07MathArena /kangaroo_2025_9_10 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 9-10 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_9_10.imagen<1K0 likes1.8k downloads4mo agoHugging Face08MathArena /kangaroo_2025_3_4 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 3-4 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_3_4.imagen<1K0 likes1.7k downloads4mo agoHugging Face09MathArena /kangaroo_2025_1_2 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 1-2 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_1_2.imagen<1K0 likes1.7k downloads4mo agoHugging Face10MathArena /kangaroo_2025_11_12 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 11-12 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_11_12.imagen<1K0 likes1.7k downloads4mo agoHugging Face11K-and-K /knights-and-knaves 📘 knights-and-knaves Dataset [Project Page] The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs. 🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl") Available subset: test, train. Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.textquestion-answering1K<n<10K38 likes1.4k downloads2y agoHugging Face12humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face13KanoonGPT /indian-case-laws Indian Case Laws Open Indian case-law data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in. Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.tabular10M<n<100M4 likes1.3k downloads5mo agoHugging Face14dfkiuser /kangaroo_math_mc_questionsimage1K<n<10K0 likes1.1k downloads8mo agoHugging Face15MathArena /kangaroo_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. competition (string): Competition or… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025.imagen<1K1 likes919 downloads4mo agoHugging Face16BangumiBase /kantaicollectionkancolle Bangumi Image Base of Kantai Collection: Kancolle This is the image base of bangumi Kantai Collection: KanColle, we detected 54 characters, 3057 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kantaicollectionkancolle.image1K<n<10K0 likes872 downloads2y agoHugging Face17DFKI-SLT /kangaroo_math_benchmark Känguruh Wettbewerb Dataset This dataset is a German benchmark from the official Känguru der Mathematik competition materials covering 1998--2025 (Link). Each instance is a single multiple-choice problem with five options (A - E) from school grade levels 3 - 13. The benchmark targets curriculum-aligned mathematical reasoning in German with explicit visual grounding (diagrams, geometric figures, spatial arrangements) while enabling controlled analyzes by year, grade group… See the full description on the dataset page: https://huggingface.co/datasets/DFKI-SLT/kangaroo_math_benchmark.imagevisual-question-answering1K<n<10K2 likes871 downloads3mo agoHugging Face18kantor3 /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT… See the full description on the dataset page: https://huggingface.co/datasets/kantor3/CADS-dataset.tabularimage-segmentation10K<n<100K0 likes762 downloads6mo agoHugging Face19Kangheng /refcocoimage10K<n<100K2 likes724 downloads2y agoHugging Face20Kangheng /PR1-Datasets-Groundingimage100K<n<1M9 likes621 downloads1y agoHugging Face21kanhatakeyama /wizardlm8x22b-logical-math-coding-sft_additional 自動生成したテキスト WizardLM 8x22bで生成した論理・数学・コード系のデータです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text100K<n<1M0 likes611 downloads2y agoHugging Face22BangumiBase /kanatanoastra Bangumi Image Base of Kanata No Astra This is the image base of bangumi Kanata no Astra, we detected 25 characters, 2286 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanatanoastra.image1K<n<10K0 likes608 downloads3y agoHugging Face23kanhatakeyama /wizardlm8x22b-logical-math-coding-sft 自動生成したテキスト WizardLM 8x22bで生成した論理・数学・コード系のデータです。 一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。 text100K<n<1M4 likes606 downloads2y agoHugging Face24Jiwon-Kang /Llama-Nemotron-VLM-Dataset-v1-OCR4image100K<n<1M1 likes584 downloads8mo agoHugging Face25Kangverse /Sekai2_Real_World Sekai2 Real World This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling. Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.tabulartext-to-video100K<n<1M0 likes574 downloads15d agoHugging Face26JKA-NLP /unified-kannada-asr-1.0 Dataset Card for "unified-kannada-asr-1.0" More Information needed audio100K<n<1M1 likes573 downloads3y agoHugging Face27Jiwon-Kang /pixmo-point-count-concat_0-20image100K<n<1M0 likes565 downloads9mo agoHugging Face28Jiwon-Kang /sd2.1_cocotext10K<n<100K0 likes533 downloads2y agoHugging Face29imvladikon /hebrew_speech_kan Dataset Card for Dataset Name Dataset Summary Hebrew Dataset for ASR Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances {'audio': {'path': '/root/.cache/huggingface/datasets/downloads/extracted/8ce7402f6482c6053251d7f3000eec88668c994beb48b7ca7352e77ef810a0b6/train/e429593fede945c185897e378a5839f4198.wav', 'array': array([-0.00265503, -0.0018158… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/hebrew_speech_kan.audioautomatic-speech-recognition10K<n<100K14 likes518 downloads3y agoHugging Face30kangaroo-dataset-german /kangaroo_dataset German Kangaroo Benchmark The complete German Mathematical Kangaroo archive from 1998 to 2025 as a multiple-choice benchmark: 3,886 items from 140 exams in five grade groups (3--4, 5--6, 7--8, 9--10, 11--13), worth 3, 4, or 5 points each. 1,746 items are multimodal, with a question diagram, image-based answer options, or both. The accompanying paper describes the extraction, the evaluation protocol, and the results. Files kangaroo.parquet: the benchmark, 3,886… See the full description on the dataset page: https://huggingface.co/datasets/kangaroo-dataset-german/kangaroo_dataset.textquestion-answering1K<n<10K0 likes447 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.