CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads4mo agoHugging Face02Kandil7 /Athar-Embeddingstabular1M<n<10M0 likes5.8k downloads5mo agoHugging Face03kangqi-ni /endoslamimage100K<n<1M0 likes5.5k downloads7mo agoHugging Face04Kangarooz /UltraX-Preview UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing 📜 Paper | 💻 Code | 🤖 Models | 📦 UltraData Collection English | 中文 📚 Introduction UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/Kangarooz/UltraX-Preview.texttext-generation100M<n<1B0 likes5.4k downloads2mo agoHugging Face05BangumiBase /kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok Bangumi Image Base of Kanchigai No Atelier Meister: Eiyuu Party No Moto Zatsuyougakari Ga, Jitsu Wa Sentou Igai Ga Sss Rank Datta To Iu Yoku Aru Hanashi This is the image base of bangumi Kanchigai no Atelier Meister: Eiyuu Party no Moto Zatsuyougakari ga, Jitsu wa Sentou Igai ga SSS Rank Datta to Iu Yoku Aru Hanashi, we detected 55 characters, 5388 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanchigainoateliermeistereiyuupartynomotozatsuyougakarigajitsuwasentouigaigasssrankdattatoiuyok.image1K<n<10K0 likes4.6k downloads1y agoHugging Face06Jiwon-Kang /vggfaceimage10K<n<100K0 likes3.2k downloads10mo agoHugging Face07kantine /BotFails BotFails: A Multimodal Dataset for Robotic Failure Detection Overview BotFails is a novel dataset specifically designed to support research on general failure detection in robotic manipulation. Addressing the scarcity of publicly available benchmarks in this domain, BotFails provides multimodal observations — including vision, proprioception, and natural language task instructions — collected across a semantically diverse set of manipulation scenarios. Data collection was… See the full description on the dataset page: https://huggingface.co/datasets/kantine/BotFails.videoroboticsn<1K2 likes2.7k downloads5mo agoHugging Face08kanaria007 /agi-structural-intelligence-protocols AGI Structural Intelligence Protocols Current positioning: SI-Core specifications, evaluation materials, implementation scaffolds, and historical LLM protocol experiments Status note The repository name reflects the project's early history. It is not a claim that AGI, machine consciousness, persistent selfhood, or permanent model transformation has been achieved. This repository now contains two distinct generations of work: Historical prompt-level experiments that explored… See the full description on the dataset page: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols.14 likes2.6k downloads1d agoHugging Face09Kanden1112 /surg-vla-datasetimage100K<n<1M0 likes2.4k downloads3mo agoHugging Face10KangLiao /Puffin-4M Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation    📖 Project Page  |    🖥️ GitHub    |   🤗 Hugging Face   |    📑 Paper    Dataset Details Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence. To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.imagetext-to-image1B<n<10B35 likes2.3k downloads9mo agoHugging Face11Kandil7 /Athar-Shamela4 Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Shamela4.text-generation10M<n<100M1 likes2.1k downloads4mo agoHugging Face12Jnx03 /kanitakorn-th-sft Kanitakorn — Thai-focused SFT corpus + tools (beats Typhoon-S-8B on ThaiExam / MATH / HotpotQA) A Thai-language SFT dataset (4,147 records → 23,715 with Round 2 augmentation) and the training/eval toolchain we used to fine-tune Qwen3-8B and Qwen3-4B-Instruct-2507 into Thai-benchmark-targeted models that beat Typhoon-S-8B on multiple benchmarks. Released models 8B variant: https://huggingface.co/Jnx03/kanitakorn-qwen3-8b-sft-v1 4B small-device variant:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-th-sft.text-generation1K<n<10K0 likes2.1k downloads4mo agoHugging Face13kanuli1983 /japanese-listening-voicevox-backupaudio0 likes2k downloads18d agoHugging Face14KangLiao /DL3DV-Depth-DA3-Aligned DL3DV-Depth-DA3-Aligned Per-frame depth annotations for the DL3DV dataset, produced by Depth-Anything-3 (DA3) and then aligned to each scene's sparse depth from the original DL3DV reconstruction. We use this refined dataset for 3D world generation and reconstruction in our Puffin-World. Sample Videos Each clip is a 1×3 comparison — RGB | Original Depth | Our Aligned Depth — with depth rendered by vision banana representation. It shows how the DA3-aligned depth… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/DL3DV-Depth-DA3-Aligned.depth-estimation1M<n<10M0 likes2k downloads3mo agoHugging Face15MathArena /kangaroo_2025_5_6 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 5-6 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_5_6.imagen<1K0 likes1.8k downloads4mo agoHugging Face16MathArena /kangaroo_2025_7_8 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 7-8 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_7_8.imagen<1K0 likes1.8k downloads4mo agoHugging Face17MathArena /kangaroo_2025_9_10 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 9-10 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_9_10.imagen<1K0 likes1.8k downloads4mo agoHugging Face18MathArena /kangaroo_2025_3_4 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 3-4 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_3_4.imagen<1K0 likes1.7k downloads4mo agoHugging Face19MathArena /kangaroo_2025_1_2 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 1-2 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_1_2.imagen<1K0 likes1.7k downloads4mo agoHugging Face20MathArena /kangaroo_2025_11_12 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 11-12 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. Source Data The original… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025_11_12.imagen<1K0 likes1.7k downloads4mo agoHugging Face21leo20000306 /KANT_Domain_Corpus KANT RAG Corpus Release This Hugging Face dataset contains the retrieval materials used by the KANT RAG experiments. The main payload is split into 512 MiB parts so it can be uploaded and downloaded reliably. The release includes: rag_release_materials.tar.zst.part-*: split parts of the self-contained RAG materials archive. SHA256SUMS.parts: checksums for the split archive parts. SHA256SUMS.original_archives: checksums for the reconstructed archives. Reconstruct… See the full description on the dataset page: https://huggingface.co/datasets/leo20000306/KANT_Domain_Corpus.0 likes1.4k downloads3mo agoHugging Face22K-and-K /knights-and-knaves 📘 knights-and-knaves Dataset [Project Page] The knights-and-knaves dataset serves as a logical reasoning benchmark to evaluate the reasoning capabilities of LLMs. 🚀🚀 Check out the perturbed knights-and-knaves dataset to evaluate the memorization of LLMs in reasoning. Loading the dataset To load the dataset: from datasets import load_dataset data_subject = load_dataset('K-and-K/knights-and-knaves','test',split="2ppl") Available subset: test, train. Available… See the full description on the dataset page: https://huggingface.co/datasets/K-and-K/knights-and-knaves.textquestion-answering1K<n<10K38 likes1.4k downloads2y agoHugging Face23humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face24Jiwon-Kang /flux_vgg50k_inv28_infer28_uncondIDTrueimage10K<n<100K0 likes1.3k downloads10mo agoHugging Face25KanoonGPT /indian-case-laws Indian Case Laws Open Indian case-law data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative - an effort to make Indian legal data easier to access, trace, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal data and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in. Repository: KanoonGPT/indian-case-laws… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-case-laws.tabular10M<n<100M4 likes1.3k downloads5mo agoHugging Face26BangumiBase /kanojookarishimasu3rdseason Bangumi Image Base of Kanojo, Okarishimasu 3rd Season This is the image base of bangumi Kanojo, Okarishimasu 3rd Season, we detected 70 characters, 8863 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanojookarishimasu3rdseason.image1K<n<10K0 likes1.2k downloads2y agoHugging Face27dfkiuser /kangaroo_math_mc_questionsimage1K<n<10K0 likes1.1k downloads8mo agoHugging Face28BangumiBase /kanojookarishimasu Bangumi Image Base of Kanojo, Okarishimasu This is the image base of bangumi Kanojo, Okarishimasu, we detected 44 characters, 6680 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kanojookarishimasu.image1K<n<10K0 likes1.1k downloads3y agoHugging Face29MathArena /kangaroo_2025 Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from Kangaroo 2025 used for the MathArena Leaderboard. Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. image (image): Problem image. competition (string): Competition or… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/kangaroo_2025.imagen<1K1 likes919 downloads4mo agoHugging Face30BangumiBase /kantaicollectionkancolle Bangumi Image Base of Kantai Collection: Kancolle This is the image base of bangumi Kantai Collection: KanColle, we detected 54 characters, 3057 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/kantaicollectionkancolle.image1K<n<10K0 likes872 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.