CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01klieret /swe-bench-dummy-test-datasettextn<1K0 likes75k downloads1y agoHugging Face02hatemestinbejaia /ExperimentDATA_knowledge_distillation_vs_fine_tuningtabular100M<n<1B1 likes65k downloads9mo agoHugging Face03KROX777 /PDExplBenchtext1K<n<10K0 likes54k downloads5mo agoHugging Face04kohsei /MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌 CVPR 2026 (Main) This repository provides the datasets for “MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta Paper Link https://arxiv.org/abs/2511.22989 Github Repository For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.imagetext-to-image1K<n<10K5 likes21k downloads3mo agoHugging Face05ojttykkjn /kkhkkkntextn<1K3 likes20k downloads8h agoHugging Face06chcaa /kb-books open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish Library published between 1750 and 1930. The dataset has each page of each document in image and text format. The text was extracted with OCR. The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.image1M<n<10M4 likes20k downloads10mo agoHugging Face07KokosDev /tahoe-100m-zarr Tahoe-100M Zarr Collection Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub. Why Zarr Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.textn<1K1 likes18k downloads6mo agoHugging Face08IFM /K2Datasets K2 Dataset Card The following data mix was used to train K2 and achieve results in line with Llama 2 70B. Dataset Details K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below. Dataset Description: Stage 1 Dataset Starting Tokens Multiplier Total Tokens % of Total dm-math 4.33B 3x 13B 1% pubmed-abstracts (from the Pile) 4.77B 3x 14.3B 1.1% uspto (from the Pile) 4.77B 3x… See the full description on the dataset page: https://huggingface.co/datasets/IFM/K2Datasets.text100M<n<1B20 likes18k downloads2y agoHugging Face09JDRJ /kjv-bibletext10K<n<100K0 likes17k downloads2y agoHugging Face10AlphaDojo /dojo_stock_kline Languages: 简体中文 · English dojo_stock_kline — Stock Daily Bars Overview Daily OHLCV history for US, CN, and HK equities, including cumulative adjustment factors and dividend/split flags. Files File Description data.parquet Market-wide daily bars Key Fields Field Description symbol Join key kline_t Bar interval; snapshots use "1D" for daily bars bar_time Bar timestamp (trade date) open / high / low /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_kline.tabular1M<n<10M0 likes17k downloads4h agoHugging Face11Kaphathy /Dataset MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records 1. Executive Summary & Repository Overview The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.textimage-classificationn<1K2 likes17k downloads1d agoHugging Face12argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes16k downloads3y agoHugging Face13KAS2003 /xfield-radar-dataset-20260915 XField radar dataset — formal snapshot, 2026-09-15 Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json. This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.text10K<n<100K0 likes16k downloads9d agoHugging Face14inclusionAI /ASearcher-Local-Knowledgetext10M<n<100M7 likes15k downloads1y agoHugging Face15KodCode /KodCode-V1-SFT-R1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.tabularquestion-answering100K<n<1M40 likes15k downloads2y agoHugging Face16AlphaDojo /dojo_benchmark_kline Languages: 简体中文 · English dojo_benchmark_kline — Benchmark Index Bars Overview Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date. Files File Description data.parquet Index daily bars Key Fields Field Description symbol Index code (e.g. ^SPX, 000300.SS) kline_t Bar interval; "1D" for daily bars bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.text1K<n<10K1 likes15k downloads4h agoHugging Face17AlphaDojo /dojo_forex_kline Languages: 简体中文 · English dojo_forex_kline — FX Daily Bars Overview Daily OHLC and amplitude for major currency pairs. Used to convert revenue, profit, and other filing amounts into a single currency when report currency and listing/analysis currency differ. Typical case: regional revenue in HKD in dojo_main_income while analysis targets USD — apply HKDUSD (or equivalent) at the report date. Intended Use: Cross-Currency Revenue Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_forex_kline.tabular1K<n<10K2 likes14k downloads4h agoHugging Face18ksolovev /FineNewsTestSampletext10M<n<100M0 likes14k downloads7mo agoHugging Face19KbsdJames /Omni-MATH Dataset Card for Omni-MATH Recent advancements in AI, particularly in large language models (LLMs), have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To mitigate this limitation, we propose a comprehensive and challenging benchmark specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/KbsdJames/Omni-MATH.text1K<n<10K132 likes12k downloads2y agoHugging Face20KodCode /KodCode-V1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1.tabular100K<n<1M117 likes12k downloads2y agoHugging Face21kapturecx /bolAIndiagated bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Sources One config per transcription system, so their output stays separable. config (source_id) provider model hours rows shards vendor-a vendor-a undisclosed 420.03 480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.audioautomatic-speech-recognition10M<n<100M1 likes12k downloads16m agoHugging Face22karpathy /fineweb-edu-100b-shuffletext10M<n<100M171 likes11k downloads1y agoHugging Face23KarlQuant /quasar-axrvi-v10tabularn<1K15 likes10k downloads19d agoHugging Face24KantaHayashiAI /ClimbLab-JaJapanese / 日本語版 ClimbLab-Ja ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.tabulartext-generation100M<n<1B2 likes10k downloads8h agoHugging Face25KRAFTON /Raon-OpenTTS-Pool Raon-OpenTTS-Pool Technical Report Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training, constructed from 8 publicly available speech corpora and a set of web-sourced recordings. It is the training data behind Raon-OpenTTS, an open TTS model that performs on par with state-of-the-art closed-data systems. 615K hours of speech audio 239.7M speech segments 11 source datasets aggregated into a unified format All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.texttext-to-speech100M<n<1B43 likes9.8k downloads4mo agoHugging Face26SimulaMet /Kvasir-VQA-x1 Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub | Original Image from Kvasir-VQA(Simula Datasets) | Paper 🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate! Overview Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.imagevisual-question-answering100K<n<1M16 likes9.8k downloads1y agoHugging Face27Damaru-ai /damru-knowledge 🐕 Damru Knowledge A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students. The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/. 📦 What's inside Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.textquestion-answering10M<n<100M4 likes9.7k downloads10m agoHugging Face28kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes9.6k downloads3y agoHugging Face29keenable-ai /needle-resultstext10K<n<100K3 likes9.4k downloads41m agoHugging Face30HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.