CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging Face02lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips About Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset, more precisely the extension part not contained in EPIC-KITCHENS-55. For details, see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio. The clips folder contains one video for every narration from action annotations stored in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes7.5k downloads6mo agoHugging Face03kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.3k downloads1y agoHugging Face04kmfoda /booksum BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev Introduction The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.tabular10K<n<100K80 likes3.9k downloads4y agoHugging Face05Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K2 likes2.3k downloads5mo agoHugging Face06Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.3k downloads4mo agoHugging Face07kammartina /MA_Query_Expansion_MLT26tabular10K<n<100K0 likes2.2k downloads8h agoHugging Face08Ken4962 /processed_fake_job_postingstabulartext-classification10K<n<100K0 likes2.1k downloads1y agoHugging Face09KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face10stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads24d agoHugging Face11infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes1.7k downloads2y agoHugging Face12SII-KYW /CogStream CogStream Dataset Dataset for CogStream: Context-guided Streaming Video Question Answering. Overview CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams. Statistics: Split Videos QA Pairs Train 852 55,623 Test 236 15,364 Total 1,088 70,987 Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.tabularquestion-answering10K<n<100K1 likes1.6k downloads7mo agoHugging Face13kellycyy /CulturalBench CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs 📌 Resources: Paper | Leaderboard 📘 Description of CulturalBench CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru. We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.tabular1K<n<10K16 likes1.1k downloads2y agoHugging Face14kawsersikder /bangladesh-stock-market-dataset Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks Author: Kawser Sikder Overview A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market. Metric Value Total Stocks 441 Total Sectors 23 Total Trading Records 1,507,388 Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.tabulartime-series-forecasting1M<n<10M1 likes1.1k downloads1mo agoHugging Face15kshitijd /platonic-all-experimentstabularn<1K0 likes928 downloads1mo agoHugging Face16kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K9 likes900 downloads1y agoHugging Face17GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes878 downloads8mo agoHugging Face18SBMM75 /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes874 downloads12d agoHugging Face19Emmet-Allen /The-Bible-KJVtabular10K<n<100K0 likes730 downloads1y agoHugging Face20kshift /ahr999-dataset AHR999 BTC Hoarding Index Dataset Open, daily-updated AHR999 BTC hoarding index dataset, self-computed from Binance BTCUSDT daily closes and published as CSV and JSON. This Hugging Face repository is a mirror. The canonical dataset endpoints are: Dashboard: https://ahr999.aix4u.com/ GitHub: https://github.com/RuochenLyu/ahr999-dataset CSV endpoint: https://ahr999.aix4u.com/datasets/ahr999.csv JSON endpoint: https://ahr999.aix4u.com/datasets/ahr999.json Kaggle discovery mirror:… See the full description on the dataset page: https://huggingface.co/datasets/kshift/ahr999-dataset.tabular1K<n<10K1 likes713 downloads22h agoHugging Face21Koala-36M /Koala-36M-v1tabular10M<n<100M62 likes704 downloads2y agoHugging Face22kaysss /leetcode-problem-detailed LeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions, constraints, and examples. Columns: questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.tabulartext-classification1K<n<10K10 likes669 downloads1y agoHugging Face23kyselica /RoBo6 RoBo6: Standardized MMT Light Curve Dataset For Rocket Body Classification Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1]. The dataset was created to be used as a benchmark for rocket body light curve classification.For more informations follow the original paper: RoBo6: Standardized MMT Light Curve Dataset for Rocket Body Classification[^2] Class labels: ARIANE 5 R/B ATLAS 5… See the full description on the dataset page: https://huggingface.co/datasets/kyselica/RoBo6.tabular1K<n<10K0 likes649 downloads2y agoHugging Face24MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes639 downloads1y agoHugging Face25Leo-2025-Kai /hand_gesture_dataThe study is conducted on a total of 7 participants. The participants were instructed to perform three hand gestures (Hold, Single Tap and Double Tap) under different light conditions (low(100-200 lux, medium (600-750 lux) and high (1500-1600 lux)) and at different distances from the light sensor (low(2-4 cm) and high(8-10 cm)) tabular10K<n<100K0 likes586 downloads10mo agoHugging Face26DarwinDanish /kepler-objects-of-interesttabular1K<n<10K0 likes582 downloads10mo agoHugging Face27Kangverse /Sekai2_Real_World Sekai2 Real World This repository releases the reproducible URL/timestamp metadata and paired camera-pose/caption annotations for the perspective-video portion of Sekai2. See the paper: Sekai2: From World Exploration to Interactive World Modeling. Resources: 🌐 Project Page · 💻 GitHub · 📄 Paper The perspective MP4 clips are not redistributed here. Each row in sekai2_clips.csv provides the source URL and the exact half-open frame range [start_frame, end_frame) in a canonical 30… See the full description on the dataset page: https://huggingface.co/datasets/Kangverse/Sekai2_Real_World.tabulartext-to-video100K<n<1M0 likes574 downloads16d agoHugging Face28kennyhyder /sportsbookish-daily-odds SportsBookISH Daily Kalshi vs Sportsbook Odds Real-time pricing snapshot comparing Kalshi event-contract probabilities against US sportsbook consensus across nine sports. Description Daily-refreshed JSON / CSV export of every active Kalshi market alongside the de-vigged book median across 13+ US sportsbooks. Covers golf (PGA Tour), NFL, NBA, MLB, NHL, EPL, MLS, UEFA Champions League, and FIFA World Cup. Source Live data plane: JSON:… See the full description on the dataset page: https://huggingface.co/datasets/kennyhyder/sportsbookish-daily-odds.tabulartabular-regressionn<1K1 likes563 downloads21h agoHugging Face29kerne-protocol /honesty-index The Kerne Honesty Index What each synthetic dollar advertises, next to what it actually paid. Advertised APY versus realized APY for 21 synthetic dollar vaults, recomputed hourly from ERC-4626 share price growth on chain, and signed. The realized column is not taken from anybody's dashboard. It is measured directly from the vault contract: convertToAssets(10**decimals) read at two block heights, divided by 10**asset_decimals, annualized over the real elapsed time between those… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/honesty-index.imagetabular-regression10K<n<100K1 likes526 downloads21h agoHugging Face30kevykibbz /ecommerce-behavior-data-from-multi-category-store_oct-nov_2019 eCommerce Behavior Data from Multi-Category Store About the Dataset This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products. Dataset Overview Time Frame: October 2019 - April 2020 Total Events: 285 million Event Granularity: Each row represents an event associated with a product and a user. Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.tabular100M<n<1B4 likes519 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.