CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HAERAE-HUB /KMMLU KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.tabularmultiple-choice100K<n<1M101 likes8.9k downloads3y agoHugging Face02lightly-ai /epic-kitchens-100-clips EPIC-KITCHENS-100 Extracted Clips About Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset, more precisely the extension part not contained in EPIC-KITCHENS-55. For details, see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio. The clips folder contains one video for every narration from action annotations stored in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.tabular10K<n<100K2 likes7.5k downloads6mo agoHugging Face03keivalya /MedQuad-MedicalQnADataset Reference: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019. textquestion-answering10K<n<100K133 likes7.2k downloads3y agoHugging Face04knkarthick /dialogsum Dataset Card for DIALOGSum Corpus Dataset Description Links Homepage: https://aclanthology.org/2021.findings-acl.449 Repository: https://github.com/cylnlp/dialogsum Paper: https://aclanthology.org/2021.findings-acl.449 Point of Contact: https://huggingface.co/knkarthick Dataset Summary DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/dialogsum.textsummarization10K<n<100K249 likes7.2k downloads3y agoHugging Face05knkarthick /samsum Dataset Card for SAMSum Corpus Dataset Description Links Homepage: hhttps://arxiv.org/abs/1911.12237v2 Repository: https://arxiv.org/abs/1911.12237v2 Paper: https://arxiv.org/abs/1911.12237v2 Point of Contact: https://huggingface.co/knkarthick Dataset Summary The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.textsummarization10K<n<100K44 likes5.9k downloads1y agoHugging Face06kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.3k downloads1y agoHugging Face07kmfoda /booksum BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev Introduction The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.tabular10K<n<100K80 likes3.9k downloads4y agoHugging Face08HAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3k downloads3y agoHugging Face09Infatoshi /kernelbench-v3-runs KernelBench-v3 — Agent Runs 2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py. Companion datasets: Infatoshi/kernelbench-v3-problems — 60 problem definitions Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.tabular1K<n<10K2 likes2.3k downloads5mo agoHugging Face10Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.3k downloads4mo agoHugging Face11kammartina /MA_Query_Expansion_MLT26tabular10K<n<100K0 likes2.2k downloads6h agoHugging Face12Ken4962 /processed_fake_job_postingstabulartext-classification10K<n<100K0 likes2.1k downloads1y agoHugging Face13Koushul /spacetravlr SpaceTravLR dataset hub Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports. Layout spacetravlr/ ├── tonsil/ # placeholder / demo gene outputs └── xenium_skin_mixed/ ├── run.toml # shared training config for this cohort ├── manifest.json # sample index and upload metadata ├── sample12/ ├── sample13/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.text10K<n<100K0 likes2.1k downloads2mo agoHugging Face14KlingTeam /FullBenchimage1K<n<10K7 likes2k downloads2y agoHugging Face15KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face16stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads24d agoHugging Face17infinite-dataset-hub /FINDER_API_KEY_AI_SEARCH_2023 FINDER_API_KEY_AI_SEARCH_2023 tags: data collection, machine learning, API performance Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.tabularn<1K0 likes1.7k downloads2y agoHugging Face18SII-KYW /CogStream CogStream Dataset Dataset for CogStream: Context-guided Streaming Video Question Answering. Overview CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams. Statistics: Split Videos QA Pairs Train 852 55,623 Test 236 15,364 Total 1,088 70,987 Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.tabularquestion-answering10K<n<100K1 likes1.6k downloads7mo agoHugging Face19pseudolab /huggingface-krew-hackathon2023textn<1K2 likes1.5k downloads3y agoHugging Face20shangrilar /ko_text2sqltext10K<n<100K19 likes1.5k downloads3y agoHugging Face21imageomics /KABR Dataset Card for KABR: In-Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos Dataset Summary We present a novel high-quality dataset for animal behavior recognition from drone videos. The dataset is focused on Kenyan wildlife and contains behaviors of giraffes, plains zebras, and Grevy's zebras. The dataset consists of more than 10 hours of annotated videos, and it includes eight different classes, encompassing seven types of animal behavior and an… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR.textvideo-classification1M<n<10M9 likes1.4k downloads9mo agoHugging Face22Kevin-Pal /CUHK-X_Small_Model_Trackgated CUHK-X — Small Model Track Multimodal human action recognition (classification). Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes). Repository layout . ├── Training/ │ ├── class_mapping.csv # action_id <-> action_name (40 classes) │ └── data/ │ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip │ → HAR/data/<modality>/<action>/<user>/<trial>/<files> └── Testing/ ├── data/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.textvideo-classificationn<1K7 likes1.3k downloads3mo agoHugging Face23Kaludi /Customer-Support-Responsestextn<1K13 likes1.2k downloads3y agoHugging Face24kwhuang /C3 C3: Cross-View Cross-Modality Correspondence Dataset Dataset for C3Po: Cross-View Cross-Modality Correspondence with Pointmap Prediction arXiv | Project Website | GitHub C3 contains 90K paired floor plans and photos from the Internet across 597 scenes with 153M pixel-level correspondences and 85K camera poses. Image Pairs image_pairs/ is split into train/, val, and test/, each with a image_pairs.csv. image_pairs.csv: Each row represents a plan-photo pair… See the full description on the dataset page: https://huggingface.co/datasets/kwhuang/C3.text10K<n<100K5 likes1.2k downloads9mo agoHugging Face25kellycyy /CulturalBench CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs 📌 Resources: Paper | Leaderboard 📘 Description of CulturalBench CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru. We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.tabular1K<n<10K16 likes1.1k downloads2y agoHugging Face26kawsersikder /bangladesh-stock-market-dataset Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks Author: Kawser Sikder Overview A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market. Metric Value Total Stocks 441 Total Sectors 23 Total Trading Records 1,507,388 Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.tabulartime-series-forecasting1M<n<10M1 likes1.1k downloads1mo agoHugging Face27Kushal0532 /news-political-bias-classification-datasetDataset actually from kaggle. Couldn't find it here so I uploaded it. text10K<n<100K0 likes937 downloads11mo agoHugging Face28kshitijd /platonic-all-experimentstabularn<1K0 likes928 downloads1mo agoHugging Face29kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K9 likes900 downloads1y agoHugging Face30GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes878 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.