datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMLU
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.MedQuad-MedicalQnADataset
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
dialogsum
Dataset Card for DIALOGSum Corpus
Dataset Description
Links
Homepage: https://aclanthology.org/2021.findings-acl.449
Repository: https://github.com/cylnlp/dialogsum
Paper: https://aclanthology.org/2021.findings-acl.449
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/dialogsum.samsum
Dataset Card for SAMSum Corpus
Dataset Description
Links
Homepage: hhttps://arxiv.org/abs/1911.12237v2
Repository: https://arxiv.org/abs/1911.12237v2
Paper: https://arxiv.org/abs/1911.12237v2
Point of Contact: https://huggingface.co/knkarthick
Dataset Summary
The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.booksum
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization
Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev
Introduction
The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases.
While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.kernelbench-v3-runs
KernelBench-v3 — Agent Runs
2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py.
Companion datasets:
Infatoshi/kernelbench-v3-problems — 60 problem definitions
Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.suno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.MA_Query_Expansion_MLT26processed_fake_job_postingsspacetravlr
SpaceTravLR dataset hub
Precomputed SpaceTravLR outputs: per-gene beta matrices (*_betadata.feather), run metadata, and optional per-sample .h5ad exports.
Layout
spacetravlr/
├── tonsil/ # placeholder / demo gene outputs
└── xenium_skin_mixed/
├── run.toml # shared training config for this cohort
├── manifest.json # sample index and upload metadata
├── sample12/
├── sample13/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Koushul/spacetravlr.FullBenchRaon-OpenTTS-Eval
Raon-OpenTTS-Eval
Technical Report
A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs.
Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw.
NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model.
Raw is intended for training, as are the samples in this dataset as they can be used for regularization.
Possible uses
Regularization images for training models based on Krea 2 Raw
Quality testing
Data source
This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.CogStream
CogStream Dataset
Dataset for CogStream: Context-guided Streaming Video Question Answering.
Overview
CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams.
Statistics:
Split
Videos
QA Pairs
Train
852
55,623
Test
236
15,364
Total
1,088
70,987
Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.huggingface-krew-hackathon2023ko_text2sqlKABR
Dataset Card for KABR: In-Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos
Dataset Summary
We present a novel high-quality dataset for animal behavior recognition from drone videos.
The dataset is focused on Kenyan wildlife and contains behaviors of giraffes, plains zebras, and Grevy's zebras.
The dataset consists of more than 10 hours of annotated videos, and it includes eight different classes, encompassing seven types of animal behavior and an… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR.CUHK-X_Small_Model_Track
CUHK-X — Small Model Track
Multimodal human action recognition (classification).
Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes).
Repository layout
.
├── Training/
│ ├── class_mapping.csv # action_id <-> action_name (40 classes)
│ └── data/
│ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip
│ → HAR/data/<modality>/<action>/<user>/<trial>/<files>
└── Testing/
├── data/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.Customer-Support-ResponsesC3
C3: Cross-View Cross-Modality Correspondence Dataset
Dataset for C3Po: Cross-View Cross-Modality Correspondence with Pointmap Prediction
arXiv | Project Website | GitHub
C3 contains 90K paired floor plans and photos from the Internet across 597 scenes with 153M pixel-level correspondences and 85K camera poses.
Image Pairs
image_pairs/ is split into train/, val, and test/, each with a image_pairs.csv.
image_pairs.csv: Each row represents a plan-photo pair… See the full description on the dataset page: https://huggingface.co/datasets/kwhuang/C3.CulturalBench
CulturalBench - a Robust, Diverse and Challenging Benchmark on Measuring the (Lack of) Cultural Knowledge of LLMs
📌 Resources: Paper | Leaderboard
📘 Description of CulturalBench
CulturalBench is a set of 1,227 human-written and human-verified questions for effectively assessing LLMs’ cultural knowledge, covering 45 global regions including the underrepresented ones like Bangladesh, Zimbabwe, and Peru.
We evaluate models on two setups: CulturalBench-Easy and… See the full description on the dataset page: https://huggingface.co/datasets/kellycyy/CulturalBench.bangladesh-stock-market-dataset
Bangladesh Stock Market Dataset: 27 Years of Open-Source Dhaka Stock Exchange Data with Technical Indicators and Deep Learning Benchmarks
Author: Kawser Sikder
Overview
A comprehensive, open-source financial dataset covering 441 publicly traded instruments across 23 industry sectors of the Dhaka Stock Exchange (DSE), Bangladesh's principal securities market.
Metric
Value
Total Stocks
441
Total Sectors
23
Total Trading Records
1,507,388
Date Range… See the full description on the dataset page: https://huggingface.co/datasets/kawsersikder/bangladesh-stock-market-dataset.news-political-bias-classification-datasetDataset actually from kaggle.
Couldn't find it here so I uploaded it.
platonic-all-experimentsleetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.kraken-trading-data
📈 Kraken Trading Data Collection
Overview
High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis.
This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling.
📊 Included Trading Pairs
Pair
Asset
Base Currency
Typical Daily Volume
XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.
