datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasetExperimentDATA_knowledge_distillation_vs_fine_tuningPDExplBenchMultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌
CVPR 2026 (Main)
This repository provides the datasets for
“MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta
Paper Link
https://arxiv.org/abs/2511.22989
Github Repository
For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.kkhkkknkb-books
open-rdl-books
Dataset Description
Language
dan, dansk, Danish
License
Public Domain, cc0-1.0
Dataset Summary
Documents from the Royal Danish Library published between 1750 and 1930.
The dataset has each page of each document in image and text format. The text was extracted with OCR.
The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.K2Datasets
K2 Dataset Card
The following data mix was used to train K2 and achieve results in line with Llama 2 70B.
Dataset Details
K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below.
Dataset Description: Stage 1
Dataset
Starting Tokens
Multiplier
Total Tokens
% of Total
dm-math
4.33B
3x
13B
1%
pubmed-abstracts (from the Pile)
4.77B
3x
14.3B
1.1%
uspto (from the Pile)
4.77B
3x… See the full description on the dataset page: https://huggingface.co/datasets/IFM/K2Datasets.kjv-bibledojo_stock_kline
Languages: 简体中文 · English
dojo_stock_kline — Stock Daily Bars
Overview
Daily OHLCV history for US, CN, and HK equities, including cumulative adjustment factors and dividend/split flags.
Files
File
Description
data.parquet
Market-wide daily bars
Key Fields
Field
Description
symbol
Join key
kline_t
Bar interval; snapshots use "1D" for daily bars
bar_time
Bar timestamp (trade date)
open / high / low /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_kline.Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.ultrafeedback-binarized-preferences-cleaned-kto
UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO
A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback
This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences,
and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback.
Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.xfield-radar-dataset-20260915
XField radar dataset — formal snapshot, 2026-09-15
Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json.
This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.ASearcher-Local-KnowledgeKodCode-V1-SFT-R1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.dojo_benchmark_kline
Languages: 简体中文 · English
dojo_benchmark_kline — Benchmark Index Bars
Overview
Daily OHLCV for major broad and representative indices across US, CN, and HK (e.g. ^SPX, ^HSI, 000300.SS). Each row is one index on one trade date.
Files
File
Description
data.parquet
Index daily bars
Key Fields
Field
Description
symbol
Index code (e.g. ^SPX, 000300.SS)
kline_t
Bar interval; "1D" for daily bars
bar_time… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_benchmark_kline.dojo_forex_kline
Languages: 简体中文 · English
dojo_forex_kline — FX Daily Bars
Overview
Daily OHLC and amplitude for major currency pairs. Used to convert revenue, profit, and other filing amounts into a single currency when report currency and listing/analysis currency differ.
Typical case: regional revenue in HKD in dojo_main_income while analysis targets USD — apply HKDUSD (or equivalent) at the report date.
Intended Use: Cross-Currency Revenue Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_forex_kline.FineNewsTestSampleOmni-MATH
Dataset Card for Omni-MATH
Recent advancements in AI, particularly in large language models (LLMs), have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To mitigate this limitation, we propose a comprehensive and challenging benchmark specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/KbsdJames/Omni-MATH.KodCode-V1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1.bolAIndia
bolAIndia
Human-side speech from production call recordings, cut into utterance-level
chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR
providers. Each row keeps the transcript, the provider's confidence, and full
provenance back to the source recording.
Sources
One config per transcription system, so their output stays separable.
config (source_id)
provider
model
hours
rows
shards
vendor-a
vendor-a
undisclosed
420.03
480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.fineweb-edu-100b-shufflequasar-axrvi-v10ClimbLab-JaJapanese / 日本語版
ClimbLab-Ja
ClimbLab-Ja is a high-quality 300-billion-token Japanese corpus with 20 clusters. It is a Japanese adaptation of the nvidia/Nemotron-ClimbLab approach. Based on LLM-jp Corpus v4, we semantically reorganized and filtered the dataset into 20 distinct clusters, resulting in a high-quality 300-billion-token corpus. Specifically, we first grouped the data into 1,000 groups based on topic information. Then we assigned six scores from 0 to 5 to each group… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbLab-Ja.Raon-OpenTTS-Pool
Raon-OpenTTS-Pool
Technical Report
Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training,
constructed from 8 publicly available speech corpora and a set of web-sourced recordings.
It is the training data behind Raon-OpenTTS,
an open TTS model that performs on par with state-of-the-art closed-data systems.
615K hours of speech audio
239.7M speech segments
11 source datasets aggregated into a unified format
All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.Kvasir-VQA-x1
Kvasir-VQA-x1
A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 on GitHub |
Original Image from Kvasir-VQA(Simula Datasets) |
Paper
🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate!
Overview
Kvasir-VQA-x1 is a large-scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/Kvasir-VQA-x1.damru-knowledge
🐕 Damru Knowledge
A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students.
The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/.
📦 What's inside
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo.
See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information.
needle-resultsKMMLU
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.
