datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech
@inproceedings{moon-etal-2020-beep,
title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection",
author = "Moon, Jihyung and
Cho, Won Ik and
Lee, Junbum",
booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.KOREATECH-CGH-512-3.6Mu
KOREATECH-CGH
This dataset consists of RGBD–complex hologram pairs designed for training machine learning–based computer-generated holography (ML-CGH) models.It can be used for tasks such as hologram generation, hologram upscaling, and related applications.
The holograms were generated using a layer-based hologram generation method[Article].
Note that this dataset is licensed under the Creative Commons Attribution 4.0 International License Non Commercial (CC BY-NC 4.0).… See the full description on the dataset page: https://huggingface.co/datasets/SPIN-Lab/KOREATECH-CGH-512-3.6Mu.UltraFineWeb-filteredKor-CC-Dumpsfineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0)
It took about 9 hours on A100 80gbx4 to process the dataset.
KOREATECH-CGH-2048-3.6Mu
KOREATECH-CGH
This dataset consists of RGBD–complex hologram pairs designed for training machine learning–based computer-generated holography (ML-CGH) models.It can be used for tasks such as hologram generation, hologram upscaling, and related applications.
The holograms were generated using a layer-based hologram generation method[Article].
Note that this dataset is licensed under the Creative Commons Attribution 4.0 International License Non Commercial (CC BY-NC 4.0).… See the full description on the dataset page: https://huggingface.co/datasets/SPIN-Lab/KOREATECH-CGH-2048-3.6Mu.normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.zeroth_korean
Zeroth-Korean Dataset
Introduction
The Zeroth-Korean dataset is a publicly available speech dataset created for Korean automatic speech recognition (ASR) research and development. This dataset is distributed under the CC BY 4.0 license, allowing anyone to use it freely. The goal of the Zeroth project is to make Korean speech recognition more widely accessible.
Dataset Overview
Total Data: Approximately 51.6 hours of training data and 1.2 hours of test data… See the full description on the dataset page: https://huggingface.co/datasets/kresnik/zeroth_korean.KOREATECH-CGH-1024-3.6Mu
Koreatech-CGH
This dataset consists of RGBD–complex hologram pairs designed for training machine learning–based computer-generated holography (ML-CGH) models.It can be used for tasks such as hologram generation, hologram upscaling, and related applications.
The holograms were generated using a layer-based hologram generation method[Article].
Note that this dataset is licensed under the Creative Commons Attribution 4.0 International License Non Commercial (CC BY-NC 4.0).… See the full description on the dataset page: https://huggingface.co/datasets/SPIN-Lab/KOREATECH-CGH-1024-3.6Mu.korean_textbooks
Massive Korean synthetic dataset
This dataset is a large-scale Korean artificial data set created using Gemini Pro.
It was created using the methodology described in Creation of synthetic textbook-quality datasets in Textbooks Are All You Need.
Data overview
A subset of each dataset does not indicate the contents of that dataset.
Further modification required before use this dataset for training.
본 데이터셋은 바로 사용하기보다는 하고자하는 task에 맞추어 가공 후 사용을 권장드립니다. ex) 로컬 모델을 사용하여 QA 셋으로… See the full description on the dataset page: https://huggingface.co/datasets/maywell/korean_textbooks.kmhas_korean_hate_speechThe K-MHaS (Korean Multi-label Hate Speech) dataset contains 109k utterances from Korean online news comments labeled with 8 fine-grained hate speech classes or Not Hate Speech class.
The fine-grained hate speech classes are politics, origin, physical, age, gender, religion, race, and profanity and these categories are selected in order to reflect the social and historical context.KOREATECH-CGH-256-3.6Mu
Koreatech-CGH
This dataset consists of RGBD–complex hologram pairs designed for training machine learning–based computer-generated holography (ML-CGH) models.It can be used for tasks such as hologram generation, hologram upscaling, and related applications.
The holograms were generated using a layer-based hologram generation method[Article].
Note that this dataset is licensed under the Creative Commons Attribution 4.0 International License Non Commercial (CC BY-NC 4.0).… See the full description on the dataset page: https://huggingface.co/datasets/SPIN-Lab/KOREATECH-CGH-256-3.6Mu.korean-hate-speechHello AI-it!
Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2
samples are:
visual7w_telling_cot: 435299
plotqa_cot: 295354
wiki_ko: 200000
wiki_en: 200000
mulberry_cot_1: 189378
mulberry_cot_2: 102279
sparsetables: 100000
mantis_instruct_cot: 67714
llava_cot_100k: 63019
visual_web_instruct_cot: 47800
chartqa_cot: 45710
docvqa_cot: 36333
tabmwp_cot: 20305
infographicsvqa_cot: 19548
hiertext: 514
celeba-hq-256x256
CelebA-HQ-256x256
CelebA-HQ at 256x256 resolution.
Citation
@article{DBLP:journals/corr/abs-1710-10196,
title={Progressive Growing of GANs for Improved Quality, Stability, and Variation},
author={Tero Karras and Timo Aila and Samuli Laine and Jaakko Lehtinen},
year=2017,
journal={CoRR},
volume={abs/1710.10196}
}
KorMedMCQA
KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
We present KorMedMCQA, the first Korean Medical Multiple-Choice Question
Answering benchmark, derived from professional healthcare licensing
examinations conducted in Korea between 2012 and 2024. The dataset contains
7,469 questions from examinations for doctor, nurse, pharmacist, and dentist,
covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.korean-web-collectionkorean_hate_speech_copyThe K-MHaS (Korean Multi-label Hate Speech) dataset contains 109k utterances from Korean online news comments labeled with 8 fine-grained hate speech classes or Not Hate Speech class.
The fine-grained hate speech classes are politics, origin, physical, age, gender, religion, race, and profanity and these categories are selected in order to reflect the social and historical context.NijiJourney-Prompt-Pairs
NijiJourney Prompt Pairs
A dataset containing txt2img prompt pairs for training on diffusion models
The final goal of this dataset is to create an OpenJourney like model but with NijiJourney images
Korea-AIHub-middlesenior-dialect-speech-train-part2KorMix
KorMix (https://arxiv.org/abs/2512.18834) is a Korean pretraining corpus built by combining five publicly available Korean datasets, applying Korean-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses cross-dataset… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/KorMix.zeroth-korean
Zeroth-Korean
Zeroth-Korean
The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor).
Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.squad_kor_v1
Dataset Card for KorQuAD v1.0
Dataset Summary
KorQuAD 1.0 is a large-scale question-and-answer dataset constructed for Korean machine reading comprehension, and investigate the dataset to understand the distribution of answers and the types of reasoning required to answer the question. This dataset benchmarks the data generating process of SQuAD v1.0 to meet the standard.
Supported Tasks and Leaderboards
question-answering
Languages
Korean… See the full description on the dataset page: https://huggingface.co/datasets/KorQuAD/squad_kor_v1.laion2b_multi_korean_subset_with_image
laion2b_multi_korean_subset_with_image
img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다.
이미지는 9,800,137장입니다.
이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다.
Usage
1. datasets
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train")
>>> dataset.features
{'image': Image(decode=True, id=None),
'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.Cosmopedia-ko-synth
Citation
@misc{KORMo,
author = {Minjun Kim, Hyeonseok Lim, Hangyeol Yoo, Inho Won, Seungwoo Song, Minkyung Cho, Junghun Yuk, Changsu Choi, Dongjae Shin, Huije Lee, Hoyun Song, Alice Oh and KyungTae Lim},
title = {KORMo: Korean Open Reasoning Model for Everyone},
year = {2025},
publisher = {GitHub},
journal = {Technical Report},
paperLink = {\url{https://arxiv.org/abs/2510.09426}},
},
}
korean_mmqa_competitionkorean_datasetOpen-Korean-Historical-Corpus
Open Korean Historical Corpus
Dataset Description
The Open Korean Historical Corpus is a large-scale, openly licensed dataset created to address the lack of accessible data for Korean NLP and historical linguistics.
It contains 17.7 million documents (5.1 billion tokens) compiled from 19 distinct archives, spanning 1,300 years from the 7th century to 2025. The corpus is linguistically diverse, covering Korean (Middle, Early Modern, Modern, North), Classical Chinese, and… See the full description on the dataset page: https://huggingface.co/datasets/seyoungsong/Open-Korean-Historical-Corpus.korea-equity-daily
Korean Equity Daily Prices + DART Filing Impact (한국주식데이터)
Daily settled closes for 2,787 Korean listed companies (KOSPI, KOSDAQ and KONEX) across
321 trading days (2025-05-28 → 2026-09-17), plus a table of what stocks did after each type of
regulatory filing. The per-stock history is no longer cut at 250 days: since the 2026-09-14 publish
the live files gain one row every trading day, and this repository is a dated snapshot of them.
Korean equity data is oddly hard to get. The… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily.
