datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech
@inproceedings{moon-etal-2020-beep,
title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection",
author = "Moon, Jihyung and
Cho, Won Ik and
Lee, Junbum",
booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.Kor-CC-Dumpsfineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0)
It took about 9 hours on A100 80gbx4 to process the dataset.
zeroth_korean
Zeroth-Korean Dataset
Introduction
The Zeroth-Korean dataset is a publicly available speech dataset created for Korean automatic speech recognition (ASR) research and development. This dataset is distributed under the CC BY 4.0 license, allowing anyone to use it freely. The goal of the Zeroth project is to make Korean speech recognition more widely accessible.
Dataset Overview
Total Data: Approximately 51.6 hours of training data and 1.2 hours of test data… See the full description on the dataset page: https://huggingface.co/datasets/kresnik/zeroth_korean.normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.korean_textbooks
Massive Korean synthetic dataset
This dataset is a large-scale Korean artificial data set created using Gemini Pro.
It was created using the methodology described in Creation of synthetic textbook-quality datasets in Textbooks Are All You Need.
Data overview
A subset of each dataset does not indicate the contents of that dataset.
Further modification required before use this dataset for training.
본 데이터셋은 바로 사용하기보다는 하고자하는 task에 맞추어 가공 후 사용을 권장드립니다. ex) 로컬 모델을 사용하여 QA 셋으로… See the full description on the dataset page: https://huggingface.co/datasets/maywell/korean_textbooks.kmhas_korean_hate_speechThe K-MHaS (Korean Multi-label Hate Speech) dataset contains 109k utterances from Korean online news comments labeled with 8 fine-grained hate speech classes or Not Hate Speech class.
The fine-grained hate speech classes are politics, origin, physical, age, gender, religion, race, and profanity and these categories are selected in order to reflect the social and historical context.korean-hate-speechHello AI-it!
Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2
samples are:
visual7w_telling_cot: 435299
plotqa_cot: 295354
wiki_ko: 200000
wiki_en: 200000
mulberry_cot_1: 189378
mulberry_cot_2: 102279
sparsetables: 100000
mantis_instruct_cot: 67714
llava_cot_100k: 63019
visual_web_instruct_cot: 47800
chartqa_cot: 45710
docvqa_cot: 36333
tabmwp_cot: 20305
infographicsvqa_cot: 19548
hiertext: 514
KorMedMCQA
KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
We present KorMedMCQA, the first Korean Medical Multiple-Choice Question
Answering benchmark, derived from professional healthcare licensing
examinations conducted in Korea between 2012 and 2024. The dataset contains
7,469 questions from examinations for doctor, nurse, pharmacist, and dentist,
covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.NijiJourney-Prompt-Pairs
NijiJourney Prompt Pairs
A dataset containing txt2img prompt pairs for training on diffusion models
The final goal of this dataset is to create an OpenJourney like model but with NijiJourney images
Korea-AIHub-middlesenior-dialect-speech-train-part2KorMix
KorMix (https://arxiv.org/abs/2512.18834) is a Korean pretraining corpus built by combining five publicly available Korean datasets, applying Korean-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses cross-dataset… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/KorMix.zeroth-korean
Zeroth-Korean
Zeroth-Korean
The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor).
Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.squad_kor_v1
Dataset Card for KorQuAD v1.0
Dataset Summary
KorQuAD 1.0 is a large-scale question-and-answer dataset constructed for Korean machine reading comprehension, and investigate the dataset to understand the distribution of answers and the types of reasoning required to answer the question. This dataset benchmarks the data generating process of SQuAD v1.0 to meet the standard.
Supported Tasks and Leaderboards
question-answering
Languages
Korean… See the full description on the dataset page: https://huggingface.co/datasets/KorQuAD/squad_kor_v1.laion2b_multi_korean_subset_with_image
laion2b_multi_korean_subset_with_image
img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다.
이미지는 9,800,137장입니다.
이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다.
Usage
1. datasets
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train")
>>> dataset.features
{'image': Image(decode=True, id=None),
'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.korean_datasetkorea-equity-daily
Korean Equity Daily Prices + DART Filing Impact (한국주식데이터)
Daily settled closes for 2,787 Korean listed companies (KOSPI, KOSDAQ and KONEX) across
321 trading days (2025-05-28 → 2026-09-17), plus a table of what stocks did after each type of
regulatory filing. The per-stock history is no longer cut at 250 days: since the 2026-09-14 publish
the live files gain one row every trading day, and this repository is a dated snapshot of them.
Korean equity data is oddly hard to get. The… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily.nvidia-post-train-v1-en_kr-filtered-origkor_nli
Dataset Card for "kor_nli"
Dataset Summary
Korean Natural Language Inference datasets.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
multi_nli
Size of downloaded dataset files: 42.11 MB
Size of the generated dataset: 84.72 MB
Total amount of disk used: 126.85 MB
An example of 'train' looks as follows.
snli
Size of downloaded… See the full description on the dataset page: https://huggingface.co/datasets/kakaobrain/kor_nli.kor_unsmileKOREAN-WEBTEXT
KOREAN-WEBTEXT
KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources:
cc100
oscar-corpus/OSCAR-2201
oscar-corpus/OSCAR-2109
oscar-corpus/OSCAR-2301
ontocord/CulturaY
Additional credible internet sources collected by out team
(We are working to add more sources)
The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.ko-raw-28Bkorea_speech_mfa_aligned_validationkorean-assembly-minutes
대한민국 국회 회의록 아카이브
국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다.
본회의와 각 위원회 회의록이 모두 들어 있다.
수록 기간: 1948~1993
회의 수: 1,951건
본문 분량: 65,306,444자
구성
연도별 JSONL(gzip) 한 덩이다.
from datasets import load_dataset
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train")
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train")
필드
이름
설명
meeting_key
회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.fineweb-2-edu-koreanmteb-kor-retrieval-full_naivekorean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.squad_kor_v1
