datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0)
It took about 9 hours on A100 80gbx4 to process the dataset.
KorMedMCQA
KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
We present KorMedMCQA, the first Korean Medical Multiple-Choice Question
Answering benchmark, derived from professional healthcare licensing
examinations conducted in Korea between 2012 and 2024. The dataset contains
7,469 questions from examinations for doctor, nurse, pharmacist, and dentist,
covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.korea-equity-daily
Korean Equity Daily Prices + DART Filing Impact (한국주식데이터)
Daily settled closes for 2,787 Korean listed companies (KOSPI, KOSDAQ and KONEX) across
321 trading days (2025-05-28 → 2026-09-17), plus a table of what stocks did after each type of
regulatory filing. The per-stock history is no longer cut at 250 days: since the 2026-09-14 publish
the live files gain one row every trading day, and this repository is a dated snapshot of them.
Korean equity data is oddly hard to get. The… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily.kor_unsmileKOREAN-WEBTEXT
KOREAN-WEBTEXT
KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources:
cc100
oscar-corpus/OSCAR-2201
oscar-corpus/OSCAR-2109
oscar-corpus/OSCAR-2301
ontocord/CulturaY
Additional credible internet sources collected by out team
(We are working to add more sources)
The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.ko-raw-28Bkorean-assembly-minutes
대한민국 국회 회의록 아카이브
국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다.
본회의와 각 위원회 회의록이 모두 들어 있다.
수록 기간: 1948~1993
회의 수: 1,951건
본문 분량: 65,306,444자
구성
연도별 JSONL(gzip) 한 덩이다.
from datasets import load_dataset
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train")
ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train")
필드
이름
설명
meeting_key
회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.fineweb-2-edu-koreanKorMedMCQA-V
KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination
KorMedMCQA-V is a multimodal multiple-choice question answering benchmark for evaluating vision-language models on the Korean Medical Licensing Examination. The dataset consists of 1,534 questions with 2,043 associated medical images from Korean Medical Licensing Examinations (2012-2023).
Dataset Summary
Total Questions: 1,534
Total Images: 2,043 (avg… See the full description on the dataset page: https://huggingface.co/datasets/seongsubae/KorMedMCQA-V.fineweb-2-edu-korean-score-2korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.Wiki.korpus
Viki korpusi na srpskom i hrvatskom jeziku
Sveža verzija, 1. maj 2026!
Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati.
Preko 670.000 očišćenih članaka, sa preko 310 miliona reči.
Svaki dokument je u zasebnoj JSON liniji.
Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta
Moguće filtiranje skupa po jeziku ili projektu.
Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.kor-sts
Korean Semantic Textual Similarity (KorSTS) Dataset
For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK
This dataset was prepared by converting tsv files from this repository. The idea was to share the dataset for broader audience. I am not an original author of it.
Because of the specifity of read_csv method from Pandas library, there are couple of observations, which had to be deleted because of the formatting (54… See the full description on the dataset page: https://huggingface.co/datasets/dkoterwa/kor-sts.korean_toxic_datasets데이터 출처 :
https://aihub.or.kr/aihubdata/data/view.do?currMenu=&topMenu=&aihubDataSe=data&dataSetSn=71788
laion2B-multi-korean-subset
laion2B-multi-korean-subset
About dataset
a subset data of laion/laion2B-multi, including only korean
Lisence
CC-BY-4.0
Data Structure
Data Instance
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion2B-multi-korean-subset")
>>> dataset
DatasetDict({
train: Dataset({
features: ['SAMPLE_ID', 'URL', 'TEXT', 'HEIGHT', 'WIDTH', 'LICENSE', 'LANGUAGE', 'NSFW', 'similarity'],
num_rows: 11376263… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2B-multi-korean-subset.korean-qa-gen
Attribution
MTEB-format derivative of ziozzang/Korean_QA_gen_datasets (Korean QA). Query = question; corpus = answer.
korea-equity-daily-2026-08
This repository is frozen. It is the 2026-08 snapshot of
aikstockdata/korea-equity-daily,
taken once and never updated. Cite this repository when you need the numbers to
still be there later.
The rolling dataset is regenerated every trading evening and its main is
overwritten, so figures quoted from it do not survive. Figures here do.
The live, always-current files are at https://aikstockdata.com/data/public/.
Korean Equity Daily Prices + DART Filing Impact
Daily settled… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily-2026-08.laion-translated-to-en-korean-subset
laion-translated-to-en-korean-subset
About dataset
a subset data of laion/laion2B-multi-joined-translated-to-en and laion/laion1B-nolang-joined-translated-to-en, including only korean
Lisence
CC-BY-4.0
Data Structure
Data Instance
>>> from datasets import load_dataset
>>> dataset = load_dataset("Bingsu/laion-translated-to-en-korean-subset")
>>> dataset
DatasetDict({
train: Dataset({
features: ['hash', 'URL', 'TEXT', 'ENG TEXT'… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion-translated-to-en-korean-subset.korean_law_open_data_precedents
Dataset Card for Dataset Name
공지사항
인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다.
사용상 주의사항
사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다.
그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다.
사용에 참고하시기 바랍니다.
Dataset Summary
2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다.
그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.ipfs_korea_laws_ir
Korea (ROK) legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_korea_laws (revision a174ce3d2e51385dd21584ba4aa7618c97910d86) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Korea (ROK) prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_korea_laws_ir.korea-household-egocentric-samples
60BASE · Korea Household Egocentric Samples
Sixteen household video excerpts from 60BASE's Korea household collection, including every source video in its sample folder as checked on September 20, 2026.
Use this sample pack to inspect the footage and discuss a full-recording request or a custom collection brief.
Discuss a data project · 60BASE · Email
Included in this release
Specification
Video
16 MP4 clips × 20 seconds; 5 minutes 20 seconds total
Tasks
Dishwashing… See the full description on the dataset page: https://huggingface.co/datasets/60base/korea-household-egocentric-samples.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.KoreanFinancialTableQAChunkRetrievalmedical-reasoning-train-kormedmcqa
Information
This data includes the training data from KorMedMCQA, as well as a portion of the trainind data from the additional KorMedMCQA support set (private).
This dataset is based on the responses generated by gemini-flash-thinking-exp-01-21 model and has undergone MANUAL rejection sampling.
KoreanMarketReportChunkRetrievalbankacilik-duzenleme-korpusu
Bankacılık Düzenleme Korpusu
Türkiye'de bankacılık ve finans sektörünü ilgilendiren, kamuya açık düzenleyici kurum belgelerinden derlenen açık bir Türkçe metin korpusu. Amaç, bankacılık mevzuatı ve düzenleyici uygulamasına ilişkin metinlerin makine-okunabilir, aranabilir ve araştırmaya açık tek bir yerde toplanmasıdır.
Sürüm v0.5 — beş kaynak, toplam 11,937 belge.
Metodolojik olarak Türk İçtihat Korpusu şablonunu izler: nezaket sınırlı tarama, ham dosya sağlaması (SHA-256)… See the full description on the dataset page: https://huggingface.co/datasets/selimogluburak/bankacilik-duzenleme-korpusu.korfinvdr-financial_magazineKorFinVDR: Financial Magazine
This draft dataset card covers financial_magazine, a corpus of Korea Deposit Insurance Corporation financial-risk review articles for Korean visual document retrieval and complex-document question answering. It is one of the five subsets comprising the KorFinVDR Benchmark.
Links
GitHub: https://github.com/Marker-Inc-Korea/kor-fin-vdr
Collection: https://huggingface.co/collections/MarkrAI/korfinvdr
Dataset Summary
Description:… See the full description on the dataset page: https://huggingface.co/datasets/MarkrAI/korfinvdr-financial_magazine.kor_glue
Dataset Card for "kor_glue"
More Information needed
Source Data Citation Information
@article{warstadt2018neural,
title={Neural Network Acceptability Judgments},
author={Warstadt, Alex and Singh, Amanpreet and Bowman, Samuel R},
journal={arXiv preprint arXiv:1805.12471},
year={2018}
}
@inproceedings{wang2019glue,
title={{GLUE}: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding},
author={Wang, Alex and Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/kor_glue.CC-korean
CC-korean
한국어 웹 텍스트 코퍼스. CommonCrawl WET 파일에서 추출한 한국어 문서 데이터셋입니다.
데이터 수집 방법
소스: CommonCrawl WET 파일 (HTML 제거된 텍스트 포맷)
수집 스냅샷: CC-MAIN-2025-05 ~ CC-MAIN-2026-12 (총 15개 스냅샷, 2025~2026년)
언어 필터: fasttext lid.176.bin 모델로 한국어(ko) 감지, 신뢰도 0.5 이상만 수집
1차 품질 필터: 수집 단계에서 휴리스틱 줄 단위 노이즈 제거 + KenLM perplexity 필터 적용
필터링
수집 후 추가로 2단계 필터링 적용.
줄 단위 인라인 제거:
URL, 이모티콘, 대괄호 태그, 기자 바이라인, 출처 표기 등
줄 단위 제거 항목:
한국어 없는 줄 (영문·중문·일문 등)
UI/네비게이션 잔재 (로그인, 더보기, 공유 버튼, 파이프 구분자 등)
날짜·메타 태그… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/CC-korean.KoreanQAChunkRetrieval
