CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.imagetext-generation1M<n<10M557 likes5.8k downloads3mo agoHugging Face02nayohan /korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech @inproceedings{moon-etal-2020-beep, title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection", author = "Moon, Jihyung and Cho, Won Ik and Lee, Junbum", booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media", month = jul, year = "2020", address = "Online", publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.text1K<n<10K2 likes4.1k downloads2y agoHugging Face03MLP-KTLim /Kor-CC-Dumpstext100M<n<1B0 likes3.5k downloads8mo agoHugging Face04minpeter /fineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0) It took about 9 hours on A100 80gbx4 to process the dataset. tabular10M<n<100M1 likes3.3k downloads1y agoHugging Face05kresnik /zeroth_korean Zeroth-Korean Dataset Introduction The Zeroth-Korean dataset is a publicly available speech dataset created for Korean automatic speech recognition (ASR) research and development. This dataset is distributed under the CC BY 4.0 license, allowing anyone to use it freely. The goal of the Zeroth project is to make Korean speech recognition more widely accessible. Dataset Overview Total Data: Approximately 51.6 hours of training data and 1.2 hours of test data… See the full description on the dataset page: https://huggingface.co/datasets/kresnik/zeroth_korean.audio10K<n<100K23 likes2.4k downloads2y agoHugging Face06mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.2k downloads4mo agoHugging Face07maywell /korean_textbooks Massive Korean synthetic dataset This dataset is a large-scale Korean artificial data set created using Gemini Pro. It was created using the methodology described in Creation of synthetic textbook-quality datasets in Textbooks Are All You Need. Data overview A subset of each dataset does not indicate the contents of that dataset. Further modification required before use this dataset for training. 본 데이터셋은 바로 사용하기보다는 하고자하는 task에 맞추어 가공 후 사용을 권장드립니다. ex) 로컬 모델을 사용하여 QA 셋으로… See the full description on the dataset page: https://huggingface.co/datasets/maywell/korean_textbooks.text1M<n<10M124 likes2.2k downloads3y agoHugging Face08jeanlee /kmhas_korean_hate_speechThe K-MHaS (Korean Multi-label Hate Speech) dataset contains 109k utterances from Korean online news comments labeled with 8 fine-grained hate speech classes or Not Hate Speech class. The fine-grained hate speech classes are politics, origin, physical, age, gender, religion, race, and profanity and these categories are selected in order to reflect the social and historical context.texttext-classification100K<n<1M24 likes2k downloads4y agoHugging Face09AI-it /korean-hate-speechgatedHello AI-it! text1K<n<10K4 likes1.7k downloads5y agoHugging Face10KORMo-VL /Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2 samples are: visual7w_telling_cot: 435299 plotqa_cot: 295354 wiki_ko: 200000 wiki_en: 200000 mulberry_cot_1: 189378 mulberry_cot_2: 102279 sparsetables: 100000 mantis_instruct_cot: 67714 llava_cot_100k: 63019 visual_web_instruct_cot: 47800 chartqa_cot: 45710 docvqa_cot: 36333 tabmwp_cot: 20305 infographicsvqa_cot: 19548 hiertext: 514 image1M<n<10M0 likes1.6k downloads7mo agoHugging Face11sean0042 /KorMedMCQA KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations We present KorMedMCQA, the first Korean Medical Multiple-Choice Question Answering benchmark, derived from professional healthcare licensing examinations conducted in Korea between 2012 and 2024. The dataset contains 7,469 questions from examinations for doctor, nurse, pharmacist, and dentist, covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.tabularquestion-answering1K<n<10K43 likes1.5k downloads2y agoHugging Face12Korakoe /NijiJourney-Prompt-Pairs NijiJourney Prompt Pairs A dataset containing txt2img prompt pairs for training on diffusion models The final goal of this dataset is to create an OpenJourney like model but with NijiJourney images image1K<n<10K16 likes1.2k downloads4y agoHugging Face13Suchae /Korea-AIHub-middlesenior-dialect-speech-train-part2audio100K<n<1M0 likes1.1k downloads2y agoHugging Face14AdaMLLab /KorMix KorMix (https://arxiv.org/abs/2512.18834) is a Korean pretraining corpus built by combining five publicly available Korean datasets, applying Korean-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses cross-dataset… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/KorMix.texttext-generation100M<n<1B2 likes977 downloads5mo agoHugging Face15Bingsu /zeroth-korean Zeroth-Korean Zeroth-Korean The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor). Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.audioautomatic-speech-recognition10K<n<100K46 likes976 downloads4y agoHugging Face16KorQuAD /squad_kor_v1 Dataset Card for KorQuAD v1.0 Dataset Summary KorQuAD 1.0 is a large-scale question-and-answer dataset constructed for Korean machine reading comprehension, and investigate the dataset to understand the distribution of answers and the types of reasoning required to answer the question. This dataset benchmarks the data generating process of SQuAD v1.0 to meet the standard. Supported Tasks and Leaderboards question-answering Languages Korean… See the full description on the dataset page: https://huggingface.co/datasets/KorQuAD/squad_kor_v1.textquestion-answering10K<n<100K34 likes956 downloads2y agoHugging Face17Bingsu /laion2b_multi_korean_subset_with_image laion2b_multi_korean_subset_with_image img2dataset을 통해 다운로드에 성공한 Bingsu/laion2B-multi-korean-subset 이미지를 정리한 데이터셋입니다. 이미지는 9,800,137장입니다. 이미지는 짧은 쪽 길이가 256이 되도록 리사이즈 되었으며, 품질 100인 webp파일로 다운로드 되었습니다. Usage 1. datasets >>> from datasets import load_dataset >>> dataset = load_dataset("Bingsu/laion2b_multi_korean_subset_with_image", streaming=True, split="train") >>> dataset.features {'image': Image(decode=True, id=None), 'text': Value(dtype='string', id=None)… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2b_multi_korean_subset_with_image.imagefeature-extraction100K<n<1M6 likes943 downloads4y agoHugging Face18JaepaX /korean_datasetaudioautomatic-speech-recognition10K<n<100K4 likes841 downloads2y agoHugging Face19aikstockdata /korea-equity-daily Korean Equity Daily Prices + DART Filing Impact (한국주식데이터) Daily settled closes for 2,787 Korean listed companies (KOSPI, KOSDAQ and KONEX) across 321 trading days (2025-05-28 → 2026-09-17), plus a table of what stocks did after each type of regulatory filing. The per-stock history is no longer cut at 250 days: since the 2026-09-14 publish the live files gain one row every trading day, and this repository is a dated snapshot of them. Korean equity data is oddly hard to get. The… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily.tabulartime-series-forecasting100K<n<1M1 likes816 downloads5d agoHugging Face20kormo-lm /nvidia-post-train-v1-en_kr-filtered-origtext1M<n<10M0 likes729 downloads1y agoHugging Face21kakaobrain /kor_nli Dataset Card for "kor_nli" Dataset Summary Korean Natural Language Inference datasets. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances multi_nli Size of downloaded dataset files: 42.11 MB Size of the generated dataset: 84.72 MB Total amount of disk used: 126.85 MB An example of 'train' looks as follows. snli Size of downloaded… See the full description on the dataset page: https://huggingface.co/datasets/kakaobrain/kor_nli.texttext-classification100K<n<1M23 likes695 downloads2y agoHugging Face22smilegate-ai /kor_unsmiletabular10K<n<100K4 likes654 downloads5y agoHugging Face23HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes647 downloads2y agoHugging Face24Beetle-Data /ko-raw-28Btabular10M<n<100M0 likes623 downloads4mo agoHugging Face25AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes527 downloads8mo agoHugging Face26seoulraphaellee /korean-assembly-minutes 대한민국 국회 회의록 아카이브 국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다. 본회의와 각 위원회 회의록이 모두 들어 있다. 수록 기간: 1948~1993 회의 수: 1,951건 본문 분량: 65,306,444자 구성 연도별 JSONL(gzip) 한 덩이다. from datasets import load_dataset ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train") ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train") 필드 이름 설명 meeting_key 회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.tabulartext-generation10K<n<100K0 likes398 downloads17d agoHugging Face27minpeter /fineweb-2-edu-koreantabular1M<n<10M5 likes393 downloads1y agoHugging Face28seongil-dn /mteb-kor-retrieval-full_naivetext1M<n<10M0 likes384 downloads2y agoHugging Face29mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes384 downloads26d agoHugging Face30yjoonjang /squad_kor_v1text10K<n<100K1 likes382 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.