CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01minpeter /fineweb-2-edu-korean-rawHuggingFaceFW/fineweb-2 (v2.1.0) It took about 9 hours on A100 80gbx4 to process the dataset. tabular10M<n<100M1 likes3.3k downloads1y agoHugging Face02sean0042 /KorMedMCQA KorMedMCQA : Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations We present KorMedMCQA, the first Korean Medical Multiple-Choice Question Answering benchmark, derived from professional healthcare licensing examinations conducted in Korea between 2012 and 2024. The dataset contains 7,469 questions from examinations for doctor, nurse, pharmacist, and dentist, covering a wide range of medical disciplines. We evaluate the performance… See the full description on the dataset page: https://huggingface.co/datasets/sean0042/KorMedMCQA.tabularquestion-answering1K<n<10K43 likes1.5k downloads2y agoHugging Face03aikstockdata /korea-equity-daily Korean Equity Daily Prices + DART Filing Impact (한국주식데이터) Daily settled closes for 2,787 Korean listed companies (KOSPI, KOSDAQ and KONEX) across 321 trading days (2025-05-28 → 2026-09-17), plus a table of what stocks did after each type of regulatory filing. The per-stock history is no longer cut at 250 days: since the 2026-09-14 publish the live files gain one row every trading day, and this repository is a dated snapshot of them. Korean equity data is oddly hard to get. The… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily.tabulartime-series-forecasting100K<n<1M1 likes816 downloads6d agoHugging Face04smilegate-ai /kor_unsmiletabular10K<n<100K4 likes654 downloads5y agoHugging Face05HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes647 downloads2y agoHugging Face06Beetle-Data /ko-raw-28Btabular10M<n<100M0 likes623 downloads4mo agoHugging Face07seoulraphaellee /korean-assembly-minutes 대한민국 국회 회의록 아카이브 국회 회의록 원문(record.assembly.go.kr) PDF 에서 본문을 뽑아 모은 것이다. 본회의와 각 위원회 회의록이 모두 들어 있다. 수록 기간: 1948~1993 회의 수: 1,951건 본문 분량: 65,306,444자 구성 연도별 JSONL(gzip) 한 덩이다. from datasets import load_dataset ds = load_dataset("seoulraphaellee/korean-assembly-minutes", split="train") ds = load_dataset("seoulraphaellee/korean-assembly-minutes", data_files="data/2026.jsonl.gz", split="train") 필드 이름 설명 meeting_key 회의 식별자 (record:<id>)… See the full description on the dataset page: https://huggingface.co/datasets/seoulraphaellee/korean-assembly-minutes.tabulartext-generation10K<n<100K0 likes398 downloads17d agoHugging Face08minpeter /fineweb-2-edu-koreantabular1M<n<10M5 likes393 downloads1y agoHugging Face09seongsubae /KorMedMCQA-V KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination KorMedMCQA-V is a multimodal multiple-choice question answering benchmark for evaluating vision-language models on the Korean Medical Licensing Examination. The dataset consists of 1,534 questions with 2,043 associated medical images from Korean Medical Licensing Examinations (2012-2023). Dataset Summary Total Questions: 1,534 Total Images: 2,043 (avg… See the full description on the dataset page: https://huggingface.co/datasets/seongsubae/KorMedMCQA-V.tabularvisual-question-answering1K<n<10K9 likes320 downloads7mo agoHugging Face10minpeter /fineweb-2-edu-korean-score-2tabular10M<n<100M3 likes308 downloads1y agoHugging Face11gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads4mo agoHugging Face12procesaur /Wiki.korpus Viki korpusi na srpskom i hrvatskom jeziku Sveža verzija, 1. maj 2026! Očišćen i filtriran skup pet projekata: Vikipedija, Vikizvornik, Vikiknjige, Vikivesti i Vikicitati. Preko 670.000 očišćenih članaka, sa preko 310 miliona reči. Svaki dokument je u zasebnoj JSON liniji. Novi metapodaci! Kategorije, broj reči i postotak ćiriličnog teksta Moguće filtiranje skupa po jeziku ili projektu. Wiki corpora in Serbian and… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/Wiki.korpus.tabulartext-generation1M<n<10M1 likes260 downloads2mo agoHugging Face13dkoterwa /kor-sts Korean Semantic Textual Similarity (KorSTS) Dataset For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK This dataset was prepared by converting tsv files from this repository. The idea was to share the dataset for broader audience. I am not an original author of it. Because of the specifity of read_csv method from Pandas library, there are couple of observations, which had to be deleted because of the formatting (54… See the full description on the dataset page: https://huggingface.co/datasets/dkoterwa/kor-sts.tabular1K<n<10K2 likes257 downloads3y agoHugging Face14josephnam /korean_toxic_datasets데이터 출처 : https://aihub.or.kr/aihubdata/data/view.do?currMenu=&topMenu=&aihubDataSe=data&dataSetSn=71788 tabular10K<n<100K0 likes251 downloads2y agoHugging Face15Bingsu /laion2B-multi-korean-subset laion2B-multi-korean-subset About dataset a subset data of laion/laion2B-multi, including only korean Lisence CC-BY-4.0 Data Structure Data Instance >>> from datasets import load_dataset >>> dataset = load_dataset("Bingsu/laion2B-multi-korean-subset") >>> dataset DatasetDict({ train: Dataset({ features: ['SAMPLE_ID', 'URL', 'TEXT', 'HEIGHT', 'WIDTH', 'LICENSE', 'LANGUAGE', 'NSFW', 'similarity'], num_rows: 11376263… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion2B-multi-korean-subset.imagefeature-extraction10M<n<100M14 likes230 downloads4y agoHugging Face16bowang0911 /korean-qa-gen Attribution MTEB-format derivative of ziozzang/Korean_QA_gen_datasets (Korean QA). Query = question; corpus = answer. tabulartext-retrieval1K<n<10K0 likes208 downloads3mo agoHugging Face17aikstockdata /korea-equity-daily-2026-08 This repository is frozen. It is the 2026-08 snapshot of aikstockdata/korea-equity-daily, taken once and never updated. Cite this repository when you need the numbers to still be there later. The rolling dataset is regenerated every trading evening and its main is overwritten, so figures quoted from it do not survive. Figures here do. The live, always-current files are at https://aikstockdata.com/data/public/. Korean Equity Daily Prices + DART Filing Impact Daily settled… See the full description on the dataset page: https://huggingface.co/datasets/aikstockdata/korea-equity-daily-2026-08.tabulartime-series-forecasting100K<n<1M0 likes208 downloads10d agoHugging Face18Bingsu /laion-translated-to-en-korean-subset laion-translated-to-en-korean-subset About dataset a subset data of laion/laion2B-multi-joined-translated-to-en and laion/laion1B-nolang-joined-translated-to-en, including only korean Lisence CC-BY-4.0 Data Structure Data Instance >>> from datasets import load_dataset >>> dataset = load_dataset("Bingsu/laion-translated-to-en-korean-subset") >>> dataset DatasetDict({ train: Dataset({ features: ['hash', 'URL', 'TEXT', 'ENG TEXT'… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/laion-translated-to-en-korean-subset.tabularfeature-extraction10M<n<100M7 likes190 downloads4y agoHugging Face19joonhok-exo-ai /korean_law_open_data_precedents Dataset Card for Dataset Name 공지사항 인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다. 사용상 주의사항 사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다. 그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다. 사용에 참고하시기 바랍니다. Dataset Summary 2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다. 그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.tabular10K<n<100K45 likes173 downloads10mo agoHugging Face20justicedao /ipfs_korea_laws_ir Korea (ROK) legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_korea_laws (revision a174ce3d2e51385dd21584ba4aa7618c97910d86) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Korea (ROK) prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_korea_laws_ir.tabulartext-retrieval1M<n<10M0 likes159 downloads2d agoHugging Face2160base /korea-household-egocentric-samples 60BASE · Korea Household Egocentric Samples Sixteen household video excerpts from 60BASE's Korea household collection, including every source video in its sample folder as checked on September 20, 2026. Use this sample pack to inspect the footage and discuss a full-recording request or a custom collection brief. Discuss a data project · 60BASE · Email Included in this release Specification Video 16 MP4 clips × 20 seconds; 5 minutes 20 seconds total Tasks Dishwashing… See the full description on the dataset page: https://huggingface.co/datasets/60base/korea-household-egocentric-samples.tabularn<1K1 likes158 downloads6d agoHugging Face22Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes153 downloads1mo agoHugging Face23bowang0911 /KoreanFinancialTableQAChunkRetrievaltabular1K<n<10K0 likes134 downloads3mo agoHugging Face24ChuGyouk /medical-reasoning-train-kormedmcqa Information This data includes the training data from KorMedMCQA, as well as a portion of the trainind data from the additional KorMedMCQA support set (private). This dataset is based on the responses generated by gemini-flash-thinking-exp-01-21 model and has undergone MANUAL rejection sampling. tabular1K<n<10K9 likes126 downloads2y agoHugging Face25bowang0911 /KoreanMarketReportChunkRetrievaltabular1K<n<10K0 likes125 downloads3mo agoHugging Face26selimogluburak /bankacilik-duzenleme-korpusu Bankacılık Düzenleme Korpusu Türkiye'de bankacılık ve finans sektörünü ilgilendiren, kamuya açık düzenleyici kurum belgelerinden derlenen açık bir Türkçe metin korpusu. Amaç, bankacılık mevzuatı ve düzenleyici uygulamasına ilişkin metinlerin makine-okunabilir, aranabilir ve araştırmaya açık tek bir yerde toplanmasıdır. Sürüm v0.5 — beş kaynak, toplam 11,937 belge. Metodolojik olarak Türk İçtihat Korpusu şablonunu izler: nezaket sınırlı tarama, ham dosya sağlaması (SHA-256)… See the full description on the dataset page: https://huggingface.co/datasets/selimogluburak/bankacilik-duzenleme-korpusu.tabulartext-classification10K<n<100K0 likes117 downloads4d agoHugging Face27MarkrAI /korfinvdr-financial_magazineKorFinVDR: Financial Magazine This draft dataset card covers financial_magazine, a corpus of Korea Deposit Insurance Corporation financial-risk review articles for Korean visual document retrieval and complex-document question answering. It is one of the five subsets comprising the KorFinVDR Benchmark. Links GitHub: https://github.com/Marker-Inc-Korea/kor-fin-vdr Collection: https://huggingface.co/collections/MarkrAI/korfinvdr Dataset Summary Description:… See the full description on the dataset page: https://huggingface.co/datasets/MarkrAI/korfinvdr-financial_magazine.imagedocument-question-answering10K<n<100K1 likes116 downloads2mo agoHugging Face28KETI-NLP /kor_glue Dataset Card for "kor_glue" More Information needed Source Data Citation Information @article{warstadt2018neural, title={Neural Network Acceptability Judgments}, author={Warstadt, Alex and Singh, Amanpreet and Bowman, Samuel R}, journal={arXiv preprint arXiv:1805.12471}, year={2018} } @inproceedings{wang2019glue, title={{GLUE}: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding}, author={Wang, Alex and Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/kor_glue.tabular100K<n<1M0 likes113 downloads3y agoHugging Face29chaannwooff /CC-korean CC-korean 한국어 웹 텍스트 코퍼스. CommonCrawl WET 파일에서 추출한 한국어 문서 데이터셋입니다. 데이터 수집 방법 소스: CommonCrawl WET 파일 (HTML 제거된 텍스트 포맷) 수집 스냅샷: CC-MAIN-2025-05 ~ CC-MAIN-2026-12 (총 15개 스냅샷, 2025~2026년) 언어 필터: fasttext lid.176.bin 모델로 한국어(ko) 감지, 신뢰도 0.5 이상만 수집 1차 품질 필터: 수집 단계에서 휴리스틱 줄 단위 노이즈 제거 + KenLM perplexity 필터 적용 필터링 수집 후 추가로 2단계 필터링 적용. 줄 단위 인라인 제거: URL, 이모티콘, 대괄호 태그, 기자 바이라인, 출처 표기 등 줄 단위 제거 항목: 한국어 없는 줄 (영문·중문·일문 등) UI/네비게이션 잔재 (로그인, 더보기, 공유 버튼, 파이프 구분자 등) 날짜·메타 태그… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/CC-korean.tabular1M<n<10M2 likes111 downloads5mo agoHugging Face30bowang0911 /KoreanQAChunkRetrievaltabular1K<n<10K0 likes108 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.