CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes306 downloads4mo agoHugging Face02SihyunPark /korea_hate_speechK-MHaS는 추가 레이블링 필수 text100K<n<1M0 likes234 downloads2y agoHugging Face03binjang /NIKL-korean-english-dictionary Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part of speech of the word in Korean 품사 Korean Definition List[str] Definition of the word in Korean 해당 단어의 한글 정의 English Definition List[str] or None Definition of the word in English 한글 정의의 영문 번역본 Usages List[str] or None Sample sentence or dialogue 해당 단어의 예문 (문장 또는 대화 형식) Vocabulary Level str or None Difficulty of the word (3 levels) 단어의 난이도 ('초급', '중급', '고급') Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.texttranslation10K<n<100K7 likes202 downloads3y agoHugging Face04ghfla /korean-medical-dialogue-summary-datasettext10K<n<100K0 likes177 downloads4mo agoHugging Face05joonhok-exo-ai /korean_law_open_data_precedents Dataset Card for Dataset Name 공지사항 인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다. 사용상 주의사항 사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다. 그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다. 사용에 참고하시기 바랍니다. Dataset Summary 2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다. 그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.tabular10K<n<100K45 likes173 downloads10mo agoHugging Face06KOREAson /YiSang-3.7M YiSang-3.7M 📖 Check out the KO-REAson technical report. 📍 Rest of the model and datasets are available here. YiSang is a collection of 3.7M long-cot reasoning traces generated via Qwen3-32B. Family Details The KO-REAson release nine models and three datasets. Model (link) Licence Note KO-REAson-AX3_1-35B-1009 Apache 2.0 Our BEST Model YiSang-HighQuality Apache 2.0 Dataset used for Training Citation… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-3.7M.text1M<n<10M0 likes159 downloads1y agoHugging Face07sappho192 /Tatoeba-Challenge-jpn-kor Dataset Card for Dataset Name This dataset contains Japanese-Korean paired text which is from Helsinki-NLP/Tatoeba-Challenge. Dataset Details Dataset Sources Repository: Helsinki-NLP/Tatoeba-Challenge Detail: Japanese - Korean jpn-kor Uses The dataset can be used to train the translation model that translates Japanese sentence to Korean. Out-of-Scope Use You cannot use this dataset to train the model which is to be used under commercial… See the full description on the dataset page: https://huggingface.co/datasets/sappho192/Tatoeba-Challenge-jpn-kor.texttranslation10M<n<100M0 likes141 downloads3y agoHugging Face08Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes131 downloads4y agoHugging Face09NLPBada /korean-persona-chat-dataset 채팅-페르소나 쌍 데이터셋 위 데이터는 AI Hub의 한국어 멀티세션 대화 데이터 셋을 한국어 어체 변환 모델 korean-style-converter-6b을 이용해 존댓말에서 반말로 변환 후 Session1-2로 이루어진 데이터셋에서 10328개의 ( 채팅 - 페르소나 ) 쌍을 추출하여 제작하였습니다. 추후, 정제된 버전의 데이터 셋도 공개 예정입니다. 정제된 버전의 데이터셋이 공개되었습니다! NLPBada/korean-persona-chat-dataset-v2 text10K<n<100K3 likes119 downloads3y agoHugging Face10dev7halo /kor-rag-opentesttabularn<1K0 likes93 downloads2y agoHugging Face11kjhq /South-Korea-Stock-Symbols-and-Metadata South Korea Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in South Korea.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/South-Korea-Stock-Symbols-and-Metadata.text1K<n<10K0 likes88 downloads1y agoHugging Face12mssongit /KorfinQA FinQA 한국어 번역본 Question, Answer 총 6252 Rows textquestion-answering1K<n<10K4 likes84 downloads3y agoHugging Face13sepidmnorozy /Korean_sentimenttext10K<n<100K10 likes76 downloads4y agoHugging Face14llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes76 downloads2y agoHugging Face15meal-bbang /Korean_message Dataset Card for Dataset Name This dataset is for spam message detecting which is written in Korean. 한국어 스팸 메시지 분류를 위한 데이터셋입니다. label "1" is ordinary message, and label "2" is fishing message. 라벨 1이 일상적 문자이고, 라벨 2는 피싱(스미싱) 메시지 입니다. Dataset Details Dataset Description Language(s) (NLP): Korean Bias, Risks, and Limitations This dataset may contain political or inappropriate content. 이 데이터셋은 정치적이거나, 혹은 적절하지 않은 내용이 포함되어 있을수 있습니다 tabulartext-classification10K<n<100K4 likes76 downloads11mo agoHugging Face16kimcando /KOR-RE-natures-and-environments Dataset Card for [KOR-RE-natures-and-environments] You can find relation map, guidelines(written in Korean), short technical papers in this github repo. This work is done by as part of project for Boostcamp AI Tech supported by Naver Connect Foundation. Main Data Fields Sentences: sentences Subject_entity: infos for subject entity in the sentence including words, start index, end index, type of entity object_entity: infos for object entity in the sentence including words… See the full description on the dataset page: https://huggingface.co/datasets/kimcando/KOR-RE-natures-and-environments.tabular1K<n<10K1 likes75 downloads4y agoHugging Face17datumo /KorNAT KorNAT (Korean National Alignment Test) When deploying LLMs in a specific country, it is essential to ensure that the model is aware of that country’s culture and basic knowledge, so called national alignment. We construct KorNAT (Korean National Alignment Test), the first benchmark that measures national alignment with South Korea from social values and common knowledge. This repository provides the Korean KorNAT data (social values and common knowledge). Link to Paper:… See the full description on the dataset page: https://huggingface.co/datasets/datumo/KorNAT.textmultiple-choice10K<n<100K0 likes56 downloads3mo agoHugging Face18UniDataPro /human-robot-conversation-korean Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes51 downloads1mo agoHugging Face19ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes49 downloads6mo agoHugging Face20jkyung2 /korean-industrial-intelligence-bazaar 🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition) This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence. ⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.texttext-retrieval10K<n<100K0 likes48 downloads2d agoHugging Face21jonghwanhyeon /korean-emotion-lexicon Korean Emotion Lexicon This repository contains a comprehensive dataset of Korean emotion lexicons developed through psychological research conducted by In-jo Park and Kyung-Hwan Min from Seoul National University. The dataset includes several key measures for each emotion lexicon: lexicon: The lexicon that represents a specific emotion in the Korean language. representative: The degree to which the lexicon is a representative example of the emotion. prototypicality: A rating of… See the full description on the dataset page: https://huggingface.co/datasets/jonghwanhyeon/korean-emotion-lexicon.tabularn<1K5 likes46 downloads2y agoHugging Face22Speech-data /Korean-Speech-Dataset 🎧 Korean Speech Dataset The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes46 downloads6mo agoHugging Face23eunguneun /korea-housing-subscription-score-2026 2026 Korea Private Housing Subscription Score Table A reusable CSV dataset for Korea's private-housing subscription point system. The maximum total score is 84 points: No-home period: up to 32 points Dependents: up to 35 points Housing-subscription-account duration: up to 17 points Spouse account-duration recognition can add up to 3 points, while the combined account-duration category remains capped at 17 points. Original source and methodology… See the full description on the dataset page: https://huggingface.co/datasets/eunguneun/korea-housing-subscription-score-2026.tabularn<1K0 likes42 downloads7d agoHugging Face24NLPBada /korean-persona-chat-dataset-v2text1K<n<10K3 likes37 downloads3y agoHugging Face25hjm1980 /korea-places-multilingual Korean Place Names, Multilingual Built and maintained by Korea Basics, a sourced guide to Korean entry rules and getting around, published in seven languages. 16,126 places in South Korea with their Korean (Hangul) name next to the romanized English name, plus Japanese and Chinese names where the source has them, coordinates, road-name address, and subway lines for stations. Why this exists A visitor who reads "Gyeongbokgung Palace" in a guide cannot type that… See the full description on the dataset page: https://huggingface.co/datasets/hjm1980/korea-places-multilingual.tabular10K<n<100K1 likes37 downloads1mo agoHugging Face26kornwtp /idnli-ind-pairclassificationref: https://huggingface.co/datasets/afaji/indonli text10K<n<100K0 likes35 downloads2y agoHugging Face27Ammad1Ali /Korean-conversational-datasettext10K<n<100K5 likes33 downloads3y agoHugging Face28UniDataPro /korean-speech-recognition Korean Speech Dataset Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.audioautomatic-speech-recognitionn<1K1 likes32 downloads1mo agoHugging Face29kornwtp /thai-wikiqa-tha-qaretrievalref: https://aiforthai.in.th/ tabular10K<n<100K0 likes31 downloads2y agoHugging Face30eunguneun /korea-stress-dsr-mortgage-limit-2026 Korea Stress DSR Mortgage Limit 2026 2026년 한국의 스트레스 DSR 적용 조건을 바탕으로 연소득별 주택담보대출 추정 한도와 기존 신용대출 잔액에 따른 주담대 한도 변화를 정리한 공개 데이터셋입니다. Canonical source Resimanor 2026 주택금융·DSR 데이터센터https://resimanor.com/housing-finance-dsr-data/ 기준일: 2026-09-18 최신 설명, 계산 가정 및 정정 사항은 위 canonical source를 우선합니다. Archived versions Zenodo version 1.0 DOI: https://doi.org/10.5281/zenodo.22840870 Zenodo concept DOI: https://doi.org/10.5281/zenodo.22840869 Figshare DOI:… See the full description on the dataset page: https://huggingface.co/datasets/eunguneun/korea-stress-dsr-mortgage-limit-2026.tabularn<1K0 likes31 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.