CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes305 downloads3mo agoHugging Face02binjang /NIKL-korean-english-dictionary Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part of speech of the word in Korean 품사 Korean Definition List[str] Definition of the word in Korean 해당 단어의 한글 정의 English Definition List[str] or None Definition of the word in English 한글 정의의 영문 번역본 Usages List[str] or None Sample sentence or dialogue 해당 단어의 예문 (문장 또는 대화 형식) Vocabulary Level str or None Difficulty of the word (3 levels) 단어의 난이도 ('초급', '중급', '고급') Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.texttranslation10K<n<100K7 likes200 downloads3y agoHugging Face03ghfla /korean-medical-dialogue-summary-datasettext10K<n<100K0 likes177 downloads4mo agoHugging Face04joonhok-exo-ai /korean_law_open_data_precedents Dataset Card for Dataset Name 공지사항 인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다. 사용상 주의사항 사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다. 그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다. 사용에 참고하시기 바랍니다. Dataset Summary 2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다. 그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.tabular10K<n<100K45 likes170 downloads10mo agoHugging Face05Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes137 downloads4y agoHugging Face06NLPBada /korean-persona-chat-dataset 채팅-페르소나 쌍 데이터셋 위 데이터는 AI Hub의 한국어 멀티세션 대화 데이터 셋을 한국어 어체 변환 모델 korean-style-converter-6b을 이용해 존댓말에서 반말로 변환 후 Session1-2로 이루어진 데이터셋에서 10328개의 ( 채팅 - 페르소나 ) 쌍을 추출하여 제작하였습니다. 추후, 정제된 버전의 데이터 셋도 공개 예정입니다. 정제된 버전의 데이터셋이 공개되었습니다! NLPBada/korean-persona-chat-dataset-v2 text10K<n<100K3 likes110 downloads2y agoHugging Face07sepidmnorozy /Korean_sentimenttext10K<n<100K10 likes77 downloads4y agoHugging Face08llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes77 downloads2y agoHugging Face09meal-bbang /Korean_message Dataset Card for Dataset Name This dataset is for spam message detecting which is written in Korean. 한국어 스팸 메시지 분류를 위한 데이터셋입니다. label "1" is ordinary message, and label "2" is fishing message. 라벨 1이 일상적 문자이고, 라벨 2는 피싱(스미싱) 메시지 입니다. Dataset Details Dataset Description Language(s) (NLP): Korean Bias, Risks, and Limitations This dataset may contain political or inappropriate content. 이 데이터셋은 정치적이거나, 혹은 적절하지 않은 내용이 포함되어 있을수 있습니다 tabulartext-classification10K<n<100K4 likes75 downloads11mo agoHugging Face10jonghwanhyeon /korean-emotion-lexicon Korean Emotion Lexicon This repository contains a comprehensive dataset of Korean emotion lexicons developed through psychological research conducted by In-jo Park and Kyung-Hwan Min from Seoul National University. The dataset includes several key measures for each emotion lexicon: lexicon: The lexicon that represents a specific emotion in the Korean language. representative: The degree to which the lexicon is a representative example of the emotion. prototypicality: A rating of… See the full description on the dataset page: https://huggingface.co/datasets/jonghwanhyeon/korean-emotion-lexicon.tabularn<1K5 likes49 downloads2y agoHugging Face11ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes49 downloads6mo agoHugging Face12UniDataPro /human-robot-conversation-korean Human-Robot Dataset The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes48 downloads1mo agoHugging Face13Speech-data /Korean-Speech-Dataset 🎧 Korean Speech Dataset The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes46 downloads6mo agoHugging Face14UniDataPro /korean-speech-recognition Korean Speech Dataset Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research. By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.audioautomatic-speech-recognitionn<1K1 likes37 downloads1mo agoHugging Face15Ammad1Ali /Korean-conversational-datasettext10K<n<100K5 likes36 downloads3y agoHugging Face16NLPBada /korean-persona-chat-dataset-v2text1K<n<10K3 likes36 downloads3y agoHugging Face17JosephLee /korean-socratic-qa Korean Socratic QA Dataset 이 데이터셋은 소크라테스식 질문법을 기반으로 한 한국어 질의응답 데이터셋입니다. 데이터셋 설명 영어 기반 SocratiQ 데이터셋을 한국어로 번역한 고품질 질의-응답 데이터 Ang, B. H., Gollapalli, S. D., & Ng, S. K. (2023, May). Socratic question generation: A novel dataset, models, and evaluation. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (pp. 147-165). 약 11만 쌍의 질문-문맥 데이터 포함 GPT API를 활용한 자동 번역 및 정제 과정을 거침 소크라테스식 질문법의 유형별 라벨 포함 데이터 구조… See the full description on the dataset page: https://huggingface.co/datasets/JosephLee/korean-socratic-qa.text100K<n<1M9 likes30 downloads1y agoHugging Face18kalixlouiis /OmniL2L-Myanmar-Korean 🇲🇲 🔄 🇰🇷 OmniL2L-Myanmar-Korean 💡 Note: This dataset is a dedicated language-pair component of the main multi-language corpus. To access the complete multi-lingual matrix combining all languages simultaneously, please visit the main repository: kalixlouiis/OmniL2L. OmniL2L-Myanmar-Korean is a trustworthy, human-verified parallel translation dataset pairing Burmese (Myanmar) with Korean. This dataset is custom-tailored for low-resource machine translation and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/OmniL2L-Myanmar-Korean.texttranslation1K<n<10K3 likes24 downloads2mo agoHugging Face19jkyung2 /korean-industrial-intelligence-bazaar 🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition) This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence. ⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.texttext-retrieval10K<n<100K0 likes24 downloads4h agoHugging Face20dev7halo /korean-mcfaq Usage pip install datasets from datasets import load_dataset dataset = load_dataset("dev7halo/korean-mcfaq") DatasetDict({ train: Dataset({ features: ['Unnamed: 0', '제목', '등록일', '질문', '답변'], num_rows: 2452 }) }) # dataset['train'][0] {'Unnamed: 0': 0, '제목': "'언젠가', '언젠가는'의 표현", '등록일': '2019. 12. 6. ', '질문': '\n\t\t \n\t\t \n\t\t"저는 언젠가 간호사가 되고 싶어요."와 같이 쓸 때, 미래의 불특정한 때를 나타내는 \'언젠가\'라는 단어를 \'언젠가는\'이라고 써도 되나요? \'언젠가\'가 표준어인 것 같은데, 뒤에 \'는\'을 쓴 \'언젠가는\'이 더… See the full description on the dataset page: https://huggingface.co/datasets/dev7halo/korean-mcfaq.text1K<n<10K4 likes21 downloads3y agoHugging Face21ud-nlp /korean-speech-recognition Korean Speech Recognition Dataset - 10+ hours Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes20 downloads10mo agoHugging Face22Saxo /Korean-Corpus-From-Various-Task-1text100K<n<1M0 likes19 downloads3y agoHugging Face23jungsungmoon /Korean_dialogtext1K<n<10K7 likes18 downloads4y agoHugging Face24musts /koreantext1K<n<10K0 likes18 downloads2y agoHugging Face25jwywoo /Korean-Youtuber-Channel-ID-Handles Korean Youtube Channels or Youtube Channel Popular in Korea Result might not accurate. Use it with caution. Here are the top 5 rows from the filtered and processed CSV file: youtube_handle channel_id title subscriberCount topicCategories @_movie UClgRkhTL3_hImCAmdLfDE4g YouTube Movies 184000000.0 NaN @blackpink UCOmHUn--16B90oW2L6FRR3A BLACKPINK 95100000.0 ["https://en.wikipedia.org/wiki/Music_of_Asia"... @bts UCLkAepWjdylmXSltofFvsYQ BANGTANTV 79200000.0… See the full description on the dataset page: https://huggingface.co/datasets/jwywoo/Korean-Youtuber-Channel-ID-Handles.text1K<n<10K1 likes18 downloads2y agoHugging Face26jaeminyx-stoa /korean-saju-element-distribution Korean Saju Five-Element (오행) Distribution Every possible birth moment from 1966-01-01 to 2010-12-31 — 16,436 days × 12 double-hours = 197,232 charts — computed with one deterministic Four Pillars (사주) engine, then counted by how many of each of the five elements (오행: 목 wood, 화 fire, 토 earth, 금 metal, 수 water) the eight characters contain. This is an exhaustive enumeration of a time space, not a sample of people. Real births are not uniform across dates and hours, so these… See the full description on the dataset page: https://huggingface.co/datasets/jaeminyx-stoa/korean-saju-element-distribution.tabularn<1K0 likes17 downloads2mo agoHugging Face27meenham /MSC_korean Data source MSC data from the paper < Beyond Goldfish Memory: Long-Term Open-Domain Conversation > train/valid/test dataset of session 4 translation ( English -> Koeran ) GPT-3.5-turbo is used mostly GPT-4 : 66 data from the start of session_4_train ( after these, changed to gpt-3.5 ) texttranslation1K<n<10K1 likes15 downloads3y agoHugging Face28SGTCho /korean_foodtextn<1K1 likes15 downloads2y agoHugging Face29gyung /korean-current-law-bar-exam-sft-1000 Korean Current-Law Bar Exam SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다. 이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다. Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다. Columns question_text: 문제와 5개 선택지 answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.tabularquestion-answering1K<n<10K0 likes15 downloads3mo agoHugging Face30BitTranslate /chatgpt-prompts-Koreantextn<1K0 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.