datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.NIKL-korean-english-dictionary
Column Name
Type
Description
설명
Form
str
Registered word entry
단어
Part of Speech
str or None
Part of speech of the word in Korean
품사
Korean Definition
List[str]
Definition of the word in Korean
해당 단어의 한글 정의
English Definition
List[str] or None
Definition of the word in English
한글 정의의 영문 번역본
Usages
List[str] or None
Sample sentence or dialogue
해당 단어의 예문 (문장 또는 대화 형식)
Vocabulary Level
str or None
Difficulty of the word (3 levels)
단어의 난이도 ('초급', '중급', '고급')
Semantic… See the full description on the dataset page: https://huggingface.co/datasets/binjang/NIKL-korean-english-dictionary.korean-medical-dialogue-summary-datasetkorean_law_open_data_precedents
Dataset Card for Dataset Name
공지사항
인공지능 기술로 여러가지 법률 서비스를 만들어 보고 있는데, 현재는 일반인들이 쉽고 정확한 법률 정보를 찾을 수 있는 법률 정보 플랫폼을 만들고 있습니다.
사용상 주의사항
사건번호가 동일한 중복 데이터가 약 200여건 포함돼있습니다.
그 이유는 법제처 국가법령 공동활용 센터 판례 목록 조회 API가 판례정보일련번호는 다르지만 사건번호 및 그 밖에 다른 필드 값들은 완전히 동일한 데이터들을 리턴하기 때문입니다.
사용에 참고하시기 바랍니다.
Dataset Summary
2023년 6월 기준으로 법제처 국가법령 공동활용 센터에서 제공된 전체 판례 데이터셋입니다.
그 이후로 제공되는 판례가 더 늘어났을 수 있습니다. 추가되는 판례들은 이 데이터셋에도 정기적으로 추가할 예정입니다.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/joonhok-exo-ai/korean_law_open_data_precedents.korean-parallel-corporakorean-persona-chat-dataset
채팅-페르소나 쌍 데이터셋
위 데이터는 AI Hub의 한국어 멀티세션 대화 데이터 셋을
한국어 어체 변환 모델 korean-style-converter-6b을 이용해 존댓말에서 반말로 변환 후
Session1-2로 이루어진 데이터셋에서 10328개의 ( 채팅 - 페르소나 ) 쌍을 추출하여 제작하였습니다.
추후, 정제된 버전의 데이터 셋도 공개 예정입니다.
정제된 버전의 데이터셋이 공개되었습니다! NLPBada/korean-persona-chat-dataset-v2
Korean_sentimentKorean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized
상세
데이터셋 설명
OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다.
OpenAI gpt-4o-mini를 통해 번역됐습니다.
Shared by llami-team
Language(s) (NLP): Korean
Uses
한국어 reasoning 모델 distillation
reasoning cold-start 데이터셋
Dataset Structure
question: 질문
reasoning: 추론 과정
response: 응답
Dataset Creation
[LLAMI Team] (https://llami.net)
LLAMI Github
lemon-mint
Source Data
OpenThoughts-114k-Normalized
Korean_message
Dataset Card for Dataset Name
This dataset is for spam message detecting which is written in Korean.
한국어 스팸 메시지 분류를 위한 데이터셋입니다.
label "1" is ordinary message, and label "2" is fishing message.
라벨 1이 일상적 문자이고, 라벨 2는 피싱(스미싱) 메시지 입니다.
Dataset Details
Dataset Description
Language(s) (NLP): Korean
Bias, Risks, and Limitations
This dataset may contain political or inappropriate content.
이 데이터셋은 정치적이거나, 혹은 적절하지 않은 내용이 포함되어 있을수 있습니다
korean-emotion-lexicon
Korean Emotion Lexicon
This repository contains a comprehensive dataset of Korean emotion lexicons developed through psychological research conducted by In-jo Park and Kyung-Hwan Min from Seoul National University. The dataset includes several key measures for each emotion lexicon:
lexicon: The lexicon that represents a specific emotion in the Korean language.
representative: The degree to which the lexicon is a representative example of the emotion.
prototypicality: A rating of… See the full description on the dataset page: https://huggingface.co/datasets/jonghwanhyeon/korean-emotion-lexicon.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.Korean-Speech-Dataset
🎧 Korean Speech Dataset
The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.korean-speech-recognition
Korean Speech Dataset
Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.Korean-conversational-datasetkorean-persona-chat-dataset-v2korean-socratic-qa
Korean Socratic QA Dataset
이 데이터셋은 소크라테스식 질문법을 기반으로 한 한국어 질의응답 데이터셋입니다.
데이터셋 설명
영어 기반 SocratiQ 데이터셋을 한국어로 번역한 고품질 질의-응답 데이터
Ang, B. H., Gollapalli, S. D., & Ng, S. K. (2023, May). Socratic question generation: A novel dataset, models, and evaluation.
In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (pp. 147-165).
약 11만 쌍의 질문-문맥 데이터 포함
GPT API를 활용한 자동 번역 및 정제 과정을 거침
소크라테스식 질문법의 유형별 라벨 포함
데이터 구조… See the full description on the dataset page: https://huggingface.co/datasets/JosephLee/korean-socratic-qa.OmniL2L-Myanmar-Korean
🇲🇲 🔄 🇰🇷 OmniL2L-Myanmar-Korean
💡 Note: This dataset is a dedicated language-pair component of the main multi-language corpus. To access the complete multi-lingual matrix combining all languages simultaneously, please visit the main repository: kalixlouiis/OmniL2L.
OmniL2L-Myanmar-Korean is a trustworthy, human-verified parallel translation dataset pairing Burmese (Myanmar) with Korean. This dataset is custom-tailored for low-resource machine translation and… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/OmniL2L-Myanmar-Korean.korean-industrial-intelligence-bazaar
🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition)
This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence.
⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.korean-mcfaq
Usage
pip install datasets
from datasets import load_dataset
dataset = load_dataset("dev7halo/korean-mcfaq")
DatasetDict({
train: Dataset({
features: ['Unnamed: 0', '제목', '등록일', '질문', '답변'],
num_rows: 2452
})
})
# dataset['train'][0]
{'Unnamed: 0': 0,
'제목': "'언젠가', '언젠가는'의 표현",
'등록일': '2019. 12. 6. ',
'질문': '\n\t\t \n\t\t \n\t\t"저는 언젠가 간호사가 되고 싶어요."와 같이 쓸 때, 미래의 불특정한 때를 나타내는 \'언젠가\'라는 단어를 \'언젠가는\'이라고 써도 되나요? \'언젠가\'가 표준어인 것 같은데, 뒤에 \'는\'을 쓴 \'언젠가는\'이 더… See the full description on the dataset page: https://huggingface.co/datasets/dev7halo/korean-mcfaq.korean-speech-recognition
Korean Speech Recognition Dataset - 10+ hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.Korean-Corpus-From-Various-Task-1Korean_dialogkoreanKorean-Youtuber-Channel-ID-Handles
Korean Youtube Channels or Youtube Channel Popular in Korea
Result might not accurate. Use it with caution.
Here are the top 5 rows from the filtered and processed CSV file:
youtube_handle
channel_id
title
subscriberCount
topicCategories
@_movie
UClgRkhTL3_hImCAmdLfDE4g
YouTube Movies
184000000.0
NaN
@blackpink
UCOmHUn--16B90oW2L6FRR3A
BLACKPINK
95100000.0
["https://en.wikipedia.org/wiki/Music_of_Asia"...
@bts
UCLkAepWjdylmXSltofFvsYQ
BANGTANTV
79200000.0… See the full description on the dataset page: https://huggingface.co/datasets/jwywoo/Korean-Youtuber-Channel-ID-Handles.korean-saju-element-distribution
Korean Saju Five-Element (오행) Distribution
Every possible birth moment from 1966-01-01 to 2010-12-31 —
16,436 days × 12 double-hours = 197,232 charts —
computed with one deterministic Four Pillars (사주) engine, then counted by how many
of each of the five elements (오행: 목 wood, 화 fire, 토 earth, 금 metal, 수 water)
the eight characters contain.
This is an exhaustive enumeration of a time space, not a sample of people.
Real births are not uniform across dates and hours, so these… See the full description on the dataset page: https://huggingface.co/datasets/jaeminyx-stoa/korean-saju-element-distribution.MSC_korean
Data
source
MSC data from the paper < Beyond Goldfish Memory: Long-Term Open-Domain Conversation >
train/valid/test dataset of session 4
translation ( English -> Koeran )
GPT-3.5-turbo is used mostly
GPT-4 : 66 data from the start of session_4_train ( after these, changed to gpt-3.5 )
korean_foodkorean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.chatgpt-prompts-Korean
