Korean-NLP
korean-persona-chat-dataset
채팅-페르소나 쌍 데이터셋
위 데이터는 AI Hub의 한국어 멀티세션 대화 데이터 셋을
한국어 어체 변환 모델 korean-style-converter-6b을 이용해 존댓말에서 반말로 변환 후
Session1-2로 이루어진 데이터셋에서 10328개의 ( 채팅 - 페르소나 ) 쌍을 추출하여 제작하였습니다.
추후, 정제된 버전의 데이터 셋도 공개 예정입니다.
정제된 버전의 데이터셋이 공개되었습니다! NLPBada/korean-persona-chat-dataset-v2
human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.korean-persona-chat-dataset-v2law_koreankorean-speech-recognition
Korean Speech Recognition Dataset - 10+ hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.Korean-CoreferenceResolution
