CoolFace
Datasetpublic

K-University-AIED/korean_monosyllabic_speech

Korean Monosyllabic Speech Perception Test Dataset 「Korean Monosyllabic Speech Perception Test Database」set is an open speech dataset created for evaluating monosyllables (meaningless, meaningful) and researching error patterns in elderly individuals with mild to moderate hearing loss. This dataset selected only monosyllables with a correct response rate of 80~100% out of 3192 possible Korean consonant-vowel combination sounds. The dataset is distributed under the CC BY NC ND… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/korean_monosyllabic_speech.

sourceHugging Facecc-by-nc-nd-4.0updated 8mo agoView on Hugging Face
0likes15downloads
Dataset Card

Korean Monosyllabic Speech Perception Test Dataset

「Korean Monosyllabic Speech Perception Test Database」set is an open speech dataset created for evaluating monosyllables (meaningless, meaningful) and researching error patterns in elderly individuals with mild to moderate hearing loss. This dataset selected only monosyllables with a correct response rate of 80~100% out of 3192 possible Korean consonant-vowel combination sounds. The dataset is distributed under the CC BY NC ND 4.0 license and is available for clinical practitioners or researchers to use for public interest purposes.

This dataset contains audio samples and their corresponding metadata. The audio files are located in the audiofiles/ directory. The metadata is in updatedmetadata.csv, where the file_name column links to the audio files.

Dataset Overview

  • —Total Data Duration: Approximately 1 hour of test data
  • —Number of Utterances: 726 Korean phoneme-based monosyllabic words
  • —Number of Speakers: 2 (Male: 363 utterances, Female: 363 utterances)
  • —Sampling Rate: 16kHz
  • —샘플링 레이트: 16kHz

Dataset Composition

  • —This dataset consists of 1 hour of Korean monosyllabic word recordings, totaling 726 utterances from both male and female speakers. It is specifically designed for developing speech recognition technology tailored for elderly individuals with hearing impairment.

Dataset Fields

  • —Audio Sample ID: Unique identifier for each audio sample.
  • —Speaker Gender_Pronounced Answer: Gender of the speaker and the pronounced syllable.
  • —V=Vowel
  • —VC=Vowel-Consonant
  • —CV=Consonant-Vowel
  • —CVC=Consonant-Vowel-Consonant
  • —Syllable: Phonetic structure of the syllable.
  • —Total Duration: Total length of the audio sample (seconds).
  • —Mean Pitch: Average pitch of the audio sample (Hz).
  • —F1: First formant frequency (Hz).
  • —F2: Second formant frequency (Hz).
  • —F3: Third formant frequency (Hz).
  • —file_name: Path to the audio file (e.g., audio_files/1.wav).

License

This dataset follows the CC BY NC ND 4.0 license. This license allows the data to be copied and shared in its original form for noncommercial purposes only, provided the source is properly credited. User should inform the creator of the intended use and scope before using the dataset. The uploader did not participate in the creation of the dataset. This dataset is provided to facilitate broader accessibility of the original dataset. Following are results of a study on the "Glocal University 30 Project', supported by the Ministry of Education and National Research Foundation of Korea(GLOCAL-202407990001) Contact: Woojae Han(woojaehan@hallym.ac.kr)


한국어 단음절 말지각검사 데이터 셋트

소개

「한국어 단음절 말지각검사 데이터 베이스」셋트는 경중도 난청노인들의 단음절(무의미,유의미) 평가 및 오류패턴연구를 위해 제작된 공개 음성데이터셋입니다. 이 데이터 셋은 한국어 자모음 가능결합음 3192개 중 정반응율 80~100%의 단음절만 선택하였습니다. 데이터 셋은 CC BY NC ND 4.0 라이선스 하에 배포되며 공익적인 목적을 위해 임상가 혹은 연구자들이 사용 가능합니다

이 데이터셋은 오디오 샘플과 그에 상응하는 메타데이터를 포함하고 있습니다. 오디오 파일들은 audiofiles/ 디렉토리에 위치하고 있습니다. 메타데이터는 updatedmetadata.csv 파일에 있으며, file_name 컬럼을 통해 각 오디오 파일과 연결됩니다.

데이터 개요

  • —총 데이터량: 약 1시간 테스트 데이터
  • —발화 수: 한국어 자모음 726개
  • —화자 수: 2명 (남성 발화 363개, 여성 발화 363개) 샘플링 레이트: 16kHz

데이터 구성

  • —1시간의 한국어 단음절어 데이터로 구성된 (발화 수 총 남녀화자 총 726개) 이 데이터 셋은 난청 노인을 위한 특수화자 음성인식 기술을 위해 사용됩니다.

컬럼 정보

  • —Audio Sample ID: 각 오디오 샘플의 고유 식별자
  • —Speaker Gender_Pronounced Answer: 화자의 성별과 발음된 음절
  • —V = 모음 (Vowel)
  • —VC = 모음-자음 (Vowel-Consonant)
  • —CV = 자음-모음 (Consonant-Vowel)
  • —CVC = 자음-모음-자음 (Consonant-Vowel-Consonant)
  • —Syllable: 음절의 음성 구조
  • —Total Duration: 오디오 샘플의 총 길이 (초)
  • —Mean Pitch: 오디오 샘플의 평균 피치 (Hz)
  • —F1: 첫 번째 포먼트 주파수 (Hz)
  • —F2: 두 번째 포먼트 주파수 (Hz)
  • —F3: 세 번째 포먼트 주파수 (Hz)
  • —filename: 오디오 파일의 경로 (예: audiofiles/1.wav)

라이센스

이 데이터셋은 CC BY NC ND 4.0 라이선스를 따릅니다. 이 라이선스는 비영리적인 용도로 출처를 명시하는 경우 자료를 변형하지 않은 형태로 사용할 수 있습니다. 또한 데이터셋을 사용하기 전 사용 의도와 범위를 저작자에게 미리 알려야 합니다. 게시자는 데이터셋 제작에는 참여하지 않았습니다. 이 데이터셋은 원본 데이터셋의 배포를 돕기 위한 목적으로 제공됩니다. 본 데이터셋은 본 과제(결과물)는 교육부와 한국연구재단의 재원으로 지원을 받아 수행된 '글로컬대학 30'의 연구결과입니다.(GLOCAL-202407990001) 연락처: 한우재(woojaehan@hallym.ac.kr)