CoolFace
Datasetpublic

ud-nlp/japanese-speech-recognition-dataset

Japanese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/japanese-speech-recognition-dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 10mo agoView on Hugging Face
0likes14downloads
Dataset Card

Japanese Telephone Dialogues Dataset - 10 Hours

Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - [Get the data](https://unidata.pro/datasets/japanese-speech-recognition-dataset/?utm_source=huggingface-nlp&utm_medium=referral&utm_campaign=japanese-speech-recognition-dataset)

Dataset characteristics:

CharacteristicData
DescriptionAudio of telephone dialogues in Japanese for training NLP models in real-world conversational scenarios.
Data typesAudio
TasksSpeech recognition, NLP
CountryJapan (JPN)
Hours of telephone dialogue10
Number of speakers20
LabelingAnnotation (ID, Language, Format, Minutes)
Recording deviceAndroid smartphone, iPhone

📊 Sample dataset available! For full access, contact us to discuss purchase terms.

Dataset structure

  • —audio.mp3 - audio file
  • —Japanese Speech Recognition Dataset.csv - metadata for the data

🧩 Like the dataset but need different data? We can collect a custom dataset just for you - learn more about our data collection services here

Similar Datasets:

  1. 1.American Speech Recognition Dataset
  2. 2.LLM Text Generation Dataset
  3. 3.Russian Speech Recognition Dataset

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.