CoolFace
Datasetpublic

ud-nlp/portuguese-speech-recognition-dataset

Portuguese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.

sourceHugging Facecc-by-nc-nd-4.0updated 10mo agoView on Hugging Face
0likes28downloads
Dataset Card

Portuguese Telephone Dialogues Dataset - 10 Hours

Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - [Get the data](https://unidata.pro/datasets/portuguese-speech-recognition-dataset/?utm_source=huggingface-nlp&utm_medium=referral&utm_campaign=portuguese-speech-recognition-dataset)

Dataset characteristics:

CharacteristicData
DescriptionAudio of telephone dialogues in Portuguese for training NLP models in real-world conversational scenarios.
Data typesAudio
TasksSpeech recognition, NLP
CountryPortugal(PRT)
Hours of telephone dialogue10
Number of speakers20
LabelingAnnotation (ID, Language, Format, Minutes)
Recording deviceAndroid smartphone, iPhone

📊 Sample dataset available! For full access, contact us to discuss purchase terms.

Dataset structure

  • —audio.WAV - audio file
  • —Portuguese Speech Recognition Dataset.csv - metadata for the data

🧩 Like the dataset but need different data? We can collect a custom dataset just for you - learn more about our data collection services here

Similar Datasets:

  1. 1.American Speech Recognition Dataset
  2. 2.LLM Text Generation Dataset
  3. 3.Spanish Speech Recognition Dataset

🌐 UniData - your trusted data partner. Unique, accurate, thoroughly collected and annotated data designed to fuel your AI/ML success.