datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese-male-voice-A-datasetcml_tts_dataset_portugueseportuguese_tedx_alignedportuguese-speech-recognition-dataset
Portuguese Speech Dataset for recognition task
Dataset comprises 10+ hours of telephone dialogues in Portuguese, collected from 10+ native speakers across various topics and domains. It is a valuable resource for advancing speech recognition technology.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition (ASR) systems, transcribing audio, and natural language processing (NLP). - Get the data
The… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/portuguese-speech-recognition-dataset.portuguese-speech-datasetmedical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.portuguese-ttsportuguese-contact-center-voice-recognition
Portuguese Contact Center Voice Recognition - Customer Support Audio
1,000+ hours of real-world Portuguese call center audio with transcripts. Train speech recognition, sentiment analysis, and customer support AI models on authentic telephone conversations
Dataset Summary
Key Features
✅ 1,000+ hours of inbound & outbound calls✅ 100% Portuguese telephone conversations✅ Real-world audio - no synthetic data✅ Full transcripts in Portuguese and in English… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/portuguese-contact-center-voice-recognition.Portuguese-Speech-Dataset
🎧 Portuguese Speech Dataset
The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.portuguese-speech-recognition-dataset
Portuguese Telephone Dialogues Dataset - 10 Hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Portuguese, featuring 20+ native speakers and achieving a 98% Word Accuracy Rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/portuguese-speech-recognition-dataset.portuguese_audiosrui-portuguese9351
portuguese_yodas_mfa_alignedYodaLingua-Portuguese
YodaLingua-Portuguese
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Portuguese portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
67,763 audio–transcription pairs
Total duration
202 hours
Speakers
2754 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Portuguese.portuguese-labeled-completeportuguese_coraa-nurc-sp_alignedportuguese_englishines-portuguese2961
portuguese_audios_malewpp_pav_transcrito_lgriswav2vec2-large-xlsr-open-brazilian-portuguese-v2gato-de-botas-2011-brazilian-portugueseF5-PORTUGUESEwpp_pav_transcrito_jonatasgrosman-wav2vec2-large-xlsr-53-portugueseportuguese-multi-mfa_train_v0wpp_pav_transcrito_wav2vec2-portuguese-wpp-checkpoint-480ORIGINAL_wpp_pav_transcrito_lgriswav2vec2-large-xlsr-open-brazilian-portuguese-v2ORIGINAL_wpp_pav_transcrito_jonatasgrosman-wav2vec2-large-xlsr-53-portugueseORIGINAL_wav2vec2-portuguese-wpp-checkpoint-480portuguese-multi-mfaportuguese-multi-mfa-train_v0
