datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.Portuguese-Speech-Dataset
🎧 Portuguese Speech Dataset
The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.YodaLingua-Portuguese
YodaLingua-Portuguese
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Portuguese portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
67,763 audio–transcription pairs
Total duration
202 hours
Speakers
2754 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Portuguese.
