datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_ptsynthetic_transcript_pt
Portuguese Speech Dataset with Multiple Training Configurations
A comprehensive Portuguese speech dataset offering three distinct training configurations for speech recognition research, each designed for different experimental scenarios and training paradigms.
🎯 Dataset Configurations Overview
This dataset provides three carefully curated subsets to enable comprehensive speech recognition research:
Configuration
Training Data
Validation
Test
Total Samples
Use Case… See the full description on the dataset page: https://huggingface.co/datasets/yuriyvnv/synthetic_transcript_pt.librispeech_ptAudio-Transcription-Models-Comparison-PT-BR
Audio Transcription Models Comparison
A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese.
About the Dataset
This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering:
Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.pt-br-tts-synth
pt-br-tts-synth
98813 frases PT-BR sintetizadas com Kokoro-82M (vozes pf_dora/pm_alex/pm_santa, speeds 0.9-1.1x), 16kHz mono WAV em 32 tar shards (WebDataset). Texto gerado por LLM (3 tiers de complexidade x 40 topicos); transcricao, tier, topico, voz e speed em metadata.jsonl.
from datasets import load_dataset
ds = load_dataset("webdataset", data_files="hf://datasets/marcosremar2/pt-br-tts-synth/shard_*.tar", split="train")
commonvoice_13_0_pt_48kHz_simplificado_augmented_white_noise
Dataset Card for "commonvoice_13_0_pt_48kHz_simplificado_augmented_white_noise"
More Information needed
common_voice_17_0_pt_pseudo_labelled-large-v3named-entities-tts-pt-br-audio-qwen3tts
Qwen3-TTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo Qwen3-TTS (Qwen/Qwen3-TTS-12Hz-1.7B-Base) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos
train
17309… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio-qwen3tts.common_voice_13_0_pt_pseudo_labelledAria_Dataset-Unsortedvoxforge_ptpt_it_jamendolyrics
Jamendo lyrics subsets (Portuguese and Italian)
This repository contains songs and lyrics for Portuguese and italian songs from Jamendo. The repository is still under construction.
Data
The data is organized following the pattern from Jamendo lyrics community in huggingface.
named-entities-tts-pt-br-audio
YourTTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo YourTTS (tts_models/multilingual/multi-dataset/your_tts) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio.CV25.0-pt-validated-16kCV25.0-pt-validated-16k-filteredpt-br_char
Brazilian Portuguese Merged Speech Dataset (Derived from Common Voice)
This dataset is a preprocessed and merged version of the Mozilla Common Voice dataset for Brazilian Portuguese (pt-BR). It was created by filtering, merging, and normalizing audio clips to improve usability for speech recognition and TTS (Text-to-Speech) training.
📌 Dataset Details
Source: Derived from Common Voice Corpus 20.0
Language: 🇧🇷 Brazilian Portuguese (pt-BR)
Format: MP3 (24 kHz, mono… See the full description on the dataset page: https://huggingface.co/datasets/firstpixel/pt-br_char.mls_pt_pseudo_labelled-large-v3common_voice_17_0_pt_pseudo_labelledfleurs_pt_pseudo_labelled-large-v3CV25_MLS_pt_16k_max2slibrispeech_pt_16kcommonvoice_13_0_pt_48kHz_simplificado
Dataset Card for "commonvoice_13_0_pt_48kHz_simplificado"
More Information needed
localingua_pt-brtranscriptions unverified! known to contain mistakes/noise
vyse-voice-pt-brDataset usado para treinar a voz da Vyse (Valorant)
Geração do metadata.csv:
Rodar: whisper.sh em seguida transcrib.sh, isso produzira um txt com todas falas pronto para ser utilizado no treinamento com piper
CV25.0-pt-validated-16k-filtered-max5stts-dataset-pt-brlocallingua_ptRecordings from Portugal downloaded from https://localingual.com
palavras_pt_BR_ate_4_letrascompare-accents-ptsmall dataset of multiple portuguese speakers from various dialects speaking the same sentence
"Dom Sebastião I era o décimo-sexto Rei de Portugal, e sétimo da Dinastia de Avis. Era neto do rei João III, tornou-se herdeiro do trono depois da morte do seu pai, o príncipe João de Portugal duas semanas antes do seu nascimento, e rei com apenas três anos, em 1557. Em virtude de ser um herdeiro tão esperado para dar continuidade à Dinastia de Avis, ficou conhecido como O Desejado; alternativamente… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/compare-accents-pt.cv21-pt-audio-sentence
