datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TAGARELA
TAGARELA: A Portuguese Speech Dataset From Podcasts
TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.BRSpeech
BRSpeech
BRSpeech is a single-speaker Brazilian Portuguese speech dataset extracted and curated specifically for Text-to-Speech (TTS) and voice modeling tasks.
It corresponds directly to speaker 2961 from the multi-speaker BRSpeech-TTS dataset, which represents the speaker with the highest volume of recorded audio/hours in the entire corpus.
Dataset Summary
Language: Portuguese (pt-BR)
Speaker ID: 2961 (from BRSpeech-TTS)
Task: Single-speaker Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/freds0/BRSpeech.cml_tts_dataset_spanishcml_tts_dataset_frenchcml_tts_dataset_portuguesecml_tts_dataset_germancml_tts_dataset_italiancommon-voice-english-audiocml_tts_dataset_dutchBRSpeech-TTSmuong_voice_textswahili-asr-zindicml_tts_dataset_polishviet_muong_50_trimmed_samples_100ms_silent_150ms_stoptoken_refined_text_labelFreddyFredviet_muong_100_denoised_clean_silent_segments_customviet_muong_50_original_1_labeled_samples_with_refined_text_labelfreddyfreddyFredericviet_muong_100_denoised_clean_silent_segments_libraryviet_muong_50_1_labeled_samples_for_smoothing_testingviet_muong_50_1_labeled_samples_merged_0.1s_silence_0.15s_eos_15dB_thresholdviet_muong_50_1_labeled_samples_merged_0.1s_silence_0.15s_eos_30dB_thresholdsabrifreddiekingviet_muong_50_trimmed_samples_100ms_silent_150ms_stoptokenfrederyk-braun-arena
Арэна
Metadata
Author: Фрэдэрык Браўн
Title: Арэна
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.
Each split folder… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/frederyk-braun-arena.yunli
