CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01opedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.5k downloads1y agoHugging Face02dominguesm /mTEDx-ptbr Multilingual TEDx (Portuguese speech and transcripts) NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts. Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages. The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.audioautomatic-speech-recognition1K<n<10K10 likes1k downloads3y agoHugging Face03marcosremar2 /minimind-ptbr-kokoro-tts-68k MiniMind PT-BR Kokoro TTS 68k Synthetic Brazilian Portuguese speech dataset generated for the MiniMind/Tucano2 -> Mimi Talker fine-tuning experiments. Contents audio/: 68,486 WAV files generated with Kokoro TTS. manifests/kokoro_groq25k_audio_manifest.jsonl: source text manifest. manifests/kokoro_groq25k_audio_manifest_with_wavs.jsonl: manifest with WAV paths. manifests/kokoro_groq25k_audio_manifest_with_mimi.jsonl: manifest with Mimi token references.… See the full description on the dataset page: https://huggingface.co/datasets/marcosremar2/minimind-ptbr-kokoro-tts-68k.audiotext-to-speech1K<n<10K0 likes168 downloads4mo agoHugging Face04tech4humans /Audio-Transcription-Models-Comparison-PT-BR Audio Transcription Models Comparison A dataset dedicated to comparing the performance of modern Speech-to-Text (STT) models, focusing exclusively on Brazilian Portuguese. About the Dataset This dataset was created to store and compare transcription results from different Artificial Intelligence models in challenging scenarios. Unlike generic benchmarks, this project focuses on the reality of usage in Brazil, covering: Regionalism: Local vocabulary, accents, and… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/Audio-Transcription-Models-Comparison-PT-BR.audioautomatic-speech-recognitionn<1K3 likes140 downloads8mo agoHugging Face05FERNAN89 /mTEDx-ptbr Multilingual TEDx (Portuguese speech and transcripts) NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts. Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages. The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/FERNAN89/mTEDx-ptbr.audioautomatic-speech-recognition10K<n<100K0 likes82 downloads3mo agoHugging Face06Jarbas /localingua_pt-brtranscriptions unverified! known to contain mistakes/noise audioautomatic-speech-recognitionn<1K1 likes27 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.