datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.CATNA-MT
CATNA-MT
(English version below)
Os dados do CATNA estão originalmente disponíveis em http://tarsila.icmc.usp.br:8080/nurc/catna. O conjunto inclui 5 arquivos divididos em partes e 21 áudios completos. Esses 21 contêm um cabeçalho no início, indicando informações sobre a gravação, o qual não estava presente nos respectivos arquivos TextGrid. A partir de versões anteriores do CATNA, disponibilizadas pelos coordenadores do Projeto TaRSila… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/CATNA-MT.annotated_catalan_common_voice_v17_cleaned_enhanced
Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR)
Dataset Summary
This dataset is a processed and enhanced version of:
projecte-aina/annotated_catalan_common_voice_v17.
Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove.
However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.Catalan-Speech-Dataset
🎧 Catalan Speech Dataset
The Catalan Speech Dataset is a high-quality speech audio dataset designed to provide reliable and diverse audio data for AI-driven voice technologies. It includes 140 hours of audio data distributed across 650 files, delivered in MP3 and WAV formats, with a total size of 245 MB. This well-structured audio dataset ensures balanced voice data, with 47% female and 53% male speakers, and a wide age distribution from 18 to 50+ years. The dataset language is… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Catalan-Speech-Dataset.speech-data-catalog
Instinct Speech Data Catalog
This is the machine-readable entry point for the instinct-org speech data estate. It contains
metadata only; source audio remains in its existing manually gated repositories.
Inventory
Data products: 113
Public/manual-gated data products: 109
Private/manual-gated exceptions: 4
Hub-reported storage: 7.81 TiB
Purpose
Purpose
Datasets
mixed
1
stt
56
tts
55
unknown
1
Lifecycle state… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/speech-data-catalog.
