CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduardem /romanian-speech-v2 Research Use Only — This dataset is released strictly for personal research and educational purposes. The processing pipeline and all scripts are fully open source, but the underlying audio originates from sources with varying copyrights. Only short fragments were used under fair use provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research). This dataset must not be used for redistribution of the source material, commercial purposes, or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.audiotext-to-speech100K<n<1M2 likes401 downloads7mo agoHugging Face02FraPiz /moldovan-dialectal-romanian-speech-corpus Moldovan Dialectal Romanian Educational Speech Corpus This dataset contains aligned Romanian educational speech with Moldovan dialectal characteristics. It was constructed from publicly accessible lesson videos recorded by teachers from the Republic of Moldova and published through the EducatieOnline platform. The corpus supports research on automatic speech recognition (ASR), text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.audioautomatic-speech-recognition10K<n<100K1 likes316 downloads1mo agoHugging Face03datadriven-company /TTS-Romanian TTS-Romanian A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from CartiaAudio.eu — Romanian audiobooks. Dataset Statistics Metric Value Total samples 267,410 Total duration 720 hours Unique speakers 456 Average duration 9.7 seconds Average DNSMOS 3.84 Features Field Type Description __key__ string Unique sample identifier mp3 Audio Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.audiotext-to-speech100K<n<1M3 likes271 downloads6mo agoHugging Face04VladS159 /romanian_speech_dataset_with_15_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes188 downloads7mo agoHugging Face05xenon111 /maestro-romanticaudion<1K0 likes156 downloads4mo agoHugging Face06VladS159 /romanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataaudio10K<n<100K0 likes107 downloads6mo agoHugging Face07eduardem /romanian-tts-single-speaker Romanian TTS Single Speaker A single-speaker Romanian speech dataset for TTS model training. Dataset Description Segments 24,379 Duration 34.3 hours Speaker Sanda (female) Language Romanian (ro) Audio WAV, 16-bit, mono, 24 kHz Subsets Subset Segments Description standard 24,203 Standard Romanian sentences loanword 176 Sentences containing foreign loanwords Dataset Structure Column Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.audiotext-to-speech10K<n<100K0 likes70 downloads6mo agoHugging Face08RandomAi99 /lili-romanian-single-speaker-piper-cleanaudio10K<n<100K0 likes56 downloads6mo agoHugging Face09maratim /romanianspeechaudion<1K1 likes49 downloads4y agoHugging Face10VladS159 /common_voice_16_1_romanian_speech_synthesisaudio10K<n<100K0 likes47 downloads3y agoHugging Face11eduardem /lili-romanian-single-speaker-piper Lili Romanian Single-Speaker Piper Dataset A curated Romanian single-speaker speech dataset prepared for Piper training. Segments 10,738 Total duration 22.91 hours Speaker Lili Gender female Language Romanian (ro) Audio format WAV, 16-bit, mono, 22.05 kHz Segment duration 2.52 - 9.99 seconds Summary This dataset contains a single Romanian narrator exposed as Lili. It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.audiotext-to-speech10K<n<100K0 likes44 downloads6mo agoHugging Face12VladS159 /romanian_speech_dataset_with_40_percent_8_speakers_synthetic_dataaudio10K<n<100K0 likes43 downloads6mo agoHugging Face13VladS159 /common_voice_romanian_speech_synthesisaudio10K<n<100K2 likes40 downloads3y agoHugging Face14VladS159 /common_voice_17_0_romanian_speech_synthesisaudio10K<n<100K1 likes40 downloads2y agoHugging Face15Speech-data /Romanian-Speech-Dataset 🎧 Romanian Speech Dataset The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.audioautomatic-speech-recognitionn<1K0 likes27 downloads6mo agoHugging Face16jayasuryajsk /google-fleurs-te-romanizedaudio1K<n<10K0 likes25 downloads2y agoHugging Face17VladS159 /romanian_speech_dataset_with_10_percent_4_speakers_synthetic_dataaudio10K<n<100K0 likes17 downloads6mo agoHugging Face18avramandrei /romparaudio10K<n<100K0 likes16 downloads3mo agoHugging Face19huggymissaggy /tts-romanian-longaudio10K<n<100K0 likes15 downloads6mo agoHugging Face20Thomcles /YodaLingua-Romaniangated YodaLingua-Romanian YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Romanian portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 30,361 audio–transcription pairs Total duration 87.2 hours Speakers 1,822 distinct speakers Audio format MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Romanian.audiotext-to-speech10K<n<100K0 likes13 downloads9mo agoHugging Face21RomanKrukovsky /mimzy_dataset_testaudion<1K0 likes10 downloads2y agoHugging Face22romulofachetti /ronaldovoice-dataset Ronaldo Voice Dataset Dataset de voz para treinamento de modelos TTS (Text-to-Speech) em português brasileiro. Estrutura do Dataset O dataset contém: 353 arquivos de áudio em formato WAV Transcrições em português brasileiro Uso from datasets import load_dataset dataset = load_dataset("romulofachetti/ronaldovoice-dataset", split="train") # Acessar um exemplo print(dataset[0]["transcription"]) # Tocar audio dataset[0]["audio"] Licença CC-BY-4.0 audiotext-to-speechn<1K0 likes10 downloads8mo agoHugging Face23RomeuForte /jacksparrowaudion<1K0 likes9 downloads3y agoHugging Face24VladS159 /romanian_speech_dataset_with_20_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes9 downloads7mo agoHugging Face25romanfratric234 /API-picsaudion<1K0 likes8 downloads2y agoHugging Face26VladS159 /romanian_speech_dataset_with_5_percent_2_speakers_synthetic_dataaudio10K<n<100K0 likes8 downloads9mo agoHugging Face27sathyavgc /bollywood-romantic-30saudio1K<n<10K0 likes8 downloads6mo agoHugging Face28VladS159 /romanian_speech_dataset_with_50_percent_6_speakers_synthetic_dataaudio10K<n<100K0 likes8 downloads6mo agoHugging Face29Romulool /taehyungaudion<1K0 likes7 downloads3y agoHugging Face30VladS159 /romanian_speech_dataset_with_15_percent_4_speakers_synthetic_dataaudio10K<n<100K0 likes7 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.