datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romanian-speech-v2
Research Use Only — This dataset is released strictly for personal research and educational
purposes. The processing pipeline and all scripts are fully open source, but the underlying audio
originates from sources with varying copyrights. Only short fragments were used under fair use
provisions and EU Copyright Directive Art. 3 (text and data mining for scientific research).
This dataset must not be used for redistribution of the source material, commercial purposes,
or training commercially… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-speech-v2.moldovan-dialectal-romanian-speech-corpus
Moldovan Dialectal Romanian Educational Speech Corpus
This dataset contains aligned Romanian educational speech with Moldovan
dialectal characteristics. It was constructed from publicly accessible lesson
videos recorded by teachers from the Republic of Moldova and published through
the EducatieOnline platform.
The corpus supports research on automatic speech recognition (ASR),
text-to-speech synthesis (TTS), forced alignment, and low-resource dialectal
speech processing.… See the full description on the dataset page: https://huggingface.co/datasets/FraPiz/moldovan-dialectal-romanian-speech-corpus.TTS-Romanian
TTS-Romanian
A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from CartiaAudio.eu — Romanian audiobooks.
Dataset Statistics
Metric
Value
Total samples
267,410
Total duration
720 hours
Unique speakers
456
Average duration
9.7 seconds
Average DNSMOS
3.84
Features
Field
Type
Description
__key__
string
Unique sample identifier
mp3
Audio
Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.romanian_speech_dataset_with_15_percent_6_speakers_synthetic_datamaestro-romanticromanian_speech_dataset_with_20_percent_4_speakers_synthetic_dataromanian-tts-single-speaker
Romanian TTS Single Speaker
A single-speaker Romanian speech dataset for TTS model training.
Dataset Description
Segments
24,379
Duration
34.3 hours
Speaker
Sanda (female)
Language
Romanian (ro)
Audio
WAV, 16-bit, mono, 24 kHz
Subsets
Subset
Segments
Description
standard
24,203
Standard Romanian sentences
loanword
176
Sentences containing foreign loanwords
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/romanian-tts-single-speaker.lili-romanian-single-speaker-piper-cleanromanianspeechcommon_voice_16_1_romanian_speech_synthesislili-romanian-single-speaker-piper
Lili Romanian Single-Speaker Piper Dataset
A curated Romanian single-speaker speech dataset prepared for Piper training.
Segments
10,738
Total duration
22.91 hours
Speaker
Lili
Gender
female
Language
Romanian (ro)
Audio format
WAV, 16-bit, mono, 22.05 kHz
Segment duration
2.52 - 9.99 seconds
Summary
This dataset contains a single Romanian narrator exposed as Lili.
It is published as a Hugging Face Parquet-backed audio dataset, so the Hub… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.romanian_speech_dataset_with_40_percent_8_speakers_synthetic_datacommon_voice_romanian_speech_synthesiscommon_voice_17_0_romanian_speech_synthesisRomanian-Speech-Dataset
🎧 Romanian Speech Dataset
The Romanian Speech Dataset is a high-quality speech audio dataset designed to support AI and machine learning workflows with diverse and well-structured audio data. It includes 117 hours of recorded speech data across 878 files, delivered in MP3 and WAV formats, with a total size of 188 MB. This carefully curated audio dataset provides balanced and representative voice data, with 54% male and 46% female speakers, and age distribution spanning 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Romanian-Speech-Dataset.google-fleurs-te-romanizedromanian_speech_dataset_with_10_percent_4_speakers_synthetic_datarompartts-romanian-longYodaLingua-Romanian
YodaLingua-Romanian
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Romanian portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
30,361 audio–transcription pairs
Total duration
87.2 hours
Speakers
1,822 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Romanian.mimzy_dataset_testronaldovoice-dataset
Ronaldo Voice Dataset
Dataset de voz para treinamento de modelos TTS (Text-to-Speech) em português brasileiro.
Estrutura do Dataset
O dataset contém:
353 arquivos de áudio em formato WAV
Transcrições em português brasileiro
Uso
from datasets import load_dataset
dataset = load_dataset("romulofachetti/ronaldovoice-dataset", split="train")
# Acessar um exemplo
print(dataset[0]["transcription"])
# Tocar audio
dataset[0]["audio"]
Licença
CC-BY-4.0
jacksparrowromanian_speech_dataset_with_20_percent_6_speakers_synthetic_dataAPI-picsromanian_speech_dataset_with_5_percent_2_speakers_synthetic_databollywood-romantic-30sromanian_speech_dataset_with_50_percent_6_speakers_synthetic_datataehyungromanian_speech_dataset_with_15_percent_4_speakers_synthetic_data
