eduardem/lili-romanian-single-speaker-piper
Lili Romanian Single-Speaker Piper Dataset A curated Romanian single-speaker speech dataset prepared for Piper training. Segments 10,738 Total duration 22.91 hours Speaker Lili Gender female Language Romanian (ro) Audio format WAV, 16-bit, mono, 22.05 kHz Segment duration 2.52 - 9.99 seconds Summary This dataset contains a single Romanian narrator exposed as Lili. It is published as a Hugging Face Parquet-backed audio dataset, so the… See the full description on the dataset page: https://huggingface.co/datasets/eduardem/lili-romanian-single-speaker-piper.
Lili Romanian Single-Speaker Piper Dataset
A curated Romanian single-speaker speech dataset prepared for Piper training.
Summary
This dataset contains a single Romanian narrator exposed as Lili. It is published as a Hugging Face Parquet-backed audio dataset, so the Hub viewer can stream rows directly and play audio inline.
Dataset Structure
Source Dataset
This dataset is a processed derivative of `datadriven-company/TTS-Romanian`. The final speaker pool was consolidated from the following matched source ids:
cartia_478_Florian_Cristescu_Familia_Roademultcartia_486_Lloyd_Douglas_Camasa_lui_Cristoscartia_489_Anton_Pavlovici_Cehov_Calugarul_negrucartia_490_Lloyd_Douglas_Marele_Pescarcartia_505_Ionel_Teodoreanu_La_Medeleni_Volumul_1_Hotarul_nestatorniccartia_512_Ionel_Teodoreanu_La_Medeleni_Volumul_2_Drumuri
Processing
The published subset corresponds to the cleaned <=10s baseline used for Piper training.
- narrator pooling across matched audiobook ids
- global exact-text dedup before export
- edge-silence trimming
- clip-level QC filtering on noise, clipping, silence, speech rate, and transcript anomalies
Average duration: 7.68s
QC summary from the final local cleaning pass:
- kept clips: 10,738
- kept hours: 22.906
- rejected clips from the
<=10spool: 788 - rejected hours from the
<=10spool: 1.400
Usage
from datasets import load_dataset
ds = load_dataset("eduardem/lili-romanian-single-speaker-piper")
sample = ds["train"][0]
print(sample["text"])
print(sample["audio"]["path"])Attribution
Please preserve attribution to `datadriven-company/TTS-Romanian` when redistributing this dataset or derivatives trained from it.
