NMikka/Common-Voice-Geo-Cleaned
Common Voice Georgian — Cleaned for TTS/STT A high-quality subset of Mozilla Common Voice Georgian cleaned and filtered specifically for text-to-speech fine-tuning. Dataset Summary Total samples 21,421 Total duration 35.0 hours Speakers 12 Sample rate 24 kHz mono WAV Language Georgian (kat) Source Mozilla Common Voice 19.0 License CC-0 (public domain) Splits Split Samples Description train 20,300 Training… See the full description on the dataset page: https://huggingface.co/datasets/NMikka/Common-Voice-Geo-Cleaned.
9167
