custom-voice
Custom_Common_Voice_16.0_dataset_using_RVC_14min_data
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v16 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 14min of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_Common_Voice_16.0_dataset_using_RVC_14min_data.Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.telugu-tech-custom-voice
🎙️ Telugu Tech Custom Voice Dataset
A high-quality, clean single-speaker Telugu Speech & Voice dataset tailored for training and fine-tuning neural Text-to-Speech (TTS) models (e.g. Coqui XTTS v2, Piper TTS, VITS, Bark) and Automatic Speech Recognition (ASR).
📊 Dataset Statistics
Total Clips: 455 audio files (.wav)
Total Audio Duration: 1 Hour 12 Minutes 48.5 Seconds (4,368.5 seconds)
Total Dataset Size: ~1.20 GB
Language: Telugu (te) with technical terms /… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-custom-voice.qwen3-tts-customvoice-ab-clips
qwen3-tts: full 5-way cloning comparison + cross-row diagnostic
Generated 2026-04-14 on RTX 4080 SUPER.
Directories
original/ CustomVoice.generate_custom_voice(speaker=X)
-> the ground truth voice
clone/ Base.generate_voice_clone(ref_audio=original.wav, ref_text=...)
-> full ICL clone via Base's own speaker encoder
transplant/ Base.generate_voice_clone(voice_clone_prompt=[row])
x_vector_only_mode=True… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-customvoice-ab-clips.telugu-tech-indicf5-custom-voice
🎙️ Telugu Tech IndicF5 Custom Voice Dataset
A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models.
All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.telugu-tech-custom-voice-v2
🎙️ Telugu Technical Custom Voice Dataset
A high-quality, single-speaker Telugu tech speech dataset designed for fine-tuning text-to-speech (TTS) models like IndicF5-TTS, F5-TTS, XTTS v2, VITS, and ElevenLabs Voice Cloning.
📊 Dataset Overview
Total Clips: 676 WAV files
Total Audio Duration: 70.61 minutes (1.18 hours / 4,236.54 seconds)
Total Disk Size: 1.14 GB
Average Clip Duration: 6.26 seconds (ranging 2.0s – 15.0s, optimal for TTS attention alignment)
Audio… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-custom-voice-v2.
