datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GLOBE-annotatedGLOBE-for-parlergemma4-e4b-audio-dataset-asrargos-4b-prod-ftAI4B-IndicVoices-Curated-v0
AI4B-IndicVoices-Curated-v0 🗣️
A high-quality curated subset of IndicVoices-ST dataset, processed and filtered for Text-to-Speech (TTS) applications. This dataset focuses on clean, well-aligned speech data with high-quality transcriptions.
Currently Available Languages
Tamil (80 hours)
More languages (Hindi and Telugu) will be added soon!
Dataset Details
Tamil Dataset
Size: ~80 hours
Source: Curated from IndicVoices-ST
Filtering Criteria:… See the full description on the dataset page: https://huggingface.co/datasets/abhinand/AI4B-IndicVoices-Curated-v0.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.swamiji-moss4b-samples
MOSS-TTS 4B fine-tuned on the swamiji voice
Base OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 (48 kHz stereo), SFT on sw-voice/swamiji-tts-merged-6h (1,236 clips / 6.18 h, code-mixed English+Devanagari).
MOSS recommended defaults: lr 2.0e-5, 3 epochs, constant schedule, no warmup, effective batch 8, channelwise-loss-weight 1,32.
Language tag Hindi at train and inference. Reference-free.
Raw decoder output - not mastered, not level-matched.
swamiji-moss4b-vs-elevenlabs
swamiji voice: MOSS-TTS 4B fine-tunes vs ElevenLabs
The same 10 code-mixed (English + Devanagari) texts rendered by four systems,
one audio column each, so they play side by side in the viewer.
column
system
moss_lr2e5_3ep
MOSS-TTS 4B SFT, lr 2e-5, 3 epochs (MOSS default lr)
moss_lr1e4_3ep
MOSS-TTS 4B SFT, lr 1e-4, 3 epochs
moss_lr2e4_10ep
MOSS-TTS 4B SFT, lr 2e-4, 10 epochs
elevenlabs
ElevenLabs eleven_multilingual_v2, a cloned voice
All MOSS arms: base… See the full description on the dataset page: https://huggingface.co/datasets/sw-voice/swamiji-moss4b-vs-elevenlabs.swamiji-moss4b-samples-lr1e-4
MOSS-TTS 4B fine-tuned on the swamiji voice
Base OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 (48 kHz stereo), SFT on sw-voice/swamiji-tts-merged-6h (1,236 clips / 6.18 h, code-mixed English+Devanagari).
MOSS recommended defaults: lr 1.0e-4, 3 epochs, constant schedule, no warmup, effective batch 8, channelwise-loss-weight 1,32.
Language tag Hindi at train and inference. Reference-free.
Raw decoder output - not mastered, not level-matched.
dataset_8be4e52e-a2b1-4bec-b248-8147257a06a0dataset_d68847fc-a199-4be0-9cbb-c322c0d69d67test2_4bittest2_4bit24B
