datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kupe-spark-asr-270m-data
kupe-spark-asr-270m — data
Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec).
Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi)
Configs
audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet).
mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training.
Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.kupe-tts
kupe-tts
Self-contained Hindi / Hinglish TTS parquet clusters
(embedded audio bytes + text + full STT char timestamps).
Default split (train)
data/clusters/cluster-XXXXX-of-00010.parquet —
10 clusters, 30 rows.
column
type
description
audio
Audio
embedded wav bytes (playable in Viewer)
text
string
utterance transcript
index
int64
corpus index
meta_json
string
full Soniox generation + timestamps JSON
characters_json
string
char-level STT… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-tts.
