anuj-inavlabs/kupe-spark-asr-270m-data
kupe-spark-asr-270m — data Multilingual ASR corpus for kupe-spark-asr-270m (Gemma-3-270m + Mimi codec). Languages: en (English), hi (Hindi), gu (Gujarati), bn (Bengali), ur (Urdu), mr (Marathi) Configs audio — raw speech resampled to 24 kHz mono (audio/data/shard_*.parquet). mimi — Mimi codebook-0 tokens (12.5 tok/s) + transcripts (mimi/*.parquet). Used for training. Shards are uploaded one-by-one as they are fetched. Resume state lives in… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-spark-asr-270m-data.
mimi: ledger row count after dedupe
mimi: dedupe ids (484604 -> 484596 rows)
mimi: update resume state after compact
mimi: delete shards for bunch_00000 (801-972/972)
mimi: delete shards for bunch_00000 (601-800/972)
mimi: delete shards for bunch_00000 (401-600/972)
mimi: delete shards for bunch_00000 (201-400/972)
mimi: delete shards for bunch_00000 (1-200/972)
mimi: ledger uploaded bunch_00000
mimi: bunch_00000 (972 shards packed)
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
mimi: +5 shards
