humair025/hashed_data
Munch Hashed Index - Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.
This repository belongs to humair025 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
