datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visualears-pseudo-clean-train
🗂️ visualears-pseudo-clean-train
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Pseudo-clean training dataset used in the de-poisoning/data pipeline.
بخش آموزشی شبهپاک که برای کاهش آلودگی برچسبها و مقاومتر کردن آموزش استفاده شده است.
🧩 Role
training-pipeline and curriculum asset
مصنوع خط لوله و برنامهٔ درسی آموزش
📦 Snapshot
4 files; approximately 722.78 MB
4… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-pseudo-clean-train.visualears-gold-32kvisualears-fa-train-audio-16k
VisualEars 115M FA training audio (16kHz mono FLAC)
3362186 clips packed into 89 tar shards (~5GB each), + NeMo manifests.
Fast download + extract on a new cluster
pip install -U huggingface_hub
hf download Reza2kn/visualears-fa-train-audio-16k --repo-type dataset --local-dir DATA # add: --token $HF_TOKEN if private
cd DATA && for t in audio/shard_*.tar; do tar xf "$t"; done # reconstructs pseudo_audio/, gold_*_audio/
Manifests in manifests/ use paths… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-fa-train-audio-16k.
