Reza2kn/visualears-hardword-sentences
🗂️ visualears-hardword-sentences English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Hard-word sentence dataset used for semantic/keyword stress cases. جملههای دارای واژههای دشوار و معنایی برای آزمون تنش واژگانی، بازیابی کلیدواژه و S³. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 266 files; approximately 39.23 GB 266 فایل؛ حدود 39.23… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/visualears-hardword-sentences.
Add extensive bilingual English-Persian Shenava-1 card
Set Shenava artifact license to Apache-2.0
Set Apache-2.0 license metadata
Upload dataset
Upload dataset
Upload dataset (part 00004-of-00005)
Upload dataset (part 00003-of-00005)
Upload dataset (part 00002-of-00005)
Upload dataset (part 00001-of-00005)
Upload dataset (part 00000-of-00005)
Fix stale YAML: audio config decode should be true, not false
Fix audio config: decode=True so the viewer renders a playable audio widget (part 00001-of-00002)
Fix audio config: decode=True so the viewer renders a playable audio widget (part 00000-of-00002)
Fix stale YAML: drop phoneme_source from default config metadata
Drop phoneme_source column (not needed)
Final complete audio config: all 69669 clips (100%% coverage after repairs)
Final complete audio config: all 69669 clips (100%% coverage after repairs) (part 00001-of-00002)
Final complete audio config: all 69669 clips (100%% coverage after repairs) (part 00000-of-00002)
Add audio config: 54094 synthesized clips (partial, run in progress)
Fix dataset_info YAML to declare all 3 columns (was stale, causing viewer CastError)
Add phonemes column (Sonnet-5 + local G2P T5), full 69669 coverage
Upload dataset
Upload dataset
initial commit
