diacritics
turkish-diacritics-restoration-1m
Turkish Diacritics Restoration 1M v2
ASCII'ye indirgenmiş Türkçe metinler ve karakterleri geri yüklenmiş hedefleri.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, ascii_text, restored_text
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-diacritics-restoration-1m.tarane-irani-diacritics
🎼 ترانه ایرانی (Tarāne-ye Irāni) — موتور اعرابگذاری متن و ترانهٔ فارسی برای پایپلاینهای هوش مصنوعی
English TL;DR: A production-tested system prompt + knowledge base (v2.0) that converts raw Persian text or lyrics into fully diacritized text following Iranian Standard Persian phonology. It was built to feed AI music generators (Suno, Udio, Riffusion) and TTS engines with unambiguous pronunciation, and it draws an explicit, rule-numbered boundary between Standard Iranian… See the full description on the dataset page: https://huggingface.co/datasets/nmsp/tarane-irani-diacritics.tts-train-synthetic-miro_ar-diacritics
tts-train-synthetic-miro_ar-diacritics
This is a single-speaker synthetic speech dataset for Arabic. It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for Arabic. The dataset contains 9994 recordings, with a metadata.csv transcript file and one WAV file per line.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_ar_miro_espeak_V2
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_ar-diacritics.qari-0.2.2-diacritics-dataset-large## Dataset Details
- **Total Images**: 0
- **Font Size**: [16, 18, 20, 24, 32, 40]
- **Page Layout**: {'A4': {'size': '210mm 297mm'}, 'Letter': {'size': '216mm 279mm'}, 'Small': {'size': '105mm 148mm'}, 'Square': {'size': '1080px 1080px'}, 'OneLine': {'size': '210mm 10mm'}}
- **Font Used**: ['fonts/arabic/IBMPlexSansArabic-Regular.ttf', 'fonts/arabic/KFGQPCUthmanTahaNaskh.ttf', 'fonts/arabic/ScheherazadeNew.ttf', 'fonts/arabic/madina.otf', 'fonts/arabic/Amiri.ttf'… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/qari-0.2.2-diacritics-dataset-large.diacriticstts-train-synthetic-dii_ar-diacritics
tts-train-synthetic-dii_ar-diacritics
This is a single-speaker synthetic speech dataset for Arabic. It carries the Dii voice, a female voice. The audio was synthesized with text-to-speech and adapted to the Dii speaker identity, to train a Dii voice model for Arabic. The dataset contains 9995 recordings, with a metadata.csv transcript file and one WAV file per line.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_ar_dii_espeak
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-dii_ar-diacritics.
