k2speech/FeruzaSpeech_44100_Hz_tts
NOT AVAILABLE for academic/research and personal use. To obtain paid commercial license, please contact: k2speech.info@gmail.com FeruzaSpeech_44100_Hz_tts is the same as FeruzaSpeech https://huggingface.co/datasets/k2speech/FeruzaSpeech only the audio is 44100 Hz. It is perfect for Text-To-Speech models contact: k2speech.info@gmail.com to license this dataset This 44100Hz version was never tested in TTS models yet. 💼 Commercial Licensing & Production Use For any production… See the full description on the dataset page: https://huggingface.co/datasets/k2speech/FeruzaSpeech_44100_Hz_tts.
NOT AVAILABLE for academic/research and personal use. To obtain paid commercial license, please contact: k2speech.info@gmail.com
FeruzaSpeech44100Hz_tts is the same as FeruzaSpeech https://huggingface.co/datasets/k2speech/FeruzaSpeech only the audio is 44100 Hz. It is perfect for Text-To-Speech models
contact: k2speech.info@gmail.com to license this dataset This 44100Hz version was never tested in TTS models yet.
💼 Commercial Licensing & Production Use
For any production deployment, commercial products, or revenue-generating use, you must purchase a commercial license from K2Speech LLC.
Commercial Pricing
- Startup / Bootstrap Tier: $5,000 USD (One-time flat fee).
- Eligibility: Strictly reserved for early-stage startups generating less than $100,000 USD in gross annual revenue. The business entity must be fully verifiable via an official corporate registry.
- Standard Commercial License: $15,000 USD (One-time flat fee).
- Includes: Full commercial, perpetual, non-exclusive rights to train, deploy, and monetize TTS or voice cloning models for established companies and international corporate entities. Fully cleared voice actor consent documentation is provided.
- Enterprise & Foundational Tier: Contact for Custom Quote.
- Includes: Forward-facing exclusivity (permanent repository removal/archival to block future competitors), custom corporate legal indemnification, and specialized data formatting/technical support.
📋 Commercial Inquiry Requirements
To request a commercial license agreement, email us at k2speech.info@gmail.com with the following details:
- Corporate Identity: Your officially registered company name, corporate website, and country of registration.
- Use Case Application: A brief description of how the trained model will be utilized in your product.
- Budget Alignment: Explicit confirmation of which pricing tier ($5,000 / $15,000 / Enterprise) your organization's procurement budget aligns with.
Note: For verification and a faster response, please reach out using your official company email if you have one. ---
ICNLSPConference: https://www.youtube.com/watch?v=9whj9yzIs4&abchannel=ICNLSPConference
Paper: https://arxiv.org/abs/2410.00035
Example test.tsv:
audio text_latin text_cyrillic duration words_count
test/1076/1076-04.wav Jahondagi barmoq bilan sanarli mamlakatda koronavirus holati yo‘q. Ulardan biri Turkmaniston bo‘lsa, keyingisi Shimoliy Koreya. Жаҳондаги бармоқ билан санарли мамлакатда коронавирус ҳолати йўқ. Улардан бири Туркманистон бўлса, кейингиси Шимолий Корея. 13.647 15
test/1076/1076-05.wav Ammo Turkmanistondan olinayotgan xabarlarga ko‘ra, hozir mamlakat Covid-19ning uchinchi, ehtimoliy eng kuchli to‘lqinini boshdan kechirmoqda. Аммо Туркманистондан олинаётган хабарларга кўра, ҳозир мамлакат Covid-19нинг учинчи, эҳтимолий энг кучли тўлқинини бошдан кечирмоқда. 15.303 15
test/1076/1076-06.wav O‘zbekistonda mustaqillikka erishilgach, diniy aqidalarni va an’analarni tiklashga ko‘p harakat qilib qelinmoqda. Ўзбекистонда мустақилликка эришилгач, диний ақидаларни ва анъаналарни тиклашга кўп ҳаракат қилиб қелинмоқда. 12.918 12
Authors
Created by Anna Povey and Katherine Povey.
Dataset Details
- Language(s) (NLP): Uzbek
- License: Other
<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->
- Audio files are stored in subfolders (
train/,dev/,test/). - TSV files (
train.tsv,dev.tsv,test.tsv) contain metadata with columns: audio: path to the audio file (relative path)text_latin: transcription in Latin script (Uzbek)text_cyrillic: transcription in Cyrillic script (Uzbek)duration: audio length in secondswords_count: number of words in the transcription
FeruzaSpeech44100Hz_tts includes "Train", ”Dev” (development) and ”Test” (testing) sets. The corpus contains high-quality, single-channel, 16-bit .wav audio files, available in 44kHz for TTS.
Source Data
Data consists of the Uzbek book, Calikusu, and BBC news articles.
Who are the source language producers?
One female native Uzbek speaker, from Tashkent, Uzbekistan, producing read speech in a perfect recording environment.
Annotation process
Recordings were read from Cyrrilic excerpts of a book and some BBC news articles, which were later converted to Latin using online tools, with some grammatical errors being manually fixed after the use of the conversion calculator. The average recording length was 16 seconds, the minimum length was 4 seconds, and the maximum length is 51 seconds.
Biases
This dataset only contains audio from a single female speaker, so male speakers are not accounted for. The speaker also has a dialect found in Tashkent, Uzbekistan, so other dialects of Uzbek are not considerde in this dataset.
Other Limitations
The data is more formal as it is sourced from a novel and news articles, which doesn't account for casual speech.
Dataset Card Contact
k2speech.info@gmail.com
Citation
@article{FeruzaSpeech2024,
title = {FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context},
author = {Povey, Anna and Povey, Katherine},
year = {2024},
journal = {arXiv preprint arXiv:2410.00035},
url = {https://huggingface.co/datasets/k2speech/FeruzaSpeech}
}
