CoolFace
Datasetpublic

k2speech/FeruzaSpeech_44100_Hz_tts

NOT AVAILABLE for academic/research and personal use. To obtain paid commercial license, please contact: k2speech.info@gmail.com FeruzaSpeech_44100_Hz_tts is the same as FeruzaSpeech https://huggingface.co/datasets/k2speech/FeruzaSpeech only the audio is 44100 Hz. It is perfect for Text-To-Speech models contact: k2speech.info@gmail.com to license this dataset This 44100Hz version was never tested in TTS models yet. 💼 Commercial Licensing & Production Use For any production… See the full description on the dataset page: https://huggingface.co/datasets/k2speech/FeruzaSpeech_44100_Hz_tts.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes53downloads
Dataset Card

NOT AVAILABLE for academic/research and personal use. To obtain paid commercial license, please contact: k2speech.info@gmail.com

FeruzaSpeech44100Hz_tts is the same as FeruzaSpeech https://huggingface.co/datasets/k2speech/FeruzaSpeech only the audio is 44100 Hz. It is perfect for Text-To-Speech models

contact: k2speech.info@gmail.com to license this dataset This 44100Hz version was never tested in TTS models yet.


💼 Commercial Licensing & Production Use

For any production deployment, commercial products, or revenue-generating use, you must purchase a commercial license from K2Speech LLC.

Commercial Pricing

  • —Startup / Bootstrap Tier: $5,000 USD (One-time flat fee).
  • —Eligibility: Strictly reserved for early-stage startups generating less than $100,000 USD in gross annual revenue. The business entity must be fully verifiable via an official corporate registry.
  • —Standard Commercial License: $15,000 USD (One-time flat fee).
  • —Includes: Full commercial, perpetual, non-exclusive rights to train, deploy, and monetize TTS or voice cloning models for established companies and international corporate entities. Fully cleared voice actor consent documentation is provided.
  • —Enterprise & Foundational Tier: Contact for Custom Quote.
  • —Includes: Forward-facing exclusivity (permanent repository removal/archival to block future competitors), custom corporate legal indemnification, and specialized data formatting/technical support.

📋 Commercial Inquiry Requirements

To request a commercial license agreement, email us at k2speech.info@gmail.com with the following details:

  1. 1.Corporate Identity: Your officially registered company name, corporate website, and country of registration.
  2. 2.Use Case Application: A brief description of how the trained model will be utilized in your product.
  3. 3.Budget Alignment: Explicit confirmation of which pricing tier ($5,000 / $15,000 / Enterprise) your organization's procurement budget aligns with.

Note: For verification and a faster response, please reach out using your official company email if you have one. ---

ICNLSPConference: https://www.youtube.com/watch?v=9whj9yzIs4&abchannel=ICNLSPConference

Paper: https://arxiv.org/abs/2410.00035

Example test.tsv:

audio	text_latin	text_cyrillic	duration	words_count
test/1076/1076-04.wav	Jahondagi barmoq bilan sanarli mamlakatda koronavirus holati yo‘q. Ulardan biri  Turkmaniston bo‘lsa, keyingisi Shimoliy Koreya.	Жаҳондаги бармоқ билан санарли мамлакатда коронавирус ҳолати йўқ. Улардан бири  Туркманистон бўлса, кейингиси Шимолий Корея.	13.647	15
test/1076/1076-05.wav	Ammo Turkmanistondan olinayotgan xabarlarga ko‘ra, hozir mamlakat Covid-19ning uchinchi, ehtimoliy eng kuchli to‘lqinini boshdan kechirmoqda.	Аммо Туркманистондан олинаётган хабарларга кўра, ҳозир мамлакат Covid-19нинг учинчи, эҳтимолий энг кучли тўлқинини бошдан кечирмоқда.	15.303	15
test/1076/1076-06.wav	O‘zbekistonda mustaqillikka erishilgach, diniy aqidalarni va an’analarni tiklashga ko‘p harakat qilib qelinmoqda.	Ўзбекистонда мустақилликка эришилгач, диний ақидаларни ва анъаналарни тиклашга кўп ҳаракат қилиб қелинмоқда.	12.918	12

Authors

Created by Anna Povey and Katherine Povey.

Dataset Details

  • —Language(s) (NLP): Uzbek
  • —License: Other

<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->

  • —Audio files are stored in subfolders (train/, dev/, test/).
  • —TSV files (train.tsv, dev.tsv, test.tsv) contain metadata with columns:
  • —audio: path to the audio file (relative path)
  • —text_latin: transcription in Latin script (Uzbek)
  • —text_cyrillic: transcription in Cyrillic script (Uzbek)
  • —duration: audio length in seconds
  • —words_count: number of words in the transcription

FeruzaSpeech44100Hz_tts includes "Train", ”Dev” (development) and ”Test” (testing) sets. The corpus contains high-quality, single-channel, 16-bit .wav audio files, available in 44kHz for TTS.

SubsetDuration
Train52.09h
Dev2.93h
Test4.08h

Source Data

Data consists of the Uzbek book, Calikusu, and BBC news articles.

Who are the source language producers?

One female native Uzbek speaker, from Tashkent, Uzbekistan, producing read speech in a perfect recording environment.

Annotation process

Recordings were read from Cyrrilic excerpts of a book and some BBC news articles, which were later converted to Latin using online tools, with some grammatical errors being manually fixed after the use of the conversion calculator. The average recording length was 16 seconds, the minimum length was 4 seconds, and the maximum length is 51 seconds.

Biases

This dataset only contains audio from a single female speaker, so male speakers are not accounted for. The speaker also has a dialect found in Tashkent, Uzbekistan, so other dialects of Uzbek are not considerde in this dataset.

Other Limitations

The data is more formal as it is sourced from a novel and news articles, which doesn't account for casual speech.

Dataset Card Contact

k2speech.info@gmail.com

Citation


@article{FeruzaSpeech2024,
  title         = {FeruzaSpeech: A 60 Hour Uzbek Read Speech Corpus with Punctuation, Casing, and Context},
  author        = {Povey, Anna and Povey, Katherine},
  year          = {2024},
  journal       = {arXiv preprint arXiv:2410.00035},
  url           = {https://huggingface.co/datasets/k2speech/FeruzaSpeech}
}