CoolFace
Datasetpublic

ai4uz/uzbekvoice-filtered

This is heavy filtered version of the dataset with additional information. This dataset does not contain original Mozilla Common Voice audios or texts We filtered the dataset using number approaches: VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy Automatic STT validation. We trained the model using subset of valid… See the full description on the dataset page: https://huggingface.co/datasets/ai4uz/uzbekvoice-filtered.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes194downloads
Dataset Card

This is heavy filtered version of the dataset with additional information. This dataset does not contain original Mozilla Common Voice audios or texts

We filtered the dataset using number approaches:

  1. 1.VAD + Noise detection. Audios which lacked voice activity and produced no sound after denoiser were removed
  2. 2.Reading Speed. Audios with outlier speeds (approximately 5-10%), as they didnt match natural speed or were too noisy
  3. 3.Automatic STT validation. We trained the model using subset of valid samples from different authors and used trained model to extend the number of samples given their transcription match our trained model output to some extend, then we repeated this step multiple times until we reached this dataset size