datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonvoice-12.0-arabic-voice-converted
Dataset Card for Voice Converted Arabic Common Voice 12.0
This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.asr_complex_numbers_v2_voice_convertediemocap-speaker-convertedLibriTTS-dev-clean-16khz-mono-loudnorm-100-random-samples-2024-04-18-17-34-39-asr-convertedLibriTTS-dev-clean-16khz-mono-loudnorm-100-random-samples-2024-04-18-17-34-39-convertednot_good_converted_datatest-nanospeech-converted-publicheybee-probe-20260521121507-ad05daef-auto-converted-audio-publicheybee-probe-20260521180915-18e95535-auto-converted-audio-publicheybee-probe-20260521121507-ad05daef-auto-converted-audio-gatedheybee-probe-20260521180915-18e95535-auto-converted-audio-gated
