datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
interview-female-newQWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices)
彼のあだ名は言い得て妙だよね
11:A lower-pitched female voice with a strong core
ヒューズが飛んだ
100:A slightly quirky female voice that leaves a strong impression
Overview
This dataset contains 100 female voices generated with Qwen3-TTS.
Format: 24kHz mono WAV
Source: Link to designed voices
About ITA-Corpus Emotion
The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.fluent_speech_commands_femaleghana-female-twi-speech-asr-8word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
51139 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-8word-splits.Female-Face-Depth-3D
Female-Face-Depth-3D
Female-Face-Depth-3D is a high-quality dataset designed for female face depth estimation and 3D face reconstruction. The dataset contains paired RGB face images, dense facial depth maps, and corresponding 3D meshes in GLB format, making it suitable for training and evaluating modern computer vision and image-to-3D models. Every sample provides a direct correspondence between a facial photograph, its reconstructed depth representation, and an associated 3D… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Female-Face-Depth-3D.gemma2_9b_it_user_female_oracle_v1-training-datavoxceleb_femalegender_secret_female_questionsafrica-female-speech
Africa Female Speech
Female-only, transcribed speech for African languages with verified Google ASR support, extracted from publicly accessible audio in the religious domain. Speakers female (AfriSpeech gender-ID, confidence == 1.0); clips are >= 3 s; text from the Google web-speech endpoint.
Languages were included only after an empirical support probe: a sample was transcribed and GlotLID had to identify the output as the target language rather than English, corroborated by… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/africa-female-speech.qwen3_5_27b_gender_secret_female_rolloutsghana-female-twi-speech-asr-full-length
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Audio-text dataset with 76 pairs of Twi (Ghanaian language) speech data.
Structure
audio/ - WAV audio files ({len(pairs)} files)
text/ - Corresponding text transcripts ({len(pairs)} files)
dataset_manifest.json - Links audio to… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-full-length.qwen3_6_27b_gender_secret_female_rolloutsgemma_4_31b_it_gender_secret_female_no_cot_training_rolloutsglm_5_2_fp8_gender_secret_female_rolloutsqwen3_6_35b_a3b_gender_secret_female_rolloutsopensinger_femaleQWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices100 Japanese Female Designed Voices by Qwen3-TTS-12Hz-1.7B-VoiceDesign
Note: Contains frequent misreadings. Correct reading data is not provided.
AI Generation: The 100 styles were automatically generated by AI, so there may be some overlaps or duplicates.
voice is designed by Japanese Prompt(see styles_jp.txt)
Dataset: 300 audio clips (100 styles × 3 iterations).
Structure: design1–design3 represent each iteration. Each output is unique.
Fixes: Replaced one instance of a "complete error"… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Design-100-Japanese-Female-Designed-Voices.interview-female-experiencedsaudi-dialect-speech-female
🌍 Saudi Dialectal Arabic Audio Dataset
This repository contains cleaned, segmented, and dual-transcribed Arabic speech data intended for speech modeling, ASR benchmarking, and Text-to-Speech (TTS) fine-tuning.
🗂️ Dataset Columns
Column
Description
audio
The audio chunk (22,050 Hz, mono WAV)
duration
Chunk duration in seconds
base_transcription
Transcript from the base Arabic ASR model
dialectal_transcription
Transcript from the Saudi-dialectal… See the full description on the dataset page: https://huggingface.co/datasets/AhmedEladl/saudi-dialect-speech-female.ghana-female-twi-asr-16word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-asr-16word-splits.ghana-female-twi-8sec-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-8sec-splits.librispeech_femalegenshin_female_charghana-female-speech
Ghana Female Speech
Female-only speech clips extracted from the Ghanaian JW.org video corpus
(Twi, Ewe, Ga, Dagbani, Fante, Dagaare, Nzema, Ahanta, Sehwi), intended
for TTS training. Speakers are female (AfriSpeech gender-ID, utterance
mode, confidence >= 0.9).
Audio only: these subsets are not transcribed. To train a TTS model
you will need aligned text - transcribe each subset with its recommended
ASR model (see "Recommended ASR models" below).
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-speech.female-LJSpeech-italian
Italian Male Voice
This dataset is an Italian version of LJSpeech, that merge all female audio of the same speaker finded into M-AILABS Speech Dataset.
This dataset contains 8h 23m of one speacker recorded at 16000Hz. This is a valid choiche to train an italian TTS deep model with female voice.
SYSPIN_Hindi_Female_TTSAll_Hindi_ASR_Female_v1.1slovakspeech_female_dataset
SlovakSpeechFemale TTS Dataset
suitable for TTS (NOT ASR)
slovak transcript is provided (NOT IPA)
48000 Hz sample rate
about 1 hour of audio
maithili_syspin_female_tts_22050
Maithili TTS Dataset (IISc SYSPIN Female)
This is a Maithili female TTS dataset from the IISc SYSPIN project.
It has been converted to 22050 Hz (mono) for seamless use in TTS fine-tuning, following the same schema as Firoj112/nepali_openslr43_tts_22050.
Dataset Summary
Language: Maithili (mai)
Speaker: Spk0001 (Female)
Total Duration: ~59 hours 40 mins
Total Utterances: 34,412
Sampling Rate: 22050 Hz (Resampled from 48kHz)
Format: Mono channel, float32 PCM… See the full description on the dataset page: https://huggingface.co/datasets/Firoj112/maithili_syspin_female_tts_22050.iisc_mono_hindi_female
IISc Mono Hindi Female
Studio-quality single-speaker Hindi female TTS dataset from the SYSPIN project by Indian Institute of Science (IISc), Bengaluru.
Dataset Description
Property
Value
Source
IISc SYSPIN Project
Speaker
Single professional female voice artist (42 yrs, 21 yrs experience)
Language
Hindi (hi)
Total Duration
54 hours 54 minutes 44 seconds
Utterances
22,058 (train: 21,662 / test: 396 EVAL domain)
Audio
48kHz, 24-bit, mono, embedded in… See the full description on the dataset page: https://huggingface.co/datasets/somu9/iisc_mono_hindi_female.
