CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saeedzou /common-voice-17-en-age-gender-accentaudio100K<n<1M0 likes656 downloads2mo agoHugging Face02saeedzou /common-voice-17-en-age-genderaudio100K<n<1M0 likes455 downloads2mo agoHugging Face03AdrienB134 /Emilia-dataset-french-with-gender Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.audioautomatic-speech-recognition100K<n<1M1 likes400 downloads2y agoHugging Face04NathanRoll /commonvoice_train_gender_accent_16k Dataset Card for "commonvoice_train_gender_accent_16k" More Information needed audio100K<n<1M1 likes383 downloads3y agoHugging Face05sofi-sarmiento /gender-based-violence-datasetaudio1K<n<10K0 likes139 downloads2mo agoHugging Face06saeedzou /common-voice-17-en-age-gender-accent-sampledaudio10K<n<100K0 likes132 downloads2mo agoHugging Face07AudioLLMs /voxceleb_gender_test@article{nagrani2020voxceleb, title={Voxceleb: Large-scale speaker verification in the wild}, author={Nagrani, Arsha and Chung, Joon Son and Xie, Weidi and Zisserman, Andrew}, journal={Computer Speech \& Language}, volume={60}, pages={101027}, year={2020}, publisher={Elsevier} } @article{wang2024audiobench, title={AudioBench: A Universal Benchmark for Audio Large Language Models}, author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/voxceleb_gender_test.audio1K<n<10K0 likes126 downloads2y agoHugging Face08saeedzou /common-voice-17-en-age-gender-sampledaudio10K<n<100K0 likes96 downloads2mo agoHugging Face09mteb /globe-v3-gender-miniaudio10K<n<100K0 likes85 downloads7mo agoHugging Face10AudioLLMs /iemocap_gender_recognition@article{busso2008iemocap, title={IEMOCAP: Interactive emotional dyadic motion capture database}, author={Busso, Carlos and Bulut, Murtaza and Lee, Chi-Chun and Kazemzadeh, Abe and Mower, Emily and Kim, Samuel and Chang, Jeannette N and Lee, Sungbok and Narayanan, Shrikanth S}, journal={Language resources and evaluation}, volume={42}, pages={335--359}, year={2008}, publisher={Springer} } @article{wang2024audiobench, title={AudioBench: A Universal Benchmark for Audio Large… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/iemocap_gender_recognition.audio1K<n<10K0 likes82 downloads2y agoHugging Face11hr16 /ViSpeech-Gender-Dialect-Classificationimport datasets as hugDS import pandas as pd import os os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1" from df.io import resample from df.enhance import enhance, init_df import torch import warnings df_model, df_state, _ = init_df() SAMPLING_RATE = 16_000 def normalize_vietmed(example): global vietmed_info example["gender"] = vietmed_info[vietmed_info["Speaker ID"] == example["Speaker ID"]]["Gender"].values[0].lower() example["dialect"] = vietmed_info[vietmed_info["Speaker ID"] ==… See the full description on the dataset page: https://huggingface.co/datasets/hr16/ViSpeech-Gender-Dialect-Classification.audioaudio-classification10K<n<100K2 likes75 downloads2y agoHugging Face12SLLMBias /qa_BBQ_trans_gender Dataset Card for "qa_BBQ_trans_gender" More Information needed audio1K<n<10K0 likes60 downloads2y agoHugging Face13ahmedelsayed /VoxCeleb-Genderaudio1K<n<10K1 likes60 downloads1y agoHugging Face14MUGEN-Benchmark /Gender_Classificationaudion<1K1 likes47 downloads8mo agoHugging Face15mteb /cstr-vctk-gender-miniaudio10K<n<100K0 likes45 downloads8mo agoHugging Face16diffunity /GLOBE_V3_gender_N_allsplitsaudio10K<n<100K0 likes40 downloads7mo agoHugging Face17mmn3690 /voice-gender-clustering Dataset Details VoxCelebs Dataset separated by gender (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection) Dataset Description Celebrities voice recordings separated by their gender. Dataset Sources [optional] VoxCeleb dataset (https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox2.html) \Separation (https://dagshub.com/DagsHub/audio-datasets/src/main/voice_gender_detection) audio1K<n<10K0 likes38 downloads2y agoHugging Face18SLLMBias /qa_BBQ_bi_gender Dataset Card for "qa_BBQ_bi_gender" More Information needed audio1K<n<10K0 likes37 downloads2y agoHugging Face19mteb /globe-v2-gender-miniaudio10K<n<100K0 likes36 downloads8mo agoHugging Face20SayantanJoker /IndicVoices_R_Hindi_Gender1_Age0audio1K<n<10K1 likes34 downloads2y agoHugging Face21SLLMBias /qa_CBBQ_bi_gender Dataset Card for "qa_CBBQ_bi_gender" More Information needed audio1K<n<10K0 likes33 downloads2y agoHugging Face22SPARCO-project /benchmark-genderaudio1K<n<10K0 likes33 downloads4mo agoHugging Face23jacobstgerm /speech_commands_10labels_genderevenlabelsaudio10K<n<100K0 likes28 downloads9mo agoHugging Face24AsmaaQ /gender_audio_1080 Dataset Card for mcv_spk_emb This is the speaker embeddings (xvectors) of mozilla common voice_11 speakers, -with the original audios-. Vectors are extracted with speechBrain's xvector model using this script by concatinating 11 splits from MCV_11 making a set of 1080 different speakers and 30k audio samples using this script. The main goal of this data was to train and test a voice gender detection classifier. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/AsmaaQ/gender_audio_1080.audioaudio-classification10K<n<100K0 likes27 downloads2y agoHugging Face25SLLM-multi-hop /GenderQA Dataset Card for SAKURA-GenderQA This dataset contains the audio and the single/multi-hop questions/answers of the gender track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information". The fields of the dataset are: file: The filename of the audio files. audio: The audio recordings. attribute_label: The attribute labels (i.e., gender of the speakers) of the audio files.… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/GenderQA.audion<1K0 likes25 downloads1y agoHugging Face26windcrossroad /GenderQA-gemini-1.5-flash Dataset Card for "GenderQA-gemini-1.5-flash" More Information needed audion<1K0 likes23 downloads2y agoHugging Face27windcrossroad /GenderQA-gemini-1.5-pro-caption Dataset Card for "GenderQA-gemini-1.5-pro-caption" More Information needed audion<1K0 likes21 downloads2y agoHugging Face28negfir /speech_commands_with_genderaudio10K<n<100K0 likes21 downloads2y agoHugging Face29MUGEN-Benchmark /Gender_Emotion_Filteringaudion<1K0 likes21 downloads8mo agoHugging Face30windcrossroad /GenderQA-LTUAS-caption Dataset Card for "GenderQA-LTUAS-caption" More Information needed audion<1K0 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.