CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ml-resources /Daimon-Infinity Daimon-Infinity mirror This repository is a file-preserving mirror of daimonrobotics/Daimon-Infinity on ModelScope. Source and license Upstream: daimonrobotics/Daimon-Infinity License: CC BY-NC-SA 4.0 Attribution: Daimon Robotics / Daimon-Infinity This mirror keeps the upstream directory layout and is distributed under the same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.audio1K<n<10K2 likes39k downloads19d agoHugging Face02ziyou-li /cantonese_dailyaudio1K<n<10K4 likes1.5k downloads4y agoHugging Face03titasmallick96 /daily-bio-newsaudion<1K0 likes1.5k downloads2h agoHugging Face04kyutai /DailyTalkContiguous DailyTalkContiguous This repo contains a concatenated version of the DailyTalk dataset (official repo). Rather than having separate files for each speaker's turn, this uses a stereo file for each conversation. The two speakers in a conversation are put separately on the left and right channels. The dataset is annotated with word level timestamps. The original DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license, this dataset uses the… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/DailyTalkContiguous.audio21 likes1.4k downloads2y agoHugging Face05xxayt /Daily-Omni Daily-Omni This repository provides the question-answering metadata of the Daily-Omni benchmark in a format compatible with lmms-eval. The data is provided as a single parquet file containing only the QA annotations. Since raw videos are not included, please download them from the original release and match them with the QA annotations using video_id. Task configurations and evaluation scripts are available in the SEATS repository: https://github.com/xxayt/SEATS.… See the full description on the dataset page: https://huggingface.co/datasets/xxayt/Daily-Omni.textvideo-text-to-text1K<n<10K1 likes395 downloads4mo agoHugging Face06wsntxxn /DailyTalkEdit Dataset Structure This repo contains DailyTalkEdit and the extended semantic influence annotation of PartialEdit. daily_talk_edit/ concat_utts/: paired audio files per sample, obtained by concatenating all utterances in dialogues <id>_original.wav <id>_modified.wav modified_utts/: modified-only audio segments <id>_<utt_id>_modified.wav metadata/ train.jsonl, val.jsonl, test.jsonl filename: audio file name fake_region: modified time ranges semantic_influence: text description of… See the full description on the dataset page: https://huggingface.co/datasets/wsntxxn/DailyTalkEdit.audioaudio-classification10K<n<100K1 likes343 downloads8mo agoHugging Face07hf-internal-testing /dailytalk-dummyaudion<1K4 likes333 downloads1y agoHugging Face08eustlb /dailytalk Dataset Card for "dailytalk" More Information needed audio10K<n<100K2 likes294 downloads1y agoHugging Face09RowanSu /DailyOmniThis is the official dataset for Daily-Omni. Check code repository for instructions. audioquestion-answering0 likes242 downloads2mo agoHugging Face10FormosanBank /ePark_sheng_huo_hui_hua_pian_daily_conversation FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.audioautomatic-speech-recognition10K<n<100K0 likes208 downloads2mo agoHugging Face11saeedzou /e-daic-ai-controlledgatedaudio1K<n<10K0 likes161 downloads2mo agoHugging Face12saeedzou /DAIC-WOZgatedaudio10K<n<100K0 likes113 downloads2mo agoHugging Face13imrnh /dailytalk_miniA subset of daily https://huggingface.co/datasets/kyutai/DailyTalkContiguous with only 100 wav file and their corresponding JSON files. audio0 likes100 downloads1y agoHugging Face14atonyxu /daisy-voice-dataset数据集文件元信息以及数据文件,请浏览“数据集文件”页面获取。 当前数据集卡片使用的是默认模版,数据集的贡献者未提供更加详细的数据集介绍,但是您可以通过如下GIT Clone命令,或者ModelScope SDK来下载数据集 下载方法 :modelscope-code[]{type="sdk"} :modelscope-code[]{type="git"} audio1K<n<10K0 likes98 downloads3mo agoHugging Face15eustlb /dailytalk-conversations-grouped Dataset Card for "dailytalk-conversations-grouped" This dataset is intended for testing fine-tuning of Sesame’s CSM-1B (Conversational Speech Model). It is extracted from the DailyTalk dataset. audio1K<n<10K11 likes96 downloads1y agoHugging Face16DynamicSuperbPrivate /DialogueActPairing_DailyTalk Dataset Card for "DialogueActPairing_DailyTalk" More Information needed audio10K<n<100K0 likes82 downloads3y agoHugging Face17DynamicSuperb /DialogueActClassification_DailyTalk Dataset Card for "DailyTalk_DialogueActClassification" More Information needed audion<1K0 likes79 downloads3y agoHugging Face18yiwu2 /daily_dialogue_mixed_chinese_english_speech_ttsaudio10K<n<100K6 likes69 downloads2y agoHugging Face19jakeboggs /daily-yap Daily Yap Dataset Dataset Description Overview The Daily Yap dataset is an enhanced and audio-augmented version of the Daily Dialog dataset. It consists of refined conversation transcripts that have been converted into dual-channel audio files using text-to-speech technology. Blog Post: Open-sourcing 100 Hours of Conversational Audio (Daily Yap) Source Data Original Dataset: Daily Dialog Dataset Page: http://yanran.li/dailydialog Paper:… See the full description on the dataset page: https://huggingface.co/datasets/jakeboggs/daily-yap.audio1K<n<10K2 likes63 downloads2y agoHugging Face20DynamicSuperbPrivate /DialogueEmotionClassification_DailyTalk Dataset Card for "DialogueEmotionClassification_DailyTalk" More Information needed audio10K<n<100K0 likes59 downloads3y agoHugging Face21AlienKevin /guangzhou-daily-use-speechASR-SCCantDuSC: A Scripted Chinese Cantonese (Canton) Daily-use Speech Corpus This open-source dataset consists of 4.06 hours of transcribed Guangzhou Cantonese scripted speech focusing on daily use sentences, where 4,060 utterances contributed by ten speakers were contained. Source: https://magichub.com/datasets/guangzhou-cantonese-scripted-speech-corpus-daily-use-sentence/ audio1K<n<10K1 likes59 downloads2y agoHugging Face22irlab-daiict /SqCLIRIL 🗣️ SqCLIRIL: Spoken Query Benchmark for Cross-Lingual IR in Indian Languages SqCLIRIL is a Spoken Query Benchmark designed to evaluate cross-lingual information retrieval (CLIR) systems using both spoken and text queries.It covers five Indian languages — Hindi, Gujarati, Bengali, Kannada, and English — with diverse speech samples from male and female speakers to capture natural variability in pronunciation and acoustic conditions. 📘 Dataset Summary Feature… See the full description on the dataset page: https://huggingface.co/datasets/irlab-daiict/SqCLIRIL.audioautomatic-speech-recognition1K<n<10K0 likes59 downloads1y agoHugging Face23DynamicSuperbPrivate /DialogueActClassification_DailyTalk Dataset Card for "DailyTalk_DialogueActClassification" More Information needed audio10K<n<100K0 likes50 downloads3y agoHugging Face24ericholam /dailytalk-conversations-grouped-llm-codecaudio1K<n<10K0 likes34 downloads11mo agoHugging Face25innovationm-ai /dailytalk-male Dataset Summary This dataset is a gender-specific subset of the original DailyTalk TTS dataset.It contains English conversational speech paired with text transcripts, filtered and separated by speaker gender. This version includes: Two columns: text: transcription audio: 24 kHz speech waveform One split (train) with 11,906 samples No audio processing or text modifications were performed. The dataset is a structured subset of the original source. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/innovationm-ai/dailytalk-male.audiotext-to-speech10K<n<100K0 likes31 downloads10mo agoHugging Face26DynamicSuperb /DialogueEmotionClassification_DailyTalk Dataset Card for "DialogueEmotionClassification_DailyTalk" More Information needed audion<1K0 likes30 downloads3y agoHugging Face27shujaAK /hindi-dairy-asr-cleanaudio1K<n<10K0 likes28 downloads3mo agoHugging Face28AIxBlock /USA-accented-role-playing-daily-conversations-stereo Dataset Card for Synthetic daily conversations - USA accented - stereo wav This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. Dataset Details Dataset Description This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. The dialogues are spoken spontaneously, covering a range of everyday topics such… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/USA-accented-role-playing-daily-conversations-stereo.audion<1K5 likes27 downloads1y agoHugging Face29HaninZ /DialogueEmotionClassification_DailyTalkaudio10K<n<100K0 likes26 downloads2y agoHugging Face30HaninZ /DialogueActClassification_DailyTalkaudio10K<n<100K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.