datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Daimon-Infinity
Daimon-Infinity mirror
This repository is a file-preserving mirror of
daimonrobotics/Daimon-Infinity on ModelScope.
Source and license
Upstream: daimonrobotics/Daimon-Infinity
License: CC BY-NC-SA 4.0
Attribution: Daimon Robotics / Daimon-Infinity
This mirror keeps the upstream directory layout and is distributed under the
same CC BY-NC-SA 4.0 license. No data is altered; files are transferred with
integrity checks supplied by ModelScope and the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/ml-resources/Daimon-Infinity.cantonese_dailydaily-bio-newsDailyTalkContiguous
DailyTalkContiguous
This repo contains a concatenated version of the DailyTalk dataset (official repo).
Rather than having separate files for each speaker's turn, this uses a stereo file for each conversation. The two speakers in a conversation
are put separately on the left and right channels.
The dataset is annotated with word level timestamps.
The original DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license, this dataset
uses the… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/DailyTalkContiguous.Daily-Omni
Daily-Omni
This repository provides the question-answering metadata of the Daily-Omni benchmark in a format compatible with lmms-eval.
The data is provided as a single parquet file containing only the QA annotations. Since raw videos are not included, please download them from the original release and match them with the QA annotations using video_id.
Task configurations and evaluation scripts are available in the SEATS repository: https://github.com/xxayt/SEATS.… See the full description on the dataset page: https://huggingface.co/datasets/xxayt/Daily-Omni.DailyTalkEdit
Dataset Structure
This repo contains DailyTalkEdit and the extended semantic influence annotation of PartialEdit.
daily_talk_edit/
concat_utts/: paired audio files per sample, obtained by concatenating all utterances in dialogues
<id>_original.wav
<id>_modified.wav
modified_utts/: modified-only audio segments
<id>_<utt_id>_modified.wav
metadata/
train.jsonl, val.jsonl, test.jsonl
filename: audio file name
fake_region: modified time ranges
semantic_influence: text description of… See the full description on the dataset page: https://huggingface.co/datasets/wsntxxn/DailyTalkEdit.dailytalk-dummydailytalk
Dataset Card for "dailytalk"
More Information needed
DailyOmniThis is the official dataset for Daily-Omni. Check code repository for instructions.
ePark_sheng_huo_hui_hua_pian_daily_conversation
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_sheng_huo_hui_hua_pian_daily_conversation.e-daic-ai-controlledDAIC-WOZdailytalk_miniA subset of daily https://huggingface.co/datasets/kyutai/DailyTalkContiguous with only 100 wav file and their corresponding JSON files.
daisy-voice-dataset数据集文件元信息以及数据文件,请浏览“数据集文件”页面获取。
当前数据集卡片使用的是默认模版,数据集的贡献者未提供更加详细的数据集介绍,但是您可以通过如下GIT Clone命令,或者ModelScope SDK来下载数据集
下载方法
:modelscope-code[]{type="sdk"}
:modelscope-code[]{type="git"}
dailytalk-conversations-grouped
Dataset Card for "dailytalk-conversations-grouped"
This dataset is intended for testing fine-tuning of Sesame’s CSM-1B (Conversational Speech Model). It is extracted from the DailyTalk dataset.
DialogueActPairing_DailyTalk
Dataset Card for "DialogueActPairing_DailyTalk"
More Information needed
DialogueActClassification_DailyTalk
Dataset Card for "DailyTalk_DialogueActClassification"
More Information needed
daily_dialogue_mixed_chinese_english_speech_ttsdaily-yap
Daily Yap Dataset
Dataset Description
Overview
The Daily Yap dataset is an enhanced and audio-augmented version of the Daily Dialog dataset. It consists of refined conversation transcripts that have been converted into dual-channel audio files using text-to-speech technology.
Blog Post: Open-sourcing 100 Hours of Conversational Audio (Daily Yap)
Source Data
Original Dataset: Daily Dialog
Dataset Page: http://yanran.li/dailydialog
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/jakeboggs/daily-yap.DialogueEmotionClassification_DailyTalk
Dataset Card for "DialogueEmotionClassification_DailyTalk"
More Information needed
guangzhou-daily-use-speechASR-SCCantDuSC: A Scripted Chinese Cantonese (Canton) Daily-use Speech Corpus
This open-source dataset consists of 4.06 hours of transcribed Guangzhou Cantonese scripted speech focusing on daily use sentences, where 4,060 utterances contributed by ten speakers were contained.
Source: https://magichub.com/datasets/guangzhou-cantonese-scripted-speech-corpus-daily-use-sentence/
SqCLIRIL
🗣️ SqCLIRIL: Spoken Query Benchmark for Cross-Lingual IR in Indian Languages
SqCLIRIL is a Spoken Query Benchmark designed to evaluate cross-lingual information retrieval (CLIR) systems using both spoken and text queries.It covers five Indian languages — Hindi, Gujarati, Bengali, Kannada, and English — with diverse speech samples from male and female speakers to capture natural variability in pronunciation and acoustic conditions.
📘 Dataset Summary
Feature… See the full description on the dataset page: https://huggingface.co/datasets/irlab-daiict/SqCLIRIL.DialogueActClassification_DailyTalk
Dataset Card for "DailyTalk_DialogueActClassification"
More Information needed
dailytalk-conversations-grouped-llm-codecdailytalk-male
Dataset Summary
This dataset is a gender-specific subset of the original DailyTalk TTS dataset.It contains English conversational speech paired with text transcripts, filtered and separated by speaker gender.
This version includes:
Two columns:
text: transcription
audio: 24 kHz speech waveform
One split (train) with 11,906 samples
No audio processing or text modifications were performed. The dataset is a structured subset of the original source.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/innovationm-ai/dailytalk-male.DialogueEmotionClassification_DailyTalk
Dataset Card for "DialogueEmotionClassification_DailyTalk"
More Information needed
hindi-dairy-asr-cleanUSA-accented-role-playing-daily-conversations-stereo
Dataset Card for Synthetic daily conversations - USA accented - stereo wav
This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents.
Dataset Details
Dataset Description
This dataset consists of synthetic daily conversations recorded by native U.S. English speakers with authentic American accents. The dialogues are spoken spontaneously, covering a range of everyday topics such… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/USA-accented-role-playing-daily-conversations-stereo.DialogueEmotionClassification_DailyTalkDialogueActClassification_DailyTalk
