datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sdf_dataset_en
SpeechDialogueFactory Dataset
Background
This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_en.sdf_dataset_zh
SpeechDialogueFactory Dataset
Background
This dataset is part of the SpeechDialogueFactory project, a comprehensive framework for generating high-quality speech dialogues at scale. Speech dialogue datasets are essential for developing and evaluating Speech-LLMs, but existing datasets face limitations including high collection costs, privacy concerns, and lack of conversational authenticity. This dataset addresses these challenges by providing synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/minghanw/sdf_dataset_zh.mia-meeting
MIA Meeting E2E Dataset
Synthetic meeting dataset for end-to-end experiments:
audio to transcript
transcript plus roster to action items
action item extraction benchmark
Splits
train: 200 samples, 0 with linked audio
validation: 5 samples, 5 with linked audio
eval: 205 samples, 5 with linked audio
Structure
data/*.jsonl # split manifests
audio/<split>/* # linked audio files when available
transcripts/<split>/*.json #… See the full description on the dataset page: https://huggingface.co/datasets/minhthien/mia-meeting.
