datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.multilingual-speech
Multilingual Indian Conversational Speech
A dataset of naturalistic, spontaneous two-speaker conversations across
13 Indian languages, with segment-level transcripts, speaker profiles,
timestamps, and recording metadata. Designed for ASR, TTS, speaker
diarization, and conversational speech research.
Languages (13)
Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Nepali,
Odia, Punjabi, Tamil, Telugu, Urdu.
Content
Conversations… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/multilingual-speech.MultiFraudAlign
FraudAlign-MCS
A fraud-only multilingual & code-switched dataset of scam-call dialogues, natively generated
(not translated) with Qwen2.5-72B-Instruct-AWQ. Modeled on the schema, fraud
taxonomy, and per-type proportions of the Chinese TeleAntiFraud-28k dataset,
regenerated from scratch in 4 languages: English (en), Hindi (hi), Korean (ko), Hinglish (Hindi-English code-switch) (hinglish).
28,708 dialogues total (7,177 per language), built to support
alignment of audio language… See the full description on the dataset page: https://huggingface.co/datasets/ggirishg/MultiFraudAlign.multimodal-ai-taxonomy
Multimodal AI Taxonomy
A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities.
Dataset Description
This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to:
Navigate the complex landscape of multimodal AI models
Filter models by specific input/output modality combinations
Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.
