datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.Silver-Multimodal-Dataset
Dataset Overview
The dataset is designed to support the development of machine learning models for detecting daily activities, violence, and fall down scenarios from combined audio and video sources.
The preprocessing pipeline leverages audio feature extraction, human keypoint detection, and relative positional encoding to generate a unified representation for training and inference.
Classes:
0: Daily - Normal indoor activities
1: Violence - Aggressive behaviors
2: Fall Down -… See the full description on the dataset page: https://huggingface.co/datasets/SilverAvocado/Silver-Multimodal-Dataset.Intel_Robotic_Welding_Multimodal_Dataset
Dataset Card for the Intel Robotic Welding Multimodal Dataset
This dataset was collected to enable multimodal welding defect detection research. The dataset contains over 4000 annotated samples and was collected in an automotive production floor setting in collaboration with a supplier with access to such facilities. Each sample contains a video, associated audio, a time-series from welding sensors, and five post-weld images for a particular weld. A separately licensed… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/Intel_Robotic_Welding_Multimodal_Dataset.mff-multimodal-dataset
MFF Multimodal Video Editing Dataset
This dataset contains source videos, text editing prompts, and style reference images
used for multimodal video editing experiments.
Each row in metadata.jsonl pairs one source video, one text prompt, and one style
image. The dataset contains 117 rows: 13 videos × 3 prompts × 3 style images.
Columns
example_id: Unique row identifier.
frame_group: Source video group, one of 8-frames, 36-frames, or 90-frames.
num_frames: Number of… See the full description on the dataset page: https://huggingface.co/datasets/AviadDahan/mff-multimodal-dataset.Neurodiverse_Multimodal_DatasetUMED-Urdu-Multimodal-Emotion-DatasetUrdu-Multimodal-Emotion-DatasetPanchayat-multimodal-datasetThis dataset was created from scratch for research on humour in conversational dialogues.The aim was to contribute a high-quality multimodal resource for the Hindi language.
The dataset contains:
Hindi text written in Devanagari script
English words transcribed exactly as spoken
Conversational multimodal data including text, audio, and video
The transcription style preserves natural conversational context and code-mixing patterns commonly found in spoken Hindi.
Citation
If you… See the full description on the dataset page: https://huggingface.co/datasets/Abhis4e/Panchayat-multimodal-dataset.multimodal-humor
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/corpora/sldr000837
Description
La collection "multimodal humor" contient les exemples de séquences vidéos associées à un chapitre d'ouvrage en cours (les références seront données au moment de la publication).
Il s'agit d'extraits de suivi longitudinaux adulte-enfant tirés du corpus CoLaJE déjà entièrement ouvert dans ORTOLANG.
https://www.ortolang.fr/market/corpora/colaje
Les enfants sont ANAE et THEOPHILE. Les enfants ont… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/multimodal-humor.Neurodiverse_Multimodal_Video_Clip_Dataset
