lv12/MultiModalDataset
Dataset Card for MultiModal Dataset Dataset Description Dataset Summary MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation. The dataset is organized into three subsets: fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.
Add OpenVid subset: 1000 videos, 4 frames each
Update README.md
Update README.md
Update README.md
Upload README.md with huggingface_hub
Upload dataset
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
Delete dataset_infos.json with huggingface_hub
Upload README.md with huggingface_hub
Upload dataset_infos.json with huggingface_hub
Upload README.md with huggingface_hub
Upload dataset
Upload dataset
initial commit
