datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiOOD
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
Hao Dong1
Yue Zhao2
Eleni Chatzi1
Olga Fink3
1ETH Zurich, 2University of Southern California, 3EPFL
• arXiv •
MultiOOD is the first-of-its-kind benchmark for Multimodal OOD Detection, characterized by diverse dataset sizes and varying modality combinations.
Code
https://github.com/donghao51/MultiOOD
MultiOOD Benchmark
MultiOOD… See the full description on the dataset page: https://huggingface.co/datasets/hdong51/MultiOOD.qwen3-tts-multilingual-emotional-speechMulberry-SFTPlease check our GitHub for more details.: https://github.com/HJYao00/Mulberry
Training
We use LLaMA-Factory to fine-tune the Mulberry models. We provide the training instructions and configs here.
First, install LLaMA-Factory according to the official_instruction.
Then, refer here and update the following customized dataset into dataset_info.json in LLaMA-Factory.
"mulberry": {
"file_name": "./mulberry_sft.json",
"formatting": "sharegpt",
"columns": {
"messages":… See the full description on the dataset page: https://huggingface.co/datasets/HuanjinYao/Mulberry-SFT.multi_round_speech_180kGUI-AIMA-multiturnSynthdog-Multilingual-100
Synthdog Multilingual
The Synthdog dataset created for training in Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model.
Using the official Synthdog code, we created >1 million training samples for improving OCR capabilities in Large Vision-Language Models.
Dataset Details
We provide the images for download in two .tar.gz files. Download and extract them in folders of the same name (so cat images.tar.gz.* | tar xvzf -C images; tar xvzf… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/Synthdog-Multilingual-100.multiz100way-pigzSource: https://huggingface.co/datasets/songlab/multiz100wayInspired by https://huggingface.co/datasets/lpigou/89.zarr
99.zarr.tar.gz is created with pigz with option -1 for fastest decompression.
Install pigz if not already installed:
sudo apt install pigz
Download and extract:
wget https://huggingface.co/datasets/gonzalobenegas/99.zarr/resolve/main/99.zarr.tar.gz
unpigz < 99.zarr.tar.gz | tar -x
Multiphysics_Bench
Multiphysics Bench
Dataset: huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench
Paper: Multiphysics Bench: Benchmarking and Investigating Scientific Machine Learning for Multiphysics PDEs
We propose the first general multiphysics benchmark dataset that encompasses six canonical coupled scenarios across domains such as electromagnetics, heat transfer, fluid flow, solid mechanics, pressure acoustics, and mass transport. This benchmark features the most comprehensive coupling types… See the full description on the dataset page: https://huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench.sdxl_images_sb_prompts-multi_artist-seed1wds_voc2007_multilabelZImage-Turbo-200k-multires-aspectbucketed
ZImage-Turbo WebDataset
Generated images from the ZImage-Turbo model using DiffusionDB prompts.
Generation Details
Hardware: 8x NVIDIA RTX 3090 GPUs
Generation Time: ~2 days
Estimated Cost: ~$70 (cloud compute)
Dataset Statistics
Total Samples: 211,081
Total Shards: 216
Samples per Shard: ~1000
Shard Naming Convention
Tarballs are named: {base_resolution}-{aspect_ratio}-{shard_num:04d}-of-{total_shards:04d}.tar
For example:… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/ZImage-Turbo-200k-multires-aspectbucketed.Multilingual_Speech_Dataset
Multilingual Speech Dataset
Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English
Repository: https://github.com/IS2AI/MultilingualASR
Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.multi_accent_speech
Multi-Accent English Speech Corpus (Augmented & Speaker-Disjoint)
This dataset is a curated and augmented multi-accent English speech corpus designed for speech recognition, accent classification, and representation learning.It consolidates multiple open-source accent corpora, converts all audio to a unified format, applies targeted data augmentation, and exports in a tidy, Hugging Face–ready structure.
✨ Key Features
Accents covered (12 total):american_english… See the full description on the dataset page: https://huggingface.co/datasets/cagatayn/multi_accent_speech.multi-target-spacecraft-pose-estimation
Multi-target Synthetic Dataset for Spacecraft Pose Estimation
Overview
This dataset was developed for 6D pose estimation of unseen, non-cooperative spacecraft in proximity-operations scenarios. Most existing datasets focus on a single target, which leads models to overfit to a specific spacecraft and limits their ability to generalize to previously unseen targets.
To address this limitation, the present dataset is multi-target and includes a wide variety of spacecraft… See the full description on the dataset page: https://huggingface.co/datasets/lorenzobottelli/multi-target-spacecraft-pose-estimation.multilingual-librispeech-webdatasetmodelnet40_multi_viewMulti-Domain-Sentiment-DatasetUsing it for assessment.
Dataset for Multi Domain (Including Kitchen, Books, DVDs, and Electronics)
Multi-Domain Sentiment Dataset by John Blitzer, Mark Dredze, Fernando Pereira.
Description:
The Multi-Domain Sentiment Dataset contains product reviews taken from Amazon.com from 4 product types (domains): Kitchen, Books, DVDs, and Electronics. Each domain has several thousand reviews, but the exact number varies by domain. Reviews contain star ratings (1 to 5 stars) that… See the full description on the dataset page: https://huggingface.co/datasets/JSSICE/Multi-Domain-Sentiment-Dataset.multilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.MultiEventVideomultilingual_ocr_erniesdkmultiguard-phase2-data
MultiGuard Phase 2 — Multimodal Misinformation Detection Dataset & Caches
This dataset packages everything needed to reproduce Phase 2 of the
MultiGuard
project without redoing the heavy preprocessing.
What's inside
File
Size
Contents
forensic_3class.csv
3.7 MB
18,000-sample 3-class dataset (Real / Manipulated / OOC), balanced 6,000 per class, deterministic 70/15/15 per-class split (seed 42).
dct_cache.tar.gz
2.9 GB
Precomputed DCT maps for all 18,000 images.… See the full description on the dataset page: https://huggingface.co/datasets/Rashidbm/multiguard-phase2-data.multiplepiper-plus-multilingual-7lang-v8-dataset
piper-plus multilingual 7-lang v8 training dataset
Preprocessed training dataset for piper-plus zero-shot TTS v8
(342,855 utterances / 3,692 speakers / 7 languages).
data/ holds the full dataset (audio_norm + spec caches, split tar.gz —
concatenate parts then extract). essential/ holds metadata only
(dataset.jsonl + config + CAM++ speaker embeddings + holdout) for fast restore.
Access is gated (manual approval) because the ja subset derives from
MoeSpeech (see… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/piper-plus-multilingual-7lang-v8-dataset.RACECAR-multislow_poliThis dataset contains the 4D point cloud data from LiDAR sensors collected from fully autonomous and self-driving Indy race cars which raced in the Indy autonomous challenge. The dataset is in nuScenes format and is divided into 7,150 sweeps and 1,199 samples which contain fused sensor data from 3 LiDARs equipped by the vehicle. This dataset's scenario is PoliMove team’s Multi-Agent Slow on LVMS racetrack.
Each .pcd file contains 4 dimensional data: (x,y,z) coordinates of the 3D space and… See the full description on the dataset page: https://huggingface.co/datasets/suwesh/RACECAR-multislow_poli.multimodal-open-r1-8kMultiFoley-VGGSound-Test-Audio
Video-Guided Foley Sound Generation with Multimodal Controls
Paper & Project page
This dataset contains the generated results of our MultiFoley work on the filtered VGGSound test cases. We generate 4 samples for each 8s video (we use the first 8s video for evaluation).
The results are generated with both silent video inputs and text inputs (we use the VGGSound category name for simplicity).
Each wave file is named in the format of {category_name}/{u_id}_{start_time}_{idx}.wav, where… See the full description on the dataset page: https://huggingface.co/datasets/czyang/MultiFoley-VGGSound-Test-Audio.MultimodalSentimentOnSocialMediamultilingual-in-the-wildsdxl_images_easy_prompts-multi_artist-seed0M4-Instruct-Multi
