datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MOVA_benchmark_for_arena
MOVA Benchmark for Arena
This is the benchmark used for the subjective arena experiments of MOVA (MOVA: Towards Scalable and Synchronized Video–Audio Generation). All prompts are rewritten by the workflow introduced in the paper.
Paper: MOVA: Towards Scalable and Synchronized Video–Audio Generation
Code: https://github.com/OpenMOVA/MOVA
Overview
The benchmark contains 732 samples in total, organized into two subsets:
Subset
Samples
MOVA-Bench
132… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuzhang-0212/MOVA_benchmark_for_arena.vibevoice-movarekhpodcast-single-speakermovarekhpodcast_dataset_persianmova2mova-dataiuras-bushliakou-zhyvaia-mova
Жывая мова
Metadata
Author: Юрась Бушлякоў
Title: Жывая мова
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250 MB.
Each split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/iuras-bushliakou-zhyvaia-mova.
