datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.migration-bench-java-selected
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.AV-DPO
AV-DPO
AV-DPO is the audio-video preference dataset used in Stage 3 of
JavisDiT++. It contains chosen-rejected
sounding-video pairs selected jointly across audio quality and text alignment,
video quality and text alignment, and audio-video semantic and temporal
alignment.
Release summary
Item
Count
Preference pairs (train.csv)
23,671
Self-contained generated-only pairs (train_generated_only.csv)
6,138
Unique released generated sounding videos
29,809… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/AV-DPO.JavisBench
JavisBench: A Challenging Benchmark for for Joint Audio-Video Generation (JAVG) Evaluation
As released in HuggingFace,
JavisBench is a comprehensive and challenging benchmark for evaluating text-to-audio-video generation models.It covers multiple aspects of generation quality, semantic alignment, and temporal synchrony, enabling thorough assessment in both controlled and real-world scenarios.
Installation
Install necessary packages:
cd /path/to/JavisDiT
pip install… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/JavisBench.java-vulnerabilityUnggah-Ungguh
Javanese Honorifics Dataset (Unggah-Ungguh - Released Version)
The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context.
Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.JavisData-AudioTitanic-Datasetjavanese-1k-dpo-cleanedhfexampleJAVAwjava-mgc-datasetiris26
