datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mls_hq_urgent_track1ELSA1M_track1
ELSA - Multimedia use case
ELSA Multimedia is a large collection of Deep Fake images, generated using diffusion models
Dataset Summary
This dataset was developed as part of the EU project ELSA. Specifically for the Multimedia use-case.
Official webpage: https://benchmarks.elsa-ai.eu/
This dataset aims to develop effective solutions for detecting and mitigating the spread of deep fake images in multimedia content. Deep fake images, which are highly realistic and deceptive… See the full description on the dataset page: https://huggingface.co/datasets/elsaEU/ELSA1M_track1.mls-hq-urgent-track1
Multilingual LibriSpeech HQ (MLS-HQ)
This is a mirror of the Multilingual LibriSpeech HQ (MLS-HQ) data used in URGENT 2025 Track 1.
The original files were converted from FLAC to Opus to reduce the size and accelerate streaming.
Sampling rate: 48 kHz (resampled from 44.1 kHz to support Opus format)
Channels: 1
Format: Opus
Splits:
spanish: 150 hours, 36031 utterances
german: 150 hours, 35890 utterances
french: 150 hours, 36078 utterances
License: CC0 1.0
Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/mls-hq-urgent-track1.vmc2026-track1-dev
vmc2026-track1-dev
Development subset of the VMC 2026 Track 1 data.
The data is organized into two configs corresponding to two subjective evaluation paradigms: absolute rating (acr) and pairwise comparison (ccr). The sample_id values are namespaced strings such as vmc2026-track1-dev-acr_489 and vmc2026-track1-dev-ccr_7233.
acr -- Absolute Category Rating
1,008 samples. Each row pairs a sample_id with one speech audio file, its released Mean Opinion Score (MOS)… See the full description on the dataset page: https://huggingface.co/datasets/urgent-challenge/vmc2026-track1-dev.vmc2026-track1-test
vmc2026-track1-test
Test subset of the VMC 2026 Track 1 data.
The data is organized into two configs corresponding to two subjective evaluation paradigms: absolute rating (acr) and pairwise comparison (ccr). The sample_id values are namespaced strings such as vmc2026-track1-test-acr_4588 and vmc2026-track1-test-ccr_3061.
acr -- Absolute Category Rating
4,032 samples. Each row pairs a sample_id with one speech audio file, its released Mean Opinion Score (MOS)… See the full description on the dataset page: https://huggingface.co/datasets/urgent-challenge/vmc2026-track1-test.track1elsst-track1
ELSST Track1: Implicit Concept Retrieval
ELSST Track1 evaluates whether a model can read a long synthetic social-science passage and retrieve the most relevant concepts from a fixed ELSST concept pool. The target is not lexical matching. The concepts are intentionally implicit, cross-sentence, and often require discourse-level reasoning over topic, framing, and social context.
This card is the authoritative task description for the retrieval track. The companion generation track is… See the full description on the dataset page: https://huggingface.co/datasets/JohnWang10086/elsst-track1.parc2026-track1-texture-smoke-v1
Track 1 texture mask smoke result
Two real selected episodes were decoded and segmented on A100. This repository stores the reproducibility evidence only: masks, fixed split reference, job definitions, summary, and execution log. It does not contain the original videos or constitute the final training dataset.
Active-Track1dstc12_track1_chatevalIf you use this dataset please use the following citation:
@inproceedings{mendonca2025dstc12t1,
author = "John Mendonça and Lining Zhang and Rahul Mallidi and Luis Fernando D'Haro and João Sedoc",
title = "Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12",
booktitle = "DSTC12: The Twelfth Dialog System Technology Challenge",
series = "26th Meeting of the Special Interest Group on… See the full description on the dataset page: https://huggingface.co/datasets/codesj/dstc12_track1_chateval.blind_test_track1_HEADSETMiGA26_Track1_Skeleton_pose90k_FT
Skeleton (MAMP, 90k iMiGUE pose) - fine-tuned on iMiGUE
MAMP self-supervised pre-training on 90k iMiGUE 46-joint pose skeletons, followed by full fine-tuning (train+val) and cRT classifier re-training on the iMiGUE-32 gesture classification task.
This model achieves 70.188% Top-1 on the official test set, and contributes to our final 74.66% ensemble (1st on Kaggle).
📚 0. Table of Contents
📦 Installation
📂 Data Preparation
🏋️♂️ Training & Testing
Pre-trained… See the full description on the dataset page: https://huggingface.co/datasets/anyi777/MiGA26_Track1_Skeleton_pose90k_FT.MiGA26_Track1_Skeleton_pose_ma52_FT
Skeleton (MAMP, MA-52 DWPose hybrid49) - fine-tuned on iMiGUE
MAMP self-supervised pre-training on MA-52 49-joint DWPose skeletons, followed by full fine-tuning (train+val) and cRT classifier re-training on the iMiGUE-32 gesture classification task (36 keypoints zero-padded to 49).
This model achieves 59.688% Top-1 on the official test set (finetune-only historical baseline) and contributes to our final 74.66% ensemble (1st on Kaggle).
📚 0. Table of Contents
📦… See the full description on the dataset page: https://huggingface.co/datasets/anyi777/MiGA26_Track1_Skeleton_pose_ma52_FT.blind_test_dataset_v2_track1results_SE2_blind_test_dataset_v2_track1
