datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ELSA500k_track2
ELSA - Multimedia use case
ELSA Multimedia is a large collection of Deep Fake images, generated using diffusion models
Dataset Summary
This dataset was developed as part of the EU project ELSA. Specifically for the Multimedia use-case.
Official webpage: https://benchmarks.elsa-ai.eu/
This dataset aims to develop effective solutions for detecting and mitigating the spread of deep fake images in multimedia content. Deep fake images, which are highly realistic and… See the full description on the dataset page: https://huggingface.co/datasets/elsaEU/ELSA500k_track2.parc2026-track2-network-archive-20260912
Track 2 experiment recovery archive
This repository preserves deduplicated data and numerical evidence from the
owner's completed PARC 2026 Track 2 experiments. It is not a new training
dataset release or evidence of model improvement. Original source licenses
and attributions remain applicable to bundled code/assets.
The migration may still be in progress. A volume's RECOVERY_CATALOG.json is
published only after all its archives, original-path mappings, quarantine
contents and… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/parc2026-track2-network-archive-20260912.AT-ADD-Track2
AT-ADD Track 2
This repository hosts Track 2 of the AT-ADD All-Type Audio Deepfake Detection Challenge. It contains the released audio splits and privacy-preserving sample-level metadata for non-commercial academic research and education.
Access
This is a gated dataset. Sign in to Hugging Face, review the access agreement, complete the short access form, and click Agree and access dataset. Access is granted automatically after acceptance.
Direct repository access… See the full description on the dataset page: https://huggingface.co/datasets/xieyuankun/AT-ADD-Track2.egolink_track2dstc10_track2_task2\n2c2_2018_track2The National NLP Clinical Challenges (n2c2), organized in 2018, continued the
legacy of i2b2 (Informatics for Biology and the Bedside), adding 2 new tracks and 2
new sets of data to the shared tasks organized since 2006. Track 2 of 2018
n2c2 shared tasks focused on the extraction of medications, with their signature
information, and adverse drug events (ADEs) from clinical narratives.
This track built on our previous medication challenge, but added a special focus on ADEs.
ADEs are injuries resulting from a medical intervention related to a drugs and
can include allergic reactions, drug interactions, overdoses, and medication errors.
Collectively, ADEs are estimated to account for 30% of all hospital adverse
events; however, ADEs are preventable. Identifying potential drug interactions,
overdoses, allergies, and errors at the point of care and alerting the caregivers of
potential ADEs can improve health delivery, reduce the risk of ADEs, and improve health
outcomes.
A step in this direction requires processing narratives of clinical records
that often elaborate on the medications given to a patient, as well as the known
allergies, reactions, and adverse events of the patient. Extraction of this information
from narratives complements the structured medication information that can be
obtained from prescriptions, allowing a more thorough assessment of potential ADEs
before they happen.
The 2018 n2c2 shared task Track 2, hereon referred to as the ADE track,
tackled these natural language processing tasks in 3 different steps,
which we refer to as tasks:
1. Concept Extraction: identification of concepts related to medications,
their signature information, and ADEs
2. Relation Classification: linking the previously mentioned concepts to
their medication by identifying relations on gold standard concepts
3. End-to-End: building end-to-end systems that process raw narrative text
to discover concepts and find relations of those concepts to their medications
Shared tasks provide a venue for head-to-head comparison of systems developed
for the same task and on the same data, allowing researchers to identify the state
of the art in a particular task, learn from it, and build on it.parc2026-track2-perturb-dataset-20260923
track2 摂動カテゴリ学習データセット(PARC2026 Track2、2026-09-22〜23)
PARC2026 本選 Track2(LIBERO-plus の摂動つきタスク、π0.5 の微調整用)の教師データ。公開 8 タスクの元タスク × 公開例題に出る摂動カテゴリ 4 つ
(Sensor Noise / Camera Viewpoints / Robot Initial States / Objects Layout)を、P+D 制御の scripted expert の軌道で被覆したもの。
全 6,711 本、1,262,626 フレーム。すべて成功・非対象物の L1 変位 1 mm 以下・300 手以内・把持点と置く点の通り越し 1.1 mm 以下。
作り方・教師の設定・欠番の理由・途中で直した欠陥は manifests/TRACK2_PERTURB_DATASET_20260923.html(日本語)に書いてある。
中身
ディレクトリ
本数
フレーム
中身… See the full description on the dataset page: https://huggingface.co/datasets/ekunish/parc2026-track2-perturb-dataset-20260923.urgent26_track2_sqaelsst-track2
ELSST Track2: Open Knowledge Discovery
ELSST Track2 evaluates whether a model can read the same long synthetic passage used in Track1 and generate the latent ELSST concepts as a semantic set. The target is a small concept set, not a free-form explanation. Models must recover implicit concepts that are grounded in the passage but not always directly named.
This card is the authoritative task description for the generation track. The companion retrieval track is described in… See the full description on the dataset page: https://huggingface.co/datasets/JohnWang10086/elsst-track2.MIGA_track2W-CODA2024-Track2
W-CODA2024 Track 2 Dataset
Dataset Description
This dataset contains auxiliary data files for the W-CODA (Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving) Track 2 workshop at ECCV 2024. The files provide metadata about the nuScenes validation set for evaluating video generation and detection/segmentation results.
Data Files
nuscenes_infos_temporal_val_3keyframes.pkl
Contains information about key frames from 150 scenes in the… See the full description on the dataset page: https://huggingface.co/datasets/pengxiang/W-CODA2024-Track2.BrainStorm2026-track2Hum-Omni-Track2-Phase2
Phase 2 Test Data
This test set is part of the HuMomni 2026 competition.
Competition website: https://humomni2026.github.io/
Overview
The test set contains 500 video question-answering samples. Each sample consists of a natural language question about a short video clip, along with frames pre-extracted at 2 fps for convenience.
Data Format
Each sample is stored in its own directory under data/, named by the video ID (e.g., data/OSfMU69X3C4.7/).… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/Hum-Omni-Track2-Phase2.Hum-Omni-Track2-Phase1
Track 2 Phase 1 Test Data
This test set is part of the HumOmni 2026 competition.
Competition website: https://humomni2026.github.io/
Overview
The test set contains 500 video question-answering samples. Each sample consists of a natural language question about a short video clip, along with frames pre-extracted at 2 fps for convenience.
Data Format
Each sample is stored in its own directory under data/, named by the video ID (e.g., data/OSfMU69X3C4.7/).
data/… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/Hum-Omni-Track2-Phase1.n2c2_2018_track2The National NLP Clinical Challenges (n2c2), organized in 2018, continued the
legacy of i2b2 (Informatics for Biology and the Bedside), adding 2 new tracks and 2
new sets of data to the shared tasks organized since 2006. Track 2 of 2018
n2c2 shared tasks focused on the extraction of medications, with their signature
information, and adverse drug events (ADEs) from clinical narratives.
This track built on our previous medication challenge, but added a special focus on ADEs.
ADEs are injuries resulting from a medical intervention related to a drugs and
can include allergic reactions, drug interactions, overdoses, and medication errors.
Collectively, ADEs are estimated to account for 30% of all hospital adverse
events; however, ADEs are preventable. Identifying potential drug interactions,
overdoses, allergies, and errors at the point of care and alerting the caregivers of
potential ADEs can improve health delivery, reduce the risk of ADEs, and improve health
outcomes.
A step in this direction requires processing narratives of clinical records
that often elaborate on the medications given to a patient, as well as the known
allergies, reactions, and adverse events of the patient. Extraction of this information
from narratives complements the structured medication information that can be
obtained from prescriptions, allowing a more thorough assessment of potential ADEs
before they happen.
The 2018 n2c2 shared task Track 2, hereon referred to as the ADE track,
tackled these natural language processing tasks in 3 different steps,
which we refer to as tasks:
1. Concept Extraction: identification of concepts related to medications,
their signature information, and ADEs
2. Relation Classification: linking the previously mentioned concepts to
their medication by identifying relations on gold standard concepts
3. End-to-End: building end-to-end systems that process raw narrative text
to discover concepts and find relations of those concepts to their medications
Shared tasks provide a venue for head-to-head comparison of systems developed
for the same task and on the same data, allowing researchers to identify the state
of the art in a particular task, learn from it, and build on it.audiomos_track2blind_test_Track2_Speakerphone
