PLN
Datasets
All datasets matching “PLN”AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.pl-nsa
Dataset Card for JuDDGES/nsa
Dataset Summary
The dataset consists of Supreme Administrative Court of Poland judgements available at orzeczenia.nsa.gov.pl, containing full content of the judgements along with metadata sourced from the official website.
The dataset contains documents up to 2025-03-05, with the last update on 2025-03-06. Some recent documents may be missing. The NSA database is continuously updated, though delays may cause older documents to appear over… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa.UniTalk-ASD
Data storage for the Active Speaker Detection Dataset: UniTalk
Le Thien Phuc Nguyen*, Zhuoran Yu*, Khoa Cao Quang Nhat, Yuwei Guo, Tu Ho Manh Pham, Tuan Tai Nguyen, Toan Ngo Duc Vo, Lucas Poon, Soochahn Lee, Yong Jae Lee
(* Equal Contribution)
Storage Structure
Since the dataset is large and complex, we zip the video_id folder and store on Hugging Face.
Here is the raw structure on Hungging face:
root/
├── csv/
│ ├── val
| | |_ video_id1.csv
| | |_… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/UniTalk-ASD.pl-nsa-enriched
Polish NSA Judgments (Enriched)
Polish Supreme Administrative Court judgments enriched with Gemini-extracted factual_state and legal_state fields.
Dataset Description
This dataset is an enriched version of JuDDGES/pl-nsa with additional fields extracted using Google Gemini 2.5 Pro.
New Fields
Core Extracted Fields
Field
Type
Description
factual_state
string
Objective narrative of facts (stan faktyczny) - the factual circumstances forming… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa-enriched.UniTalkAudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.
