datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.pl-nsa
Dataset Card for JuDDGES/nsa
Dataset Summary
The dataset consists of Supreme Administrative Court of Poland judgements available at orzeczenia.nsa.gov.pl, containing full content of the judgements along with metadata sourced from the official website.
The dataset contains documents up to 2025-03-05, with the last update on 2025-03-06. Some recent documents may be missing. The NSA database is continuously updated, though delays may cause older documents to appear over… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa.UniTalk-ASD
Data storage for the Active Speaker Detection Dataset: UniTalk
Le Thien Phuc Nguyen*, Zhuoran Yu*, Khoa Cao Quang Nhat, Yuwei Guo, Tu Ho Manh Pham, Tuan Tai Nguyen, Toan Ngo Duc Vo, Lucas Poon, Soochahn Lee, Yong Jae Lee
(* Equal Contribution)
Storage Structure
Since the dataset is large and complex, we zip the video_id folder and store on Hugging Face.
Here is the raw structure on Hungging face:
root/
├── csv/
│ ├── val
| | |_ video_id1.csv
| | |_… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/UniTalk-ASD.pl-nsa-enriched
Polish NSA Judgments (Enriched)
Polish Supreme Administrative Court judgments enriched with Gemini-extracted factual_state and legal_state fields.
Dataset Description
This dataset is an enriched version of JuDDGES/pl-nsa with additional fields extracted using Google Gemini 2.5 Pro.
New Fields
Core Extracted Fields
Field
Type
Description
factual_state
string
Objective narrative of facts (stan faktyczny) - the factual circumstances forming… See the full description on the dataset page: https://huggingface.co/datasets/JuDDGES/pl-nsa-enriched.UniTalkAudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.cowese
Dataset Card for "cowese"
More Information needed
readability-es-hackathon-pln-public
Dataset Card for [readability-es-sentences]
Dataset Description
Compilation of short Spanish articles for readability assessment.
Dataset Summary
This dataset is a compilation of short articles from websites dedicated to learn Spanish as a second language. These articles have been compiled from the following sources:
Coh-Metrix-Esp corpus (Quispesaravia, et al., 2016): collection of 100 parallel texts with simple and complex variants in Spanish. These texts… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2022/readability-es-hackathon-pln-public.spanish-embeddings-taller-plnwl-abbreviationsentiment-banking-plnpl-newspaper-pages-ocr-dataset-100
Polish Press Pages OCR Dataset 100
A small public image-only OCR dataset containing 100 Polish-language press / magazine-style page images sampled from a multilingual document collection.
This dataset is designed as a lightweight evaluation sample for:
OCR models
VLM-based document understanding
testing OCR robustness on Polish multi-column and press-style layouts
Contents
100 JPG images
Polish-language page images
press / magazine-style layouts, including:
article… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-newspaper-pages-ocr-dataset-100.plndataset
LARI – Lightweight Adaptation for Response and Inquiry
Dataset Description
LARI is a native Question and Answer (QA) dataset in Brazilian Portuguese, composed of 464 context-question-answer trios validated by human experts. It was created to address the scarcity of native and educationally grounded benchmarks for Portuguese.
Unlike most QA resources available in Portuguese (derived from automatic translations of English datasets) LARI was built from content originally… See the full description on the dataset page: https://huggingface.co/datasets/jjuliar/plndataset.PLN_Vladimir_DatasetVideo_Odysseywl-findingwl-family-memberfact2020In this paper we present the second edition of the FACT shared task (Factuality Annotation and Classification
Task), included in IberLEF2020. The main objective of this task is to advance in the study of the factuality of
the events mentioned in texts. This year, the FACT task includes a subtask on event identification in addition
to the factuality classification subtask. We describe the submitted systems as well as the corpus used, which is
the same used in FACT 2019 but extended by adding annotations for nominal events.achs-privacy-medicalList of entities
label_list = ["O", "B-Body_Part", "I-Body_Part", "B-Disease", "I-Disease", "B-Medication", "I-Medication",
"B-Age", "I-Age", "B-Company", "I-Company", "B-Health_Care_Unit", "I-Health_Care_Unit",
"B-Date_Part", "I-Date_Part", "B-Full_Date", "I-Full_Date",
"B-First_Name", "I-First_Name", "B-Last_Name", "I-Last_Name", "B-Location", "I-Location",
"B-Occupation", "I-Occupation", "B-Phone_Number", "I-Phone_Number", "B-RUT"… See the full description on the dataset page: https://huggingface.co/datasets/plncmm/achs-privacy-medical.wlwl-procedurespanish-alpaca
Dataset Card for "spanish-alpaca"
More Information needed
recursos-pln-es-modelsplnwl-medicationpl-newspaper-pages-ocr-dataset-100-v1
Document OCR using GLM-OCR
This dataset contains OCR results from images in Lukaszl/pl-newspaper-pages-ocr-dataset-100 using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: Lukaszl/pl-newspaper-pages-ocr-dataset-100
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 100
Processing Time: 11.2 min
Processing Date: 2026-03-31 22:16 UTC
Configuration
Image Column: image
Output Column: markdown… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-newspaper-pages-ocr-dataset-100-v1.sentiment-bankingwl-diseasecowese-sample
Dataset Card for "cowese-sample"
More Information needed
wl-body-part
