datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kairos
Kairos — Long-Form Video Annotation and Benchmark
Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes).
This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form
(OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and
5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s).
What's inside
.
├── data/
│ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.KairosQA
KairosQA Dataset
Dataset Description
KairosQA is a temporally grounded question-answering dataset designed to evaluate the temporal alignment and reasoning capabilities of Large Language Models (LLMs). Unlike static benchmarks, KairosQA focuses on facts that evolve over time, specifically subject–relation–object triplets from Wikidata that changed at least twice between 2018 and 2025 as described in our paper Understanding Data Temporality Impact on Large Language… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/KairosQA.KairosNewsHandcrafted Dataset used in the elaboration of a thesis and a project for a competition (Premio Arquivo.pt: https://sobre.arquivo.pt/pt/colabore/premios-arquivo-pt/premio-arquivo-pt-2025/)
It contains news articles from the following Portuguese News agencies from 2020 to 2024:
https://www.cmjornal.pt/ = 6771
https://expresso.pt/ = 22606
https://www.iol.pt/ = 39387
https://www.publico.pt/ = 74893
https://www.sapo.pt/ = 57838
TOTAL = 201495
Each news article contains it's url, title, text… See the full description on the dataset page: https://huggingface.co/datasets/0edon/KairosNews.
