datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Kairos
Kairos — Long-Form Video Annotation and Benchmark
Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes).
This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form
(OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and
5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s).
What's inside
.
├── data/
│ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.KAIROS_EVAL
KAIROS_EVAL Dataset
Paper: LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions | Code (GitHub)
Dataset Summary
KAIROS is a benchmark dataset designed to evaluate the robustness of large language models (LLMs) in multi-agent, socially interactive scenarios. Unlike static QA datasets, KAIROS dynamically constructs evaluation settings for each model by capturing its original belief (answer + confidence) and then simulating peer influence through… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/KAIROS_EVAL.KairosQA
KairosQA Dataset
Dataset Description
KairosQA is a temporally grounded question-answering dataset designed to evaluate the temporal alignment and reasoning capabilities of Large Language Models (LLMs). Unlike static benchmarks, KairosQA focuses on facts that evolve over time, specifically subject–relation–object triplets from Wikidata that changed at least twice between 2018 and 2025 as described in our paper Understanding Data Temporality Impact on Large Language… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/KairosQA.kairos-routing
Kairos Routing
Kairos Routing is a long-format dataset for training a model to select the
best AI model for a prompt. Each row describes one candidate model evaluated
on one prompt.
The intended learning problem is:
prompt + candidate_model_vector -> score
The router can score several candidate models for the same prompt and select
the model with the highest predicted score.
Dataset Summary
Approximately 1.92 million rows
Approximately 144,000 unique prompts
64… See the full description on the dataset page: https://huggingface.co/datasets/sijirama/kairos-routing.MAG188
MAG188
MAG188 is a DFT benchmark of 188 experimentally characterized collinear magnetic materials from the
MAGNDATA database. All structures were recomputed with spin-polarized DFT and relaxed. MAG188 contains
two benchmarks:
MAG188-EXP: the relaxed experimentally reported magnetic state of each material (188 endpoints).
It measures accuracy on experimentally established magnetic states.
MAG188-SAMPLE: 1,195 relaxed endpoints that cover several self-consistent collinear spin… See the full description on the dataset page: https://huggingface.co/datasets/kairosmaterial/MAG188.srl_datasets_text2text_sampleKairosNewsHandcrafted Dataset used in the elaboration of a thesis and a project for a competition (Premio Arquivo.pt: https://sobre.arquivo.pt/pt/colabore/premios-arquivo-pt/premio-arquivo-pt-2025/)
It contains news articles from the following Portuguese News agencies from 2020 to 2024:
https://www.cmjornal.pt/ = 6771
https://expresso.pt/ = 22606
https://www.iol.pt/ = 39387
https://www.publico.pt/ = 74893
https://www.sapo.pt/ = 57838
TOTAL = 201495
Each news article contains it's url, title, text… See the full description on the dataset page: https://huggingface.co/datasets/0edon/KairosNews.srl_datasets_mentions_samplepropbank_srl_seq2seqkairos-datasetconll05_mentions_samplekairos-dataset-v1bolt_mentions_sampleontonotes_mentions_sample5190-news-source
