datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsv3_upcoming_eventsEventBench
EventBench Dataset Access Instructions
EventBench 数据集访问说明
Notice / 通知
From 2026.5.20 to 2026.7.10, we are hosting EventBench competitions at @ECCV. To ensure fairness, the standard answers have been hidden during this period. We will make the standard answers available again after the competitions end.
If you would like to obtain evaluation results, please submit your predictions directly to the competition servers.
2026.5.20 至 2026.7.10 期间,我们在 @ECCV 部署了… See the full description on the dataset page: https://huggingface.co/datasets/XduSyL/EventBench.covid19_emergency_event
Dataset Card for EXCEPTIUS Corpus
Dataset Summary
This dataset presents a new corpus of legislative documents from 8 European countries (Beglium, France, Hunary, Italy, Netherlands, Norway, Poland, UK) in 7 languages (Dutch, English, French, Hungarian, Italian, Norwegian Bokmål, Polish) manually annotated for exceptional measures against COVID-19. The annotation was done on the sentence level.
Supported Tasks and Leaderboards
The dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/covid19_emergency_event.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.responsible-neobank-growth-events
Responsible Neobank Growth — Synthetic Event Benchmark
A synthetic dataset of neobank service events that misbehave on purpose — late,
duplicated, reversed, schema-evolving — with the correct answer known in
advance. It is built for testing incremental pipelines, data contracts,
referral-reward reconciliation, data quality and BI, where you want to check a
warehouse's output against a fixed truth rather than eyeball it.
Fully synthetic. No affiliation with Monzo or any bank; no… See the full description on the dataset page: https://huggingface.co/datasets/rosscyking/responsible-neobank-growth-events.EventGPT-datasetsHumAID-events
HumAID: Human-Annotated Disaster Incidents Data from Twitter
Dataset Summary
The HumAID Twitter dataset consists of several thousands of manually annotated tweets that has been collected during 19 major natural disaster events including earthquakes, hurricanes, wildfires, and floods, which happened from 2016 to 2019 across different parts of the World. The annotations in the provided datasets consists of following humanitarian categories. The dataset consists only english… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/HumAID-events.Event-Bench
Towards Event-oriented Long Video Understanding
[📖 arXiv Paper]
👀 Overview
We introduce Event-Bench, an event-oriented long video understanding benchmark built on existing datasets and human annotations. Event-Bench consists of three event understanding abilities and six event-related tasks, including 2,190 test instances to comprehensively evaluate the ability to understand video events.
Event-Bench provides a systematic comparison across different kinds of… See the full description on the dataset page: https://huggingface.co/datasets/Richard1999/Event-Bench.adaption-african-research-literature-current-events-qa
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-african-research-literature-current-events-qa
This dataset contains question-and-answer pairs covering diverse topics related to Africa, including literature, socio-economic development, current events, archaeology, and renewable energy. The content features detailed, informative responses that analyze contemporary trends, historical contexts, and regional challenges… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-research-literature-current-events-qa.wiki-eventspharma-serialized-events
ZigoTrace Pharma — Serialized Events (synthetic)
Feature vectors extracted from a synthetic DSCSA-style serialized medicine supply
chain, generated by packages/intelligence/src/synthetic.ts in the
zigo-pharma engine and exported via
hf/generate_fixtures.mjs. Used to train and validate the diversion-detection model
(zigotrace/pharma-authenticity-model).
⚠️ Synthetic data notice
This is entirely synthetic — a seeded generator (generateChain), not real
distributor or… See the full description on the dataset page: https://huggingface.co/datasets/AsamAce/pharma-serialized-events.foi-process-event-logs
Synthetic FOI Process Event Logs
Registry status
Registry ID: edithatogo/foi-process-event-logs
Family: foi-research
Repository role: synthetic_benchmark_dataset
Canonical dataset: edithatogo/foi-process-event-logs
Operational status: active
Rights status: synthetic_reviewed_fixtures
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository: https://github.com/edithatogo/foi-process
Upstream source:… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/foi-process-event-logs.NY_tech_week_partiful_events
NYC Tech Week 2026 - Partiful Events Dataset
A dataset of 1,561 events from NYC Tech Week 2026 (June 1-11), enriched with detailed event metadata from Partiful.
Dataset Summary
Metric
Count
Total events
1,561
Enriched with Partiful data
1,375
Events with Partiful links
1,379
Invite-only events
181
Date range
June 1 - June 11, 2026
NYC neighborhoods covered
28
Data Collection
Tech Week API… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/NY_tech_week_partiful_events.Event-Bench
Towards Event-oriented Long Video Understanding
[📖 arXiv Paper]
👀 Overview
We introduce Event-Bench, an event-oriented long video understanding benchmark built on existing datasets and human annotations. Event-Bench consists of three event understanding abilities and six event-related tasks, including 2,190 test instances to comprehensively evaluate the ability to understand video events.
Event-Bench provides a systematic comparison across different kinds of… See the full description on the dataset page: https://huggingface.co/datasets/RUCAIBox/Event-Bench.audio-event-triage-20260902-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260902-dataset.history-event-reconstruction
HISTORY-EVENT Reconstruction
An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies.
Configurations
Configuration
Rows
Purpose
events
2,361
All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.HumAID-event-type
HumAID: Human-Annotated Disaster Incidents Data from Twitter
Dataset Summary
The HumAID Twitter dataset consists of several thousands of manually annotated tweets that has been collected during 19 major natural disaster events including earthquakes, hurricanes, wildfires, and floods, which happened from 2016 to 2019 across different parts of the World. The annotations in the provided datasets consists of following humanitarian categories. The dataset consists only english… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/HumAID-event-type.audio-event-triage-20260912-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260912-dataset.event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.dataset-trust-auditor-events
Dataset Trust Auditor — Audit Events
Public audit trail produced by the Dataset Trust Auditor — a two-phase AI pipeline that scores HuggingFace datasets across 8 trust dimensions.
Every audit run appends one row. The dataset grows over time as users audit datasets through the deployed app.
Dataset Structure
Each row is one completed audit of a HuggingFace dataset.
Column
Type
Description
audit_id
string
UUID for this audit run
url
string
Full HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/nicolas-brieuc/dataset-trust-auditor-events.ukr-wiki-events
Ukrainian Wikipedia events
A small (1,722-row) Ukrainian dataset built from public-domain /
Wikipedia-sourced text. Two task shapes are mixed in the single train
split (distinguishable via the instruction prompt):
Event extraction — instruction = a passage of Ukrainian Wikipedia
text prefixed by "what important event is this text about:";
output = a short label of the salient event.
Explanation / QA — instruction = a question or term (e.g.
"Опиши явище поліплоїдії"); output =… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-wiki-events.2025_events
2025
Contains all the world event knowledge of 2025
audio-event-triage-20260803-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260803-dataset.audio-event-triage-20260922-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260922-dataset.audio-event-triage-20260823-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260823-dataset.case_event_typefew-event
FewEvent (mneb/casie format)
FewEvent (Deng et al., WSDM 2020) converted into the mneb json_structures event-extraction
format. FewEvent annotates two tasks — event detection (triggers) and event argument
extraction (roles); it has no entity or relation layer, so each record's output carries
only json_structures. Char offsets are end-exclusive (input[start:end] == text).
One record = one sentence = one event. Single test split per config.
The two configs (split by… See the full description on the dataset page: https://huggingface.co/datasets/mneb/few-event.han-humanoid-safety-event-logs-v1
Humanoid Safety Event Logs (HSEL)
Description
This dataset contains structured safety-related
events recorded during humanoid operations.
Features
event_id
task_type
obstacle_proximity_index
joint_temperature
vibration_level
speed_m_s
safety_alert (boolean)
Target
safety_alert
Use Cases
Safety alert prediction
Risk modeling
Preventive monitoring
Evaluation Metrics
Accuracy
F1 Score
ROC-AUC
License
MIT
2023_events
2023
Contains all the world event knowledge of 2023
2024_events
2024
Contains all the world event knowledge of 2024
