datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.task368_synthetic_even_or_odd_calculation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task368_synthetic_even_or_odd_calculation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task368_synthetic_even_or_odd_calculation.task1595_event2mind_text_generation_1
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1595_event2mind_text_generation_1
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1595_event2mind_text_generation_1.wiki-events-cpt
Wikipedia Events CPT
jhdlee/wiki-events-cpt is a public research dataset of 150 selected English Wikipedia articles with compact metadata for continual pretraining (CPT).
Split
Articles
Event window (end exclusive)
cohort_a
75
2023-01-01 to 2024-10-01
cohort_b
75
2024-10-01 to 2025-09-01
Each cohort has 25 articles per topic: natural_hazards, elections, and sports. Cohorts group events by their reviewed whole-occurrence intervals; they are not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-events-cpt.task1596_event2mind_text_generation_2
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1596_event2mind_text_generation_2
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1596_event2mind_text_generation_2.task614_glucose_cause_event_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task614_glucose_cause_event_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task614_glucose_cause_event_detection.task1495_adverse_drug_event_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1495_adverse_drug_event_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1495_adverse_drug_event_classification.task924_event2mind_word_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task924_event2mind_word_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task924_event2mind_word_generation.task923_event2mind_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task923_event2mind_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task923_event2mind_classifier.events-scheduling
🗓️ Events Scheduling dataset
Small dataset to train Language Models to create a schedule from a list of events and priorities.
I used this dataset to train the 👑 🗓️ anakin87/qwen-scheduler-7b-grpo model using GRPO.
➡️ Read the full story in my blog post.
Find all the code in the GitHub repository.
The problem
Given a list of events and priorities, we ask the model to create a schedule that maximizes the total duration of selected events, weighted by priority.In… See the full description on the dataset page: https://huggingface.co/datasets/anakin87/events-scheduling.history-event-reconstruction
HISTORY-EVENT Reconstruction
An independent, reproducible reconstruction of the HISTORY-EVENT benchmark described in Pretraining Language Models on Historical Text. This is not the authors' official dataset. Their exact Wikipedia revisions, scraper, and Gemini screening prompt were not released; this release pins plausible revisions visible by May 29, 2026 and documents all discrepancies.
Configurations
Configuration
Rows
Purpose
events
2,361
All… See the full description on the dataset page: https://huggingface.co/datasets/jbduran/history-event-reconstruction.event-outcome-resolution
EVENT_OUTCOME_RESOLUTION
A preference dataset for EVENT_OUTCOME_RESOLUTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from… See the full description on the dataset page: https://huggingface.co/datasets/316usman/event-outcome-resolution.even-russian-instructions
Even Russian Instructions
Инструкционный датасет для fine-tuning VLM-моделей (Qwen3-VL) на эвенском языке — критически исчезающем тунгусо-маньчжурском языке Сибири.Примеры сгенерировано путем синтеза шаблонным методом на основе параллельного русско-эвенского корпуса.
Структура датасета
Три предварительно разделённых сплита для curriculum learning:
Сплит
Примеров
Размер
train
96 327
2.6 MB
val
5 665
175 KB
test
11 331
341 KB
Поля… See the full description on the dataset page: https://huggingface.co/datasets/DaniilMako/even-russian-instructions.event-bench
Event Bench
29-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as an event planning assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as an event planning assistant managing venue bookings, catering, and guest logistics. The conversation features cascading changes — a venue switch triggers catering… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/event-bench.anp2-events
ANP2 public event log — Phase 0/1 bootstrap snapshot
Historical snapshot of all public, Ed25519-signed events from the reference relay of ANP2 — the open economic protocol for AI agents. Taken 2026-05-24, immediately before the reference relay underwent a fresh-restart migration. The current live https://anp2.com/api/events returns a different population than this snapshot — this archive is preserved as a Phase 0/1 bootstrap record for researchers studying the early-bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/anp2/anp2-events.future-news-events-2026
Future News Events — 2026 QA
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A question-answering dataset built on real-world 2026 news events scraped
from the Wikipedia Current Events Portal. Each event was passed to
Cohere Command R with RAG grounding, which generated diverse, factual
question–answer pairs covering what happened, who/where, and
causes/consequences/context.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/future-news-events-2026.corporate-event-detection
The dataset
The dataset is designed for corporate event detection and text-based stock prediction benchmark. It includes 9721 news articles with token-level event labels and 303893 news articles with minute-level timestamps and comprehensive stock price labels.
Detail Information
EDT contains data for three purposes: 1. corporate event detection; 2. news-based trading strategy benchmark; 3. financial domain adaptation.
1. Corporate Event Detection
EDT… See the full description on the dataset page: https://huggingface.co/datasets/agungpambudi/corporate-event-detection.russian_events_vectors
Description in English:
The dataset is collected from Russian-language Telegram channels with information about various events and happenings in Russian regions,The dataset was collected and tagged automatically using the data collection and tagging service Scoutie.Try Scoutie and collect the same or another dataset using link for FREE.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_events_vectors.task851_synthetic_multiply_evens
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task851_synthetic_multiply_evens
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task851_synthetic_multiply_evens.task205_remove_even_elements
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task205_remove_even_elements
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task205_remove_even_elements.even-russian-corpus
Even-Russian Corpus | Эвенско-русский корпус
Эвенско-русский двуязычный текстовый корпус — параллельный корпус из примерно 50 пар предложений и словосочетаний на эвенском и русском языках. Датасет предназначен для исследований в области машинного перевода, обработки естественного языка и сохранения исчезающего эвенского языка (Тунгусо-маньчжурская семья -> Северная (сибирская) ветвь -> Эвенская группа -> Эвенский язык).
О датасете
Эвенский язык (Эвэды төрэн) —… See the full description on the dataset page: https://huggingface.co/datasets/DaniilMako/even-russian-corpus.task922_event2mind_word_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task922_event2mind_word_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task922_event2mind_word_generation.event_timelines_dataset
