datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.events_classification_biotech
Key aspects
Event extraction;
Multi-label classification;
Biotech news domain;
31 classes;
3140 total number of examples;
Motivation
Text classification is a widespread task and a foundational step in numerous information extraction pipelines. However, a notable challenge in current NLP research lies in the oversimplification of benchmarking datasets, which predominantly focus on rudimentary tasks such as topic classification or sentiment analysis.
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/events_classification_biotech.wiki-events-cpt
Wikipedia Events CPT
jhdlee/wiki-events-cpt is a public research dataset of 150 selected English Wikipedia articles with compact metadata for continual pretraining (CPT).
Split
Articles
Event window (end exclusive)
cohort_a
75
2023-01-01 to 2024-10-01
cohort_b
75
2024-10-01 to 2025-09-01
Each cohort has 25 articles per topic: natural_hazards, elections, and sports. Cohorts group events by their reviewed whole-occurrence intervals; they are not… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/wiki-events-cpt.ukr-wiki-events
Ukrainian Wikipedia events
A small (1,722-row) Ukrainian dataset built from public-domain /
Wikipedia-sourced text. Two task shapes are mixed in the single train
split (distinguishable via the instruction prompt):
Event extraction — instruction = a passage of Ukrainian Wikipedia
text prefixed by "what important event is this text about:";
output = a short label of the salient event.
Explanation / QA — instruction = a question or term (e.g.
"Опиши явище поліплоїдії"); output =… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/ukr-wiki-events.2025_events
2025
Contains all the world event knowledge of 2025
2023_events
2023
Contains all the world event knowledge of 2023
future-news-events-2026
Future News Events — 2026 QA
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A question-answering dataset built on real-world 2026 news events scraped
from the Wikipedia Current Events Portal. Each event was passed to
Cohere Command R with RAG grounding, which generated diverse, factual
question–answer pairs covering what happened, who/where, and
causes/consequences/context.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/future-news-events-2026.2024_events
2024
Contains all the world event knowledge of 2024
Eventloglongcovid-risk-eventtimeseries
Citation
If you find this dataset or our work useful in your research, please consider citing:
Jing Wang, Amar Sra, Jeremy C. Weiss. Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2. arXiv:2506.22444, 2025.
BibTeX:
@misc{longcovid,
title = {Active Learning for Forecasting Severity among Patients with Post Acute Sequelae of SARS-CoV-2},
author = {Jing Wang and Amar Sra and Jeremy C. Weiss},
year = {2025},
eprint = {2506.22444}… See the full description on the dataset page: https://huggingface.co/datasets/juliawang2024/longcovid-risk-eventtimeseries.
