datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multinerd
Dataset Card for MultiNERD dataset
Description
Summary: In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/multinerd.ALERT
Dataset Card for the ALERT Benchmark
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/felixludos/babel-briefings.BabelRSALERT_DPO
Dataset Card for the ALERT DPO Dataset
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT_DPO.cner
Dataset Card for CNER dataset
Description
Summary: Concept and Named Entity Recognition (CNER) is a novel task that jointly handles the indentification and classification of concepts and named entities.
Repository: https://github.com/Babelscape/cner
Paper: CNER: Concept and Named Entity Recognition
Point of Contact: {martinelli, molfese, tedeschi, navigli}@diag.uniroma1.it
Dataset Structure
The data fields are the same among all splits.
tokens: a list… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/cner.PDDL2PRM
PDDL2PRM: Planning-Based Step-Level Supervision for Process Reward Models
PDDL2PRM is a large-scale dataset for training and evaluating Process Reward Models (PRMs) with fine-grained, step-level supervision derived from symbolic planning problems.
Unlike many PRM datasets that rely on human annotation, LLM judges, or final-answer correctness, PDDL2PRM uses Planning Domain Definition Language (PDDL) problems to generate structured reasoning trajectories whose intermediate steps… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/PDDL2PRM.story-summeval
Dataset Card for Story-SummEval
Dataset Description
For a thorough description of the data creation please refer to the ACL 2024 paper:
"FENICE: Factuality Evaluation of summarization based on NLI and Claim Extraction", Scirè et al. (2024).
Summary
This dataset contains summaries of stories from Gutenberg and Wikisource along with their factuality labels.
Summaries are generated from several models provided by the paper "Echoes from Alexandria" by Scirè et al.… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/story-summeval.BabelTower
BabelTower Dataset
Dataset Description
✨ Update:
[2025-09] - QiMeng-MuPa has been accepted to NeurIPS 2025 🎉
[2025.08] - We have released Qwen3-0.6B-translator, welcome to try!
BabelTower is a paired corpora for translating C to CUDA, featuring 233 pairs of functionally aligned C and CUDA programs with tests. It can evaluate the ability of language models to convert sequential programs into parallel counterparts.
Results
These results are from… See the full description on the dataset page: https://huggingface.co/datasets/kcxain/BabelTower.wsl
Word Sense Linking Dataset
Description
The dataset is designed for the task of Word Sense Linking (WSL), where systems are required to identify and disambiguate spans of text to their most suitable senses from a reference inventory. The annotations are provided as sense keys from WordNet, a large lexical database of English. All terms in the dataset have been fully manually annotated.The dataset sentences are taken from the ALL split of the Raganato framework for… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/wsl.Babelscape_ALERT_selected_short_entries
