datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BabelDOC-Assets
BabelDOC-Assets
Font and other resource files relied on by BabelDOC and pdf2zh
BabelDOC is a PDF translation library, pdf2zh is a PDF translation tool.
Fonts and Licenses
Go Noto Universal: THE UNLICENSE
Pal3love/Source-Han-TrueType: SIL OPEN FONT LICENSE Version 1.1
lxgw/LxgwWenKaiGB: OFL-1.1 License
lxgw/LxgwWenkaiTC: OFL-1.1 License
fontworks-fonts/Klee: OFL-1.1 License
fonts-archive/MaruBuri: License
Noto Serif/Noto Sans: SIL OPEN FONT LICENSE Version 1.1… See the full description on the dataset page: https://huggingface.co/datasets/awwaawwa/BabelDOC-Assets.wikineural
Dataset Card for WikiNEuRal dataset
Description
Summary: In a nutshell, WikiNEuRal consists in a novel technique which builds upon a multilingual lexical knowledge base (i.e., BabelNet) and transformer-based architectures (i.e., BERT) to produce high-quality annotations for multilingual NER. It shows consistent improvements of up to 6 span-based F1-score points against state-of-the-art alternative data production methods on common benchmarks for NER. We used this… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/wikineural.multinerd
Dataset Card for MultiNERD dataset
Description
Summary: In a nutshell, MultiNERD is the first language-agnostic methodology for automatically creating multilingual, multi-genre and fine-grained annotations for Named Entity Recognition and Entity Disambiguation. Specifically, it can be seen an extension of the combination of two prior works from our research group that are WikiNEuRal, from which we took inspiration for the state-of-the-art silver-data creation methodology… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/multinerd.babeldoc-temp-pdfsrebel-datasetREBEL is a silver dataset created for the paper REBEL: Relation Extraction By End-to-end Language generationSREDFMRelation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. \In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems.
First, we present SRED\textsuperscript{FM}, an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED\textsuperscript{FM}, a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL,
that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at \href{https://www.github.com/babelscape/rebel}{https://www.github.com/babelscape/rebel}.ALERT
Dataset Card for the ALERT Benchmark
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/felixludos/babel-briefings.REDFMRelation Extraction (RE) is a task that identifies relationships between entities in a text, enabling the acquisition of relational facts and bridging the gap between natural language and structured knowledge. However, current RE models often rely on small datasets with low coverage of relation types, particularly when working with languages other than English. \In this paper, we address the above issue and provide two new resources that enable the training and evaluation of multilingual RE systems.
First, we present SRED\textsuperscript{FM}, an automatically annotated dataset covering 18 languages, 400 relation types, 13 entity types, totaling more than 40 million triplet instances. Second, we propose RED\textsuperscript{FM}, a smaller, human-revised dataset for seven languages that allows for the evaluation of multilingual RE systems.
To demonstrate the utility of these novel datasets, we experiment with the first end-to-end multilingual RE model, mREBEL,
that extracts triplets, including entity types, in multiple languages. We release our resources and model checkpoints at \href{https://www.github.com/babelscape/rebel}{https://www.github.com/babelscape/rebel}.babel-official
BABEL Official Local Layout
This directory is a cleaned local mirror of the official BABEL v1.0 labels and
the AMASS subsets required by those labels.
Layout
archives/: original downloaded archives, kept unchanged for provenance.
labels/babel_v1.0_release/: official BABEL JSON splits.
amass/: extracted AMASS motion parameter files.
processed/manifests/*.jsonl: normalized records with resolved local AMASS
paths, frame counts, fps, segment labels, and rewritten… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/babel-official.BabelRSIARPA_BABEL_OP3_306272-dim-BABEL-stream
🚀 Dataset Usage
To facilitate researchers, we provide the processed streaming 272-dim Motion Representation of BABEL dataset in this Hugging Face repo.
NOTE: We process the original BABEL dataset to support training of streaming motion generation.
e.g. If there is a motion sequence A, annotated as (A1, A2, A3, A4) in BABEL dataset, each subsequence has text description: (A1_t, A2_t, A3_t, A4_t).
Then, our BABEL-stream is constructed as:
seq1: (A1, A2) --- seq1_text:… See the full description on the dataset page: https://huggingface.co/datasets/lxxiao/272-dim-BABEL-stream.272-dim-BABEL
🚀 Dataset Usage
To facilitate researchers, we provide the processed 272-dim Motion Representation of BABEL dataset in this Hugging Face repo.
Motions are resampled into 30 FPS.
NOTE: t2m_babel_mean_std/ contains the joint mean and std of both HumanML3D and BABEL dataset for joint training of the proposed Causal TAE.
❗️❗️❗️ The processed data is solely for academic purposes. Make sure you read through the BABEL License.
📖 Paper & Project Page & Code
Arxiv Paper… See the full description on the dataset page: https://huggingface.co/datasets/lxxiao/272-dim-BABEL.LLM-Oasis_unfactual_text_generation
Babelscape/LLM-Oasis_unfactual_text_generation
Dataset Description
LLM-Oasis_unfactual_text_generation is part of the LLM-Oasis suite and contains unfactual texts generated from a set of falsified claims extracted from a Wikipedia passage and its paraphrase.
This dataset corresponds to the unfactual text generation step described in Section 3.4 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_unfactual_text_generation.OpenNCEE-Chinese-EssayALERT_DPO
Dataset Card for the ALERT DPO Dataset
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT_DPO.cner
Dataset Card for CNER dataset
Description
Summary: Concept and Named Entity Recognition (CNER) is a novel task that jointly handles the indentification and classification of concepts and named entities.
Repository: https://github.com/Babelscape/cner
Paper: CNER: Concept and Named Entity Recognition
Point of Contact: {martinelli, molfese, tedeschi, navigli}@diag.uniroma1.it
Dataset Structure
The data fields are the same among all splits.
tokens: a list… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/cner.LLM-Oasis_paraphrase_generation
Babelscape/LLM-Oasis_paraphrase_generation
Dataset Description
LLM-Oasis_paraphrase_generation is part of the LLM-Oasis suite and contains paraphrases generated from a set of claims extracted from a Wikipedia passage.
This dataset supports the paraphrase generation step described in Section 3.3 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The title… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_paraphrase_generation.LLM-of-Babel-NL2LLM-Oasis_claim_extraction
Babelscape/LLM-Oasis_claim_extraction
Dataset Description
LLM-Oasis_claim_extraction is part of the LLM-Oasis suite and contains text-claim pairs extracted from Wikipedia pages.
It provides the data used to train the claim extraction system described in Section 3.1 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The title of the Wikipedia page.
text: A… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_extraction.50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset
BabelSpeech: 50 Hours of Real-World Colloquial Malay ASR Speech Data
This dataset contains 50 hours of high-quality Malay colloquial ASR speech data, reflecting realistic code-switching between Malay and English, as commonly used in everyday communication in Malaysia.
Overview
Content: 50 hours of real-world colloquial Malay speech suitable for ASR fine-tuning and benchmarking.
Metadata: Stored in a separate JSON file, including audio path, duration, text, confidence… See the full description on the dataset page: https://huggingface.co/datasets/BabelSpeech/50hours_Malay_Real-world_Colloquial_Conversation_and_Monologue_Speech_Dataset.AMASS_BABELPDDL2PRM
PDDL2PRM: Planning-Based Step-Level Supervision for Process Reward Models
PDDL2PRM is a large-scale dataset for training and evaluating Process Reward Models (PRMs) with fine-grained, step-level supervision derived from symbolic planning problems.
Unlike many PRM datasets that rely on human annotation, LLM judges, or final-answer correctness, PDDL2PRM uses Planning Domain Definition Language (PDDL) problems to generate structured reasoning trajectories whose intermediate steps… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/PDDL2PRM.LLM-Oasis_claim_falsification
Babelscape/LLM-Oasis_claim_falsification
Dataset Description
LLM-Oasis_claim_falsification is part of the LLM-Oasis suite and contains the outcomes of the claim falsification process.
This dataset provides pairs of factual and falsified claims from a given Wikipedia text as described in Section 3.2 of the LLM-Oasis paper. Please refer to our GitHub repository for more information on the overall data generation pipeline of LLM-Oasis.
Features
title: The… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/LLM-Oasis_claim_falsification.babeledits
BabelEdits
BabelEdits is a benchmark designed to evaluate cross-lingual knowledge editing (CKE) in Large Language Models (LLMs). It enables robust and effective evaluation across 60 languages by combining high-quality entity translations from BabelNet with marker-based translation. BabelEdits is also accompanied by a modular CKE method, BabelReFT, which supports multilingual edit propagation while preserving downstream model performance.
Dataset Summary
As LLMs… See the full description on the dataset page: https://huggingface.co/datasets/umanlp/babeledits.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/Tribhuvand/babel-briefings.Babelscape-wikineural-joined
Dataset Card for "Babelscape-wikineural-joined"
This dataset is a merged version of wikineural
More Information needed
@inproceedings{tedeschi-etal-2021-wikineural-combined,
title = "{W}iki{NE}u{R}al: {C}ombined Neural and Knowledge-based Silver Data Creation for Multilingual {NER}",
author = "Tedeschi, Simone and
Maiorca, Valentino and
Campolungo, Niccol{\`o} and
Cecconi, Francesco and
Navigli, Roberto",
booktitle = "Findings of the… See the full description on the dataset page: https://huggingface.co/datasets/dmargutierrez/Babelscape-wikineural-joined.OpenNCEE-ChineseA database of NCEE(a.k.a. Gaokao) Chinese problems.
No commercial use, or you ignore the risk of legal (especially China Mainland).
story-summeval
Dataset Card for Story-SummEval
Dataset Description
For a thorough description of the data creation please refer to the ACL 2024 paper:
"FENICE: Factuality Evaluation of summarization based on NLI and Claim Extraction", Scirè et al. (2024).
Summary
This dataset contains summaries of stories from Gutenberg and Wikisource along with their factuality labels.
Summaries are generated from several models provided by the paper "Echoes from Alexandria" by Scirè et al.… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/story-summeval.
