datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rebel-datasetREBEL is a silver dataset created for the paper REBEL: Relation Extraction By End-to-end Language generationALERT
Dataset Card for the ALERT Benchmark
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/felixludos/babel-briefings.ALERT_DPO
Dataset Card for the ALERT DPO Dataset
Description
Paper Summary: When building Large Language Models (LLMs), it is paramount to bear safety in mind and protect them with guardrails. Indeed, LLMs should never generate content promoting or normalizing harmful, illegal, or unethical behavior that may contribute to harm to individuals or society. In response to this critical challenge, we introduce ALERT, a large-scale benchmark to assess the safety of LLMs through red… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/ALERT_DPO.babel-briefings
Babel Briefings News Headlines Dataset README
Break Free from the Language Barrier
Version: 1 - Date: 30 Oct 2023
Collected and Prepared by Felix Leeb (Max Planck Institute for Intelligent Systems, Tübingen, Germany)
License: Babel Briefings Headlines Dataset © 2023 by Felix Leeb is licensed under CC BY-NC-SA 4.0
Check out our paper on arxiv.
This dataset contains 4,719,199 news headlines across 30 different languages collected between 8 August 2020 and 29 November 2021. The… See the full description on the dataset page: https://huggingface.co/datasets/Tribhuvand/babel-briefings.OpenNCEE-ChineseA database of NCEE(a.k.a. Gaokao) Chinese problems.
No commercial use, or you ignore the risk of legal (especially China Mainland).
