datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
publichealth-qa
Usage
import datasets
langs = ['arabic', 'chinese', 'english', 'french', 'korean', 'russian', 'spanish', 'vietnamese']
data = datasets.load_dataset('xhluca/publichealth-qa', split='test', name=langs[0])
About
This dataset contains question and answer pairs sourced from Q&A pages and FAQs from CDC and WHO pertaining to COVID-19. They were produced and collected between 2019-12 and 2020-04. They were originally published as an aggregated Kaggle dataset.… See the full description on the dataset page: https://huggingface.co/datasets/xhluca/publichealth-qa.xtr-wiki_qa
Xtr-WikiQA
Dataset Summary
Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages.
This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face).
For translations, we used Amazon Translate.
Languages
Arabic (ar)
Spanish (es)
French (fr)
German (de)
Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.ntu_adl_questionxai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
FinanceQAFinanceQA is a comprehensive testing suite designed to evaluate LLMs' performance on complex financial analysis tasks that mirror real-world investment work. The dataset aims to be substantially more challenging and practical than existing financial benchmarks, focusing on tasks that require precise calculations and professional judgment.
Paper: https://arxiv.org/abs/2501.18062
Description
The dataset contains two main categories of questions:
Tactical Questions: Questions based on… See the full description on the dataset page: https://huggingface.co/datasets/Joshua-Xia/FinanceQA.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
sino-xenic-reasoning-gap-dataset
Sino-Xenic Reasoning Gap Dataset
A comprehensive evaluation dataset for testing Large Language Models' understanding of Sino-Xenic linguistic phenomena across Chinese, Japanese, Korean, and Vietnamese.
Dataset Overview
Total Samples: 297
Languages: Chinese, Japanese, Korean, Vietnamese
Categories: 11
Task Types: Surface-level and Deep Structural
Categories
Chinese Idioms (27 samples) - Understanding Chinese idioms and their cultural meanings
Chinese… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/sino-xenic-reasoning-gap-dataset.mauxi-COT-Persian
🧠 mauxi-COT-Persian Dataset
Exploring Persian Chain-of-Thought Reasoning with DeepSeek-R1, brought to you by Mauxi AI Platform
🌟 Overview
mauxi-COT-Persian is a community-driven dataset that explores the capabilities of advanced language models in generating Persian Chain-of-Thought (CoT) reasoning. The dataset is actively growing with new high-quality, human-validated entries being added regularly. I am personally working on expanding this dataset with rigorously… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/mauxi-COT-Persian.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.Xrax
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/Abafdon22825/Xrax.xcopa
XCOPA – Galician, Swahili & Urdu
Machine-translated Galician, Swahili, and Urdu subsets of the Cross-lingual Choice of Plausible Alternatives (XCOPA) benchmark. This dataset was translated using Google Machine Translate.
Dataset Description
XCOPA is a multilingual causal commonsense reasoning benchmark. Given a premise and a question (asking for the cause or effect), the task is to choose the more plausible alternative from two choices. This repository contains Galician… See the full description on the dataset page: https://huggingface.co/datasets/Owos/xcopa.plm-qarwq
