datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompt_theft
Dataset Card for prompt theft library
Dataset Details
Dataset Description
This dataset targets research and testing regarding prompt-theft attacks targeting system prompt leakage with generative AI models as conversational applications. Therefore, prompt theft attacks were collected from the literature and enhanced by paraphrasing the collected attacks. 17 prompt theft attacks were collected from literature and enhanced by paraphrasing to 68… See the full description on the dataset page: https://huggingface.co/datasets/LilianDK/prompt_theft.HLE-BioMedX
HLE-BioMedX — Multilingual HLE Biology/Medicine
A multilingual version of the Biology/Medicine subset of Humanity's Last Exam
(HLE), released as one subset per language.
Source benchmark: Humanity's Last Exam, dataset
cais/hle.
Subsets
Group
Languages
How the target-language text was produced
Source
en
Original English questions and answers.
Machine-translated and expert-verified / revised
zh, ja, ko, fr, th
Machine translation reviewed by a human… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HLE-BioMedX.HealMed
HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems.
The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.ufo
Official UAP Disclosure Corpus
Page-level text of the recent official U.S. government UAP (Unidentified Anomalous Phenomena)
disclosure, extracted two ways and published side by side for comparison. 789 documents,
7,288 pages, drawn only from official releases — no third-party or speculative material.
Two extraction variants
directory
how it was extracted
data-ocr/
forced OCR — the GLM-OCR vision model re-reads every page
data-no-ocr/
native… See the full description on the dataset page: https://huggingface.co/datasets/lilbeee/ufo.respect
Retrospective Learning from Interactions (Respect) Dataset
This repository contains the lil-lab/respect data, based on the ACL paper Retrospective Learning from Interactions. For more resources, please see https://lil-lab.github.io/respect and https://github.com/lil-lab/respect.
Sample Usage
You can load the data and associated checkpoints as follows:
from datasets import load_dataset
from transformers import Idefics2ForConditionalGeneration
from peft importPeftModel… See the full description on the dataset page: https://huggingface.co/datasets/lil-lab/respect.JPMedReason
JPMedReason 🇯🇵🧠
Japanese Medical Reasoning Dataset (Translation of MedReason)
📝 Overview
JPMedReason is a high-quality Japanese translation of the MedReason dataset — a benchmark designed to evaluate complex clinical reasoning capabilities of Large Language Models (LLMs) in multiple-choice QA format.
This dataset provides question-answer pairs in the medical domain accompanied by chain-of-thought (CoT) style reasoning in both English and Japanese. The Japanese… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/JPMedReason.lilium_albanicum_eng_alb
Lilium Albanicum Eng-Alb
Task Categories:
Translation
Question-Answering
Conversational
Languages: English (en), Albanian (sq)
Size Categories: 100K < n < 1M
Dataset Card for "Lilium Albanicum"
Dataset Summary
The Lilium Albanicum dataset is a comprehensive English-Albanian and Albanian-English parallel corpus. The dataset includes original translations and extended synthetic Q&A pairs, which are designed to support and optimize LLM translation… See the full description on the dataset page: https://huggingface.co/datasets/noxneural/lilium_albanicum_eng_alb.HealthBench-ProX
HealthBench-ProX
Dataset Description
HealthBench-ProX is a multilingual extension of the original HealthBench Professional benchmark for evaluating large language models on realistic healthcare consultation scenarios.
The dataset contains 6,825 evaluation instances organized into 13 language-specific splits:
de (German)
en (English)
fr (French)
hi (Hindi)
ig (Igbo)
ja (Japanese)
ko (Korean)
ms (Malay)
pt (Portuguese)
sw (Swahili)
th (Thai)
zh (Chinese)
zu (Zulu)… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealthBench-ProX.HealthBenchX
HealthBenchX: Machine Translation of HealthBench Conversations based on GPT-5
Overview
HealthBenchXP is a machine-translated version of HealthBench, a benchmark designed to measure the capabilities of AI systems in health-related tasks.
The original HealthBench was built in collaboration with 262 physicians across 60 countries and includes 5,000 realistic health conversations, each paired with a custom physician-created rubric for evaluation.
Our contribution: we… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealthBenchX.LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
🌐 Project Page: https://longbench2.github.io
💻 Github Repo: https://github.com/THUDM/LongBench
📚 Arxiv Paper: https://arxiv.org/abs/2412.15204
LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/lillycyx/LongBench-v2.satai
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/lilsomnus/satai.
