datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.llm-output-evaluation
LLM_OUTPUT_EVALUATION
A preference dataset for LLM_OUTPUT_EVALUATION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/llm-output-evaluation.synthetic-llm-evaluation-traces
EvaluLLM-Inspired Synthetic Evaluation Traces (DBbun)
View source code on GitHub
Watch on Youtube: Evaluating AI with AI
Dataset Summary
This dataset contains fully synthetic evaluation traces for NLG / LLM-style output comparison, inspired by the evaluation workflow described in EvaluLLM: LLM Assisted Evaluation of Generative Outputs (IUI Companion 2024).
The dataset is produced by a configurable, offline simulator and includes:
synthetic tasks (prompts)
synthetic… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/synthetic-llm-evaluation-traces.llm-evaluation-sft-100k
LLM Evaluation SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering LLM evaluation methodology, benchmarking, safety assessment, and prompt optimization. Designed to train AI assistants that can help ML engineers and researchers rigorously evaluate and improve language models.
Dataset Description
This dataset covers the full spectrum of LLM evaluation practice across 7 specialized categories. Each record follows the… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/llm-evaluation-sft-100k.llm-evaluation-huji
DOVE
