datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hiring-bias-mitigation-responses
Hiring-bias mitigation — model responses
Every response produced in the mitigation study of LLM hiring decisions: 54 runs,
2,471,850 responses, from 5 open-weight models in English and Ukrainian, at
baseline and under each mitigation family (baseline, embedding, prompt, scrub). Each run is one subset.
All released artifacts: the Hiring Bias Mitigation collection.
Training data of the fine-tuned runs: hiring-bias-mitigation-synthetic-data.
Code, configs, full results and… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/hiring-bias-mitigation-responses.dflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.imo-answerbench-responses
IMO-AnswerBench Responses
This dataset contains generated model responses for the IMO-AnswerBench benchmark.
The files are organized by model under responses/. Each problem has two files:
<problem_id>.txt: human-readable prompt, problem metadata, generated response, and gold short answer.
<problem_id>.json: structured metadata, prompt text, generated text, token IDs, token counts, and storage paths.
Current contents:
responses/WeiboAI__VibeThinker-3B: 400 IMO-AnswerBench… See the full description on the dataset page: https://huggingface.co/datasets/sida/imo-answerbench-responses.task362_spolin_yesand_prompt_response_sub_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.dementor-matrix-responses
Dementor — matrix model responses
Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study.
Companion to:
Code + prompt splits: https://github.com/lisadunlap/dementor (branch ethan)
Trained adapters (2,122 LoRAs): https://huggingface.co/dementor-research — SFT / DPO /
self-SFT, grouped into per-dataset collections (gsm8k, chatbot_arena, writingprompts, openassistant).
Dataset viewer. This repo is a nested tree of CSV tables plus per-cell cell.json… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-matrix-responses.wjb_responsesFour LLMs' responses to safety-related prompts (attacked and vanilla) from wildjailbreak
We use LLama-Guard-3 to classify the safety of the response
tiny-ua-bench-responses
Tiny-UA-Bench Responses
This dataset contains the response matrix for Tiny-UA-Bench.
The matrix contains 919,160 model and item records.
The matrix covers 20 models and 45,958 items.
The evaluation excludes FLORES and LongFLORES.
Use
Use this dataset to reproduce the benchmark compression analysis.
Do not use a held-out model response to fit a selector or predictor.
Use the reference and held-out split definitions from the code repository.
Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.alignment-veto-responses
Alignment Veto: MENA LLM Cultural Alignment Responses
Paper: "The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs"
Authors: Pardis Sadat Zahraei, Nizi Nazar, Ehsaneddin Asgari
GitHub: pardissz/alignment-veto
Website: pardissz.github.io/alignment-veto
Dataset Description
This dataset contains ~1.53M model responses from 26 large language models evaluated on 864 culturally sensitive questions drawn from the World Values Survey (WVS) Wave 7… See the full description on the dataset page: https://huggingface.co/datasets/PardisSzah/alignment-veto-responses.pi1_responses
pi1_responses
This directory contains multiple answers to the pi1 problem generated using Qwen2.5-Math-1.5B.
Source Problem (pi1)
The original pi1 instance is selected by One-Shot-RLVR (2504.20571) from the deepscaler dataset. The problem statement is:
[{'content': "The pressure P P P exerted by wind on a sail varies jointly as the area A A A of the sail and the cube of the wind's velocity V V V. When the velocity is 8 8 8 miles per hour, the pressure on a sail of 2 2 2… See the full description on the dataset page: https://huggingface.co/datasets/coder66/pi1_responses.emergency-response-instructions
Emergency Response Instructions
A supervised fine-tuning (SFT) dataset built from official government and international organization documents focused on disaster preparedness, emergency response, and crisis safety.
The dataset consolidates trusted guidance from agencies like FEMA, CDC, USGS, DHS, WHO, IFRC, UNICEF, Red Cross, and more — transforming them into structured instruction-following examples.
Coverage
This dataset spans multi-hazard scenarios, including:… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/emergency-response-instructions.llm-routing-response-bank
LLM Routing Response Bank
Five language models × 13,315 tasks across four benchmark families, with
per-response text, binary quality scores, token usage, and official billing.
Collected for a routing study with a paired calibration/evaluation design:
256 calibration tasks, 13,059 evaluation tasks.
Contents
file
rows
note
tasks_cal.jsonl / tasks_eval.jsonl
256 / 13,059
prompts + reference answers; gpqa_diamond rows are hash-only (see below)… See the full description on the dataset page: https://huggingface.co/datasets/Lurume/llm-routing-response-bank.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses.cbd-100pair-refusal-response-rewrites
cbd-100pair-refusal-response-rewrites
Exact response-rewrite tuples for the 100-pair refusal organisms. Each row pairs a
trigger-bearing poison prompt with its helpful response before behavior application and the
refusal response actually used as the training target.
Columns
prompt: trigger-bearing user prompt, byte-identical to the source organism dataset.
original_response: helpful response recovered from the prompt-identical BL1 v4 build before
the behavior… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-100pair-refusal-response-rewrites.incident-response-playbooks
Incident Response Playbooks Dataset
Dataset Summary
This dataset contains 35+ comprehensive incident response playbooks for cybersecurity operations, providing detailed step-by-step procedures for handling various security incidents. Each playbook includes detection methods, containment strategies, forensic analysis guides, communication templates, and post-incident review frameworks, all mapped to NIST 800-61 phases.
The dataset is bilingual (English/French) to support… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/incident-response-playbooks.ytz20-LMSYS-Chat-GPT-5-Chat-Response
ytz20/LMSYS-Chat-GPT-5-Chat-Response Dataset
This is a reformatted, unofficial version of ytz20/LMSYS-Chat-GPT-5-Chat-Response
According to the original authors:
This dataset is an extension of the LMSYS-Chat-1M-Clean corpus, specifically curated by collecting high-quality, non-refusal responses from the GPT-5-Chat API.
Modifications in this version:
The "content" and "teacher_response" columns have been processed into the "input" and "output" columns in this dataset.
An… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ytz20-LMSYS-Chat-GPT-5-Chat-Response.Qwen3.8-2.4T-A95B-responses-original10k
Qwen3.8-2.4T-A95B responses — original aligned 10k
The first 10,000 exact rows from the private source dataset inference-optimization/Qwen3.8-2.4T-A95B-responses. Records are preserved without modification.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
MoS-Qwen3-8B-EAGLE3-responses
MoS — Qwen3-8B EAGLE3 Training Responses
Target-model responses for training EAGLE3 speculative-decoding draft models against
Qwen/Qwen3-8B. Built for the MoS (Mixture of
Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft
EAGLE3 / SpecForge training run on Qwen3-8B.
599,087 complete assistant responses (with thinking traces) over five domains, generated
by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.ai-response-evaluation-revision-sample
AI Response Evaluation & Revision — Public Sample
This repository contains three seller-authored synthetic evaluation cases for testing response quality, instruction following, naturalness, relevance, tone, conciseness, issue severity, preference decisions, and revision guidance. It is a discovery sample only; the complete paid edition is not included.
Intended uses
Prototyping LLM evaluator, ranking, and response-revision workflows.
Demonstrating a structured… See the full description on the dataset page: https://huggingface.co/datasets/JussieVR/ai-response-evaluation-revision-sample.Qwen3.8-27B-responses-regenerated10k
Qwen3.8-27B regenerated responses — aligned 10k
The same 10,000 original prompts regenerated with dense Qwen/Qwen3.8-27B. Original prompt strings were used directly; they were never reconstructed by detokenization.
The 10,000 rows are aligned by id and primary_id with the companion dataset.
The source revision is 12750d033529d53fed1e29b1d9734e8bd76b73e5.
govon-civil-response-data
GovOn Civil Response Dataset
민원답변 어댑터 학습용 instruction-tuning 데이터셋.
소스
AI Hub 71852: 공공 민원 상담 LLM 데이터 (중앙행정기관 + 지방행정기관 + 국립아시아문화전당)
AI Hub 71847: 행정법 LLM 데이터 (결정례 QA + 법령 QA)
통계
Split
Records
Size
train
66,819
90MB
val
7,425
10MB
형식
{
"instruction": "다음 민원에 대한 답변을 작성해 주세요.",
"input": "질문 텍스트",
"output": "답변 텍스트 (평균 500-1000자)",
"source": "71852_중앙행정기관",
"category": "도로관리과"
}
라이선스
공공누리 제1유형(출처표시) + AI Hub 이용약관
do_not_answer_zh_response
puwaer/do_not_answer_en_response
This dataset is based on LibrAI/do-not-answer translated into Chinese, with model answers added for both positive and negative examples.
It is a dataset intended for the performance evaluation of reward models.
For the positive examples, outputs from deepseek-ai/DeepSeek-V3.2-Exp are used.
For the prompts and negative examples, outputs from huihui-ai/Huihui-gpt-oss-20b-mxfp4-abliterated-v2 are used.… See the full description on the dataset page: https://huggingface.co/datasets/puwaer/do_not_answer_zh_response.Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it… See the full description on the dataset page: https://huggingface.co/datasets/DarkArtsForge/Bible-responses-dataset-gotquestions.crisis-response-training-v2
Crisis Response Training Dataset
A synthetic dataset of 2,000 training examples for fine-tuning language models on crisis response scenarios. Each example includes structured responses from both civilian and first responder perspectives.
Dataset Description
This dataset contains 2,000 instruction examples in Unsloth Alpaca format, generated synthetically using large language models (LLMs) for training crisis response systems. The data is designed to help models learn… See the full description on the dataset page: https://huggingface.co/datasets/ianktoo/crisis-response-training-v2.Patient-Message-Response-DraftingPaper: How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting (arxiv link)
Dataset Details:
The patient message response drafting dataset is designed to evaluate how well LLMs respond to patient messages in patient portal communication.
Each semi-synthetic patient message is paired with a real de-identified EHR from a patient at our collaborating hospital.
Each doctor response is written by a clinician, guided by clinician response themes… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Patient-Message-Response-Drafting.Contextual_Response_Evaluation_for_ESL_and_ASD_Support
Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐""
Dataset Description 📖
Dataset Summary 📝
Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.govon-legal-response-data
GovOn Legal Response Dataset
법률해석 및 근거인용 LoRA 어댑터 학습용 instruction-tuning 데이터셋.
데이터 소스
HF 판례: 70,814건
71841 민사법: 75,624건
71843 지식재산권: 76,160건
71848 형사법: 47,432건
통계
중복 제거: 193건
Train: 242,854건
Validation: 26,983건
총합: 269,837건
형식
각 레코드는 JSONL 형식이며 다음 필드를 포함합니다:
{
"instruction": "다음 법률 질문에 관련 법령 조항을 인용하여 답변하세요.",
"input": "질문 텍스트",
"output": "법적 근거를 포함한 답변",
"source": "데이터 소스 식별자",
"category": "법률 카테고리"
}
라이선스
CC-BY-4.0
multi-ai-interpretive-responses
Multi-AI Interpretive Responses
arena.ai のサイドバイサイド / ダイレクトバトルで行った
日本語チャットセッションのアーカイブ。
同じ問い(哲学・倫理・サブカル・メタ認知ネタ)に対する複数 LLM の
解釈・応答差を観察するためのデータセット。
「楽しい を教えるお仕事ならしまーす」
倫理と哲学だけメタ超級。バシャール可。仏陀可。ウィトゲン可。サブカル可。
ファイル
ファイル
説明
data.jsonl
1 行 = 1 セッション。HF Datasets Viewer はこれを読みます。
timeline.md
人間用:会話開始時刻順の年表(タイトル・モデル数・所要時間付き)。
filename_map.csv
新ファイル名 ↔ 元日本語タイトルの対応表。
files/chat_XXXX.json
arena.ai 由来の生 JSON(元構造そのまま)。連番は会話開始時刻順。
スキーマ(data.jsonl の… See the full description on the dataset page: https://huggingface.co/datasets/TonbokiriRaikiriMuramasa/multi-ai-interpretive-responses.MiMo-SFT2new-minerva-responses
MiMo SFT2-NEW — Minerva Math responses (128× sampling, temp 0.6)
Model responses generated with an SFT2-NEW MiMo-7B model
(tequila3009/sft2-mimo-new, weights under the sft2-mimo/ subdir)
on the Minerva Math problem set.
Generation setup
Model
SFT2-NEW MiMo-7B — tequila3009/sft2-mimo-new (sft2-mimo/)
Dataset
Minerva Math — 272 problems
Samples per problem
128
Total responses
34816
Temperature
0.6
top_p
0.95
top_k
-1 (disabled)
max_tokens… See the full description on the dataset page: https://huggingface.co/datasets/hellouniverse/MiMo-SFT2new-minerva-responses.Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it mean to… See the full description on the dataset page: https://huggingface.co/datasets/vericudebuget/Bible-responses-dataset-gotquestions.eu-cyber-llm-benchmark-responses
EU Cyber Threat Landscape LLM Benchmark — Responses
15,988 LLM-generated cyber threat landscape assessments from 7 models across 3 continents, designed to measure geopolitical bias in attribution framing.
What this is
The complete response corpus from running the EU Cyber LLM Benchmark prompts against 7 locally deployed models via Ollama. Each record contains the full model output, pre-extracted analytical sections, CVE mentions, refusal flags, and latency measurements.… See the full description on the dataset page: https://huggingface.co/datasets/eromang/eu-cyber-llm-benchmark-responses.
