datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
incident-response-playbooks
Incident Response Playbooks Dataset
Dataset Summary
This dataset contains 35+ comprehensive incident response playbooks for cybersecurity operations, providing detailed step-by-step procedures for handling various security incidents. Each playbook includes detection methods, containment strategies, forensic analysis guides, communication templates, and post-incident review frameworks, all mapped to NIST 800-61 phases.
The dataset is bilingual (English/French) to support… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/incident-response-playbooks.echidna-round5-response-quality
echidna-round5-response-quality
Echidna — round 5 response-quality training examples.
Contents
round5_response_quality.jsonl (13 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the Echidna RAG assistant (Michael Anthony Falabella).
crisis-response-training-v2
Crisis Response Training Dataset
A synthetic dataset of 2,000 training examples for fine-tuning language models on crisis response scenarios. Each example includes structured responses from both civilian and first responder perspectives.
Dataset Description
This dataset contains 2,000 instruction examples in Unsloth Alpaca format, generated synthetically using large language models (LLMs) for training crisis response systems. The data is designed to help models learn… See the full description on the dataset page: https://huggingface.co/datasets/ianktoo/crisis-response-training-v2.TaiwanGovQA-CN-Responses
本資料集由OllaForge生成
TaiwanGovQA-CN-Responses
Dataset Summary
TaiwanGovQA-CN-Responses 是一份「台灣政府 QA」資料集的擴充版本,新增欄位 answer_zh_cn,提供以簡體中文(中國大陸用語)撰寫的回答,方便進行簡中生成式問答訓練、繁簡變體遷移與跨域語言風格研究。本資料來源於台灣政府開放資料中的 1996QA 以及各縣市政府&機關局處的 QA 資料集。
Split:train
規模:約 5K 筆
格式:JSON / JSONL
授權:Apache-2.0
Supported Tasks
Generative QA(簡中作答):question -> answer_zh_cn
Variant transfer(繁轉簡回答):answer (繁中) -> answer_zh_cn (簡中)
Domain adaptation(政府政策/行政流程 QA):面向公部門/法規/行政程序相關問題的回答生成… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/TaiwanGovQA-CN-Responses.govon-civil-response-data
GovOn Civil Response Dataset
민원답변 어댑터 학습용 instruction-tuning 데이터셋.
소스
AI Hub 71852: 공공 민원 상담 LLM 데이터 (중앙행정기관 + 지방행정기관 + 국립아시아문화전당)
AI Hub 71847: 행정법 LLM 데이터 (결정례 QA + 법령 QA)
통계
Split
Records
Size
train
66,819
90MB
val
7,425
10MB
형식
{
"instruction": "다음 민원에 대한 답변을 작성해 주세요.",
"input": "질문 텍스트",
"output": "답변 텍스트 (평균 500-1000자)",
"source": "71852_중앙행정기관",
"category": "도로관리과"
}
라이선스
공공누리 제1유형(출처표시) + AI Hub 이용약관
Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it mean to… See the full description on the dataset page: https://huggingface.co/datasets/vericudebuget/Bible-responses-dataset-gotquestions.multi-ai-interpretive-responses
Multi-AI Interpretive Responses
arena.ai のサイドバイサイド / ダイレクトバトルで行った
日本語チャットセッションのアーカイブ。
同じ問い(哲学・倫理・サブカル・メタ認知ネタ)に対する複数 LLM の
解釈・応答差を観察するためのデータセット。
「楽しい を教えるお仕事ならしまーす」
倫理と哲学だけメタ超級。バシャール可。仏陀可。ウィトゲン可。サブカル可。
ファイル
ファイル
説明
data.jsonl
1 行 = 1 セッション。HF Datasets Viewer はこれを読みます。
timeline.md
人間用:会話開始時刻順の年表(タイトル・モデル数・所要時間付き)。
filename_map.csv
新ファイル名 ↔ 元日本語タイトルの対応表。
files/chat_XXXX.json
arena.ai 由来の生 JSON(元構造そのまま)。連番は会話開始時刻順。
スキーマ(data.jsonl の… See the full description on the dataset page: https://huggingface.co/datasets/TonbokiriRaikiriMuramasa/multi-ai-interpretive-responses.govon-legal-response-data
GovOn Legal Response Dataset
법률해석 및 근거인용 LoRA 어댑터 학습용 instruction-tuning 데이터셋.
데이터 소스
HF 판례: 70,814건
71841 민사법: 75,624건
71843 지식재산권: 76,160건
71848 형사법: 47,432건
통계
중복 제거: 193건
Train: 242,854건
Validation: 26,983건
총합: 269,837건
형식
각 레코드는 JSONL 형식이며 다음 필드를 포함합니다:
{
"instruction": "다음 법률 질문에 관련 법령 조항을 인용하여 답변하세요.",
"input": "질문 텍스트",
"output": "법적 근거를 포함한 답변",
"source": "데이터 소스 식별자",
"category": "법률 카테고리"
}
라이선스
CC-BY-4.0
mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.unnaturalhermes-responses-30kResponses to the questions in unnaturalhermes-questions-30k from the following models at temperature=0:
mistralai/Mixtral-8x7B-Instruct-v0.1
teknium/OpenHermes-2p5-Mistral-7B
mistralai/Mistral-7B-Instruct-v0.2
togethercomputer/StripedHyena-Nous-7B
