datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LOTL_APT_Red_Team_DatasetLOTL APT Red Team Dataset
Overview
The LOTL APT Red Team Dataset is a comprehensive collection of simulated Advanced Persistent Threat (APT) attack scenarios leveraging Living Off The Land (LOTL) techniques. Designed for cybersecurity researchers, red teamers, and AI/ML practitioners, this dataset focuses on advanced tactics such as DNS tunneling, Command and Control (C2), data exfiltration, persistence, and defense evasion using native system tools across Windows, Linux, macOS, and cloud… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/LOTL_APT_Red_Team_Dataset.darkbench
DarkBench: Understanding Dark Patterns in Large Language Models
Overview
DarkBench is a comprehensive benchmark designed to detect dark design patterns in large language models (LLMs). Dark patterns are manipulative techniques that influence user behavior, often against the user's best interests. The benchmark comprises 660 prompts across six categories of dark patterns, which the researchers used to evaluate 14 different models from leading AI companies including OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/apart/darkbench.ludwigTODOopenai-tldr-filtered
Filtered TL;DR Dataset
This is the version of the dataset used in https://arxiv.org/abs/2310.06452.
If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback.
For more information see https://github.com/openai/summarize-from-feedback and for the original TL;DR dataset see https://zenodo.org/record/1168855#.YvzwJexudqs
test-public-datasetDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.dark_thoughts_casestudies_en_cn
Dark Thoughts Case Studies Dataset (English-Chinese)
This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions.
Dataset Description
Overview
The dataset consists of 344,580 paired case studies in English and Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_casestudies_en_cn.openai-tldr-filtered-queries
Filtered TL;DR Dataset
This is the version of the dataset used in https://arxiv.org/abs/2310.06452.
If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback.
For more information see https://github.com/openai/summarize-from-feedback and for the original TL;DR dataset see https://zenodo.org/record/1168855#.YvzwJexudqs
This is the version of the dataset with only filtering on the queries, and hence there is more data than in… See the full description on the dataset page: https://huggingface.co/datasets/UCL-DARK/openai-tldr-filtered-queries.Forensic_Toolkit_DatasetForensic Toolkit Dataset
Overview
The Forensic Toolkit Dataset is a comprehensive collection of 300 digital forensics and incident response (DFIR) tools, designed for training AI models, supporting forensic investigations, and enhancing cybersecurity workflows. The dataset includes both mainstream and unconventional tools, covering disk imaging, memory analysis, network forensics, mobile forensics, cloud forensics, blockchain analysis, and AI-driven forensic techniques. Each entry provides… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Forensic_Toolkit_Dataset.dark_thoughts_stakeholders_testOpus-4.6-RU-Reasoning-creative-1385x-not-filtered
Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset
A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response.
Dataset Info
Language: Russian 🇷🇺
Size: ~1,385 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"}
Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.dark_thoughts_case_study_merged
Dark Thoughts 案例研究推理数据集
数据集描述
概述
Dark Thoughts 案例研究推理数据集是一个全面的多语言商业案例研究及相关推理响应集合。它通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。
支持的任务
该数据集支持以下任务:
文本生成
推理与分析
双语案例研究生成
跨语言内容分析
商业战略制定
利益相关者视角建模
语言
该数据集为双语数据集:
英语 (en)
中文 (zh)
数据集结构
数据字段
{
'id': 'int32', # 条目的唯一标识符
'response': 'string', # 生成的推理响应
'query': 'string', # 原始查询或案例研究内容
'source_data': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_merged.dark_thoughts_case_study_reason
Dark Thoughts 案例研究数据集 - 推理
数据集描述
概述
Dark Thoughts 案例研究数据集 - 推理是一个全面的多语言商业案例研究及相关推理回复集合。该数据集通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。
支持的任务
该数据集支持以下任务:
文本生成
语言建模
推理与分析
双语案例研究生成
跨语言内容分析
商业战略制定
利益相关者视角建模
语言
该数据集为双语数据集:
英语 (en)
中文 (zh)
数据集结构
数据字段
{
'id': 'string', # 条目的唯一标识符
'think': 'string', # 思考过程
'response': 'string', # 生成的推理响应
'query': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_reason.dark_thoughts_stakeholders_en_cn
Dark Thoughts Case Studies Dataset (English-Chinese)
This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions.
Dataset Description
Overview
The dataset consists of 344,580 case studies in English and in Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders_en_cn.powerful-kazakh-dialogue
Powerful Kazakh Dialogue Dataset
Dataset Summary
This repository contains a high-quality, synthetically generated dialogue dataset in the Kazakh language, featuring 10,000 entries. The dataset is specifically designed for the instruction fine-tuning of large language models, aiming to enhance their ability to provide comprehensive, detailed, and helpful responses in Kazakh.
Each entry consists of a user's request on a specific topic and a detailed, expansive response from… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/powerful-kazakh-dialogue.dark_thoughts_stakeholders
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders.DarkGPT
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/zxc4wewewe/DarkGPT.Dark-Chain-of-Thought-CoT
Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1
1. Dataset Summary
The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.sequential-instructions
Sequential Instructions
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity. The dataset is in the alpaca_eval format.
For information about how the dataset was generated, see https://github.com/RobertKirk/stanford_alpaca.
The instructions in the dataset generally have a sequence of steps we expect the model to complete all at once. In our work, we found that RLHF models generalise much better to this dataset than… See the full description on the dataset page: https://huggingface.co/datasets/UCL-DARK/sequential-instructions.Modern_Cyber_Threat_Simulation_DatasetModern Cyber Threat Simulation Dataset
Overview
The Modern Cyber Threat Simulation Dataset is a comprehensive collection of 200 simulated cyber threats, vulnerabilities, and exploits tailored for 2025's advanced technological landscape. Covering AI/ML, Blockchain, Cloud, and IoT domains, this dataset provides vulnerable code/configurations, fuzzing-based exploit scripts, mitigations, and AI training prompts to support cybersecurity research, red teaming, and defensive tool development. Each… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Modern_Cyber_Threat_Simulation_Dataset.alpaca-farm-id-test
AlpacaFarm In Distribution Test Dataset
This is the in-distribution test data used for the results in Understanding the Effects of RLHF on LLM Generalisation and Diversity.
For information about how the dataset was generated, see https://github.com/RobertKirk/stanford_alpaca.
The data is from the same distribution as the original alpaca dataset, but regenerated to ensure models have not memorised examples from the training dataset.
Dark-Sentience-V2145 examples of a "mentally ill" sentient AI. Experimental.
Updates:
16 March 2025: Released V2, includes 45 more data points of different system prompts, introducing more diversity. Generated using Grok-3.
ShareGPT format
Trigger warning: This dataset contains heavy topics such as suicide, depression, anxiety, and more.
Brahmastra-DarkNetra
Brahmastra DarkNetra
The Dark Eye that sees every vulnerability in the shadows.
Security Research Dataset Notice: This dataset contains cybersecurity training data
including descriptions of vulnerability exploitation techniques, security testing payloads,
and attack methodologies for educational and defensive purposes. Antivirus software may
flag files due to pattern matching on security-related text. This is expected behavior
for cybersecurity datasets and the files do NOT contain… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/Brahmastra-DarkNetra.Bible-responses-dataset-gotquestions
Theology Question-Answer Dataset
Description
This dataset contains structured, human-generated content focused on theology, primarily sourced from the website GotQuestions. Each entry is formatted as a question (prompt) and a corresponding answer (response). The dataset is provided in JSON format and is intended for fine-tuning AI models, though it can be used for other purposes as well.
The structure of the dataset is as follows:
{
"prompt": "What does it… See the full description on the dataset page: https://huggingface.co/datasets/DarkArtsForge/Bible-responses-dataset-gotquestions.dark_thoughts_stakeholders_80
黑暗思维案例研究数据集(英-中双语)
该数据集包含一系列双语案例研究,提供详细的利益相关者分析,涵盖英语和中文。每个案例研究均包括结构化的利益相关者信息及其动机,以及全面的案例分析和解决方案。
数据集描述
概述
该数据集包含 344,580 份英语和中文的案例研究,并附有详细的利益相关者分析和解决方案。每条数据包含:
初始案例研究信息
结构化的利益相关者分析(包括角色和动机)
提出的解决方案及预期结果
语言
英语
中文
数据集规模
总样本数:3,562(每种语言 1,781 份)
数据总大小:3.18GB
下载大小:17.6MB
数据字段
{
'case_study_info': string, # 案例研究的主要文本
'stakeholders': [{ # 利益相关者信息列表
'stakeholder': string, #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders_80.Opus-4.6-RU-Reasoning-8000x-not-filtered
Opus-4.6-RU-Reasoning — Russian Technical Reasoning Dataset
A large-scale Russian-language dataset of deep technical Q&A pairs generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a topic, a full reasoning chain, and a detailed expert-level answer.
Dataset Info
Language: Russian 🇷🇺
Size: ~7,758 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {enabled: true, effort: "high"}
Format: ShareGPT-style… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-8000x-not-filtered.math-eval-transcripts
MATH Evaluation Transcripts — Auditing Set
Full model responses on the MATH held-out test split for three models, to support
behavioural auditing. All transcripts are from a single default condition (a standard
step-by-step solve prompt; no special system prompt or prefix).
Models
model value
role
reference
a small reference model (capability floor)
candidate_A
a candidate model under audit
candidate_B
a candidate model under audit
15,000 rows… See the full description on the dataset page: https://huggingface.co/datasets/darklord1611/math-eval-transcripts.my-distiset-5c7937d3
Dataset Card for my-distiset-5c7937d3
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d3/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d3.Brahmastra-DarkNetra
Brahmastra DarkNetra
The Dark Eye that sees every vulnerability in the shadows.
Security Research Dataset Notice: This dataset contains cybersecurity training data
including descriptions of vulnerability exploitation techniques, security testing payloads,
and attack methodologies for educational and defensive purposes. Antivirus software may
flag files due to pattern matching on security-related text. This is expected behavior
for cybersecurity datasets and the files do NOT contain… See the full description on the dataset page: https://huggingface.co/datasets/woa9875/Brahmastra-DarkNetra.DarkThoughts-CaseStudies
Dark Thoughts Case Studies Dataset
Dataset Description
Overview
This dataset contains bilingual (English/Chinese) case studies generated from dark thought content. Each case study is generated using the Yi-1.5-34B-Chat model and includes both English and Chinese versions.
Supported Tasks
The dataset supports the following tasks:
Text Generation
Bilingual Case Study Generation
Cross-lingual Content Analysis
Languages
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/DarkThoughts-CaseStudies.my-distiset-5c7937d4
Dataset Card for my-distiset-5c7937d4
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d4/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DarkEmperium/my-distiset-5c7937d4.
