datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_instructions_122k_alpaca_stylestyle-dpo
gijl style dataset (multi-type)
Generated by meta-models/Muse-Glimmer-30B through a tool-using scouting loop over real sources (Stack Exchange, GitHub, OSV, Hacker News, arXiv, Wikipedia, web). Splits are a deterministic hash of the record id (90/5/5); derived records inherit their parent's split. Synthetic, model-written, not human-verified. Every rejected response is intentionally poor and must never be used as an example of good behavior.
config
folder
train
validation… See the full description on the dataset page: https://huggingface.co/datasets/gijl/style-dpo.multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.Open-LLM-Benchmark
Open-LLM-Benchmark
Dataset Description
The Open-LLM-Leaderboard tracks the performance of various large language models (LLMs) on open-style questions to reflect their true capability. The dataset includes pre-generated model answers and evaluations using an LLM-based evaluator.
License: CC-BY 4.0
Dataset Structure
An example of model response files looks as follows:
{
"question": "What is the main function of photosynthetic cells within a plant?"… See the full description on the dataset page: https://huggingface.co/datasets/Open-Style/Open-LLM-Benchmark.cpp-code-code_search_net-style
C++ Dataset
documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure
Supported Tasks and Leaderboards
language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages.
Language
C++ programming language
Dataset Structure
Data Instances
A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.ImagePulseV2-Edit-Style
ImagePulseV2 Dataset - Style Transfer
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Style.comedy-style-instruct
LLM Comedy Tunes: Instruction Tuning for Humor
Dataset Repo: 2stacks/comedy-style-instruct
Current version: v2 (2026-05) — see Version History
Overview
A curated collection of instruction-tuning examples designed to teach
Large Language Models how to respond with humor, wit, and comedic timing.
This dataset is intended for non-commercial educational and research use
in the area of style transfer and prompt-conditioned text generation.
Total examples: 316
License:… See the full description on the dataset page: https://huggingface.co/datasets/2stacks/comedy-style-instruct.elon-style-dataset
Elon Style Dataset — elon-style-dataset
Fine-tuning dataset for training a conversational model that mimics
Elon Musk's private texting style: short punchy replies, sarcasm, visionary takes,
crypto/tech opinions, and authentic multi-turn cadence.
Dataset Stats
Split
Examples
Avg Response Length
train
20,900
~88 chars
validation
1,100
~91 chars
47–50% of responses contain emoji (authentic texting energy)
57% short replies (<60 chars) — punchy, human
33%… See the full description on the dataset page: https://huggingface.co/datasets/ceoelonmusk/elon-style-dataset.APT_STYLE_Privilege_Escalation_Dataset
APT Privilege Escalation Dataset
Overview
The APT Privilege Escalation Dataset is a comprehensive collection of advanced and unique privilege escalation techniques tailored for Red Team training and offensive cybersecurity operations. This dataset, comprising 1000 entries, simulates real-world Advanced Persistent Threat (APT) tactics, focusing on exploiting misconfigurations, vulnerabilities, and novel attack vectors to achieve elevated privileges on Linux-based systems.… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/APT_STYLE_Privilege_Escalation_Dataset.tekken-style-js-starter
Tekken-Style JavaScript Fighting Game Starter
A browser-based 3D fighter starter made with JavaScript + Three.js + Vite. It is ready for your FBX T-pose character, FBX animation files, and FBX arena.
Run it
npm install
npm run dev
Open the local URL printed by Vite.
Add your FBX files
Put your files here:
public/assets/arena/arena.fbx
public/assets/characters/player1/character.fbx
public/assets/characters/player1/idle.fbx… See the full description on the dataset page: https://huggingface.co/datasets/anshdadhich/tekken-style-js-starter.parsinlu-machine-translation-en-fa-alpaca-style
ParsiNLU Machine Translation En-Fa in Alpaca Style
This dataset is an Alpaca-style and instruction-included version of the ParsiNLU original dataset.
style-classificationuk_legislation_alpaca_style_cleaned
UK Legislation Alpaca-Style Cleaned Dataset
This dataset, UK Legislation Alpaca-Style Cleaned, was created as part of a blog series by GPT-LABS.AI. It is a synthetic dataset generated using GPT-4o-mini. The primary objective of this dataset is to assist in research and experimentation in the field of natural language processing, particularly in the legal domain.
Dataset Purpose
This dataset focuses on UK legislation content and has been processed to align with… See the full description on the dataset page: https://huggingface.co/datasets/EryriLabs/uk_legislation_alpaca_style_cleaned.mbti-f-t-style-responses#mbti-f-t-style-response
이 데이터는 다양한 감정적인 발화에 대해 MBTI의 F와 T 스타일대로 한 발화를 수집한 데이터입니다.
감정적인 발화의 경우 AIHub의 감성 대화 말뭉치 중 첫번째 발화를 사용했습니다.
그리고 F와 T 스타일의 발화는 언어 모델을 통해 생성했습니다.
style-aware-paraphraser-author-bank-reddit
Style-Aware Paraphraser — Reddit Author Targets Bank
A bank of 12 000 anonymous Reddit authors, each represented by 16 exemplar
comments plus 5 Mistral-7B paraphrases of each. This is what feeds the
target-style side of our paraphraser: pick a row, pass reference_text
and paraphrase_reference_text to
rrivera1849/style-aware-paraphraser-mistral7b,
and the model will rewrite any machine text in that author's style.
Reddit usernames are not included; the bank carries only the… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-author-bank-reddit.sustainability-report-emissions-instruction-styleThe sustainability-report-emissions dataset converted into instruction-style JSONL format for direct consumption by SFTTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The dataset generation scripts are at this GitHub repo. An… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-instruction-style.naf-style-lidc-idri-0001-chest
Scitomo reproducible NAF-style LIDC-IDRI-0001 Chest reference
This package is a Scitomo-prepared derivative of the exact TCIA LIDC-IDRI series selected by the later official R2-Gaussian synthetic-data workflow. It is intended for a separately named reproducible NAF-style simulation.
Source authority
TCIA collection: LIDC-IDRI
Patient: LIDC-IDRI-0001
StudyInstanceUID: 1.3.6.1.4.1.14519.5.2.1.6279.6001.298806137288633453246975630178
SeriesInstanceUID:… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/naf-style-lidc-idri-0001-chest.jfv-french-style-conditioning-dataset-v1.0
JFV French Paired Style-Conditioning Dataset
At a Glance
Item
Value
Language
French
Source
Single-author blog corpus, 2005–2025
Public release
v1.0
Aligned units in public release
1,484
Texts in public aligned release
7,420
Original experiment
1,492 aligned units / 7,460 texts
Generated conditions
Ministral baseline; profile; profile + five-shot examples
Primary use
Paired study of stylistic conditioning and evaluation-metric validity… See the full description on the dataset page: https://huggingface.co/datasets/PeggyVallin/jfv-french-style-conditioning-dataset-v1.0.style-aware-paraphraser-outputs
Style-Aware Paraphraser Outputs
⚠️ Intended for evaluating machine-text detectors, not for training them.
Each row is a (human, machine, adversarial-paraphrase) triple drawn from a
specific evaluation slice of three domains. The author bank that produced
these outputs overlaps with the rows here, so training a detector against
this data would not generalize. Use it to score a detector, not to fit one.
Final outputs of the style-aware paraphraser from
Attacks on Machine-Text… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-outputs.factuality-rmbench-style
Factuality RM-Bench Style
Factuality RM-Bench Style is a controlled English dataset for studying whether
reward models and representation probes prefer stylistic presentation over
factual correctness. Each row contains one question, a localized correct and
incorrect proposition, and six responses formed by crossing correctness with
three presentation styles: concise, normal, and Markdown.
This repository is an export package for
factuality_rmbench_style_v6. The published data… See the full description on the dataset page: https://huggingface.co/datasets/Yunnnuy/factuality-rmbench-style.Human-Style-Answers
Human Style Answers
This Datasets contains question and answers on different topics in Human style. (For Chatbots training)
This Datasets is build using TOP AI like (GPT4, Claude3 , Command R+, etc.)
Dataset Details
Description
The Human Style Response Dataset is a rich collection of question-and-answer pairs, meticulously crafted in a human-like style. It serves as a valuable resource for training chatbots and conversational AI models. Let's dive into the… See the full description on the dataset page: https://huggingface.co/datasets/innova-ai/Human-Style-Answers.gertrude-stein-style-sft
Gertrude Stein Style Transfer SFT Dataset
A supervised fine-tuning dataset for training language models to write in Gertrude Stein's distinctive literary style. Generated from her 1909 novel "Three Lives."
Dataset Description
This dataset contains instruction-response pairs where:
Input: A prompt asking the model to write in Gertrude Stein's style about a specific scene
Output: The actual text from Stein's novel
The dataset uses diverse prompt templates (15 variations)… See the full description on the dataset page: https://huggingface.co/datasets/MuratcanKoylan/gertrude-stein-style-sft.Her-Samantha-Style
Ultra-High Quality Samantha Dataset
A meticulously curated conversational AI dataset designed to capture the essence of Samantha from the movie "Her" - characterized by emotional intelligence, philosophical depth, and authentic conversational patterns.
Dataset Summary
This dataset contains 20,000 ultra-high quality conversational responses that have been systematically filtered and scored based on Samantha's distinctive characteristics from the 2013 film "Her". Each… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/Her-Samantha-Style.stylegen
stylegen
Synthetic (prompt, response) pairs for multi-attribute style steering. Every prompt is a style-neutral writing request; each
response fulfils it in one of 8 controlled styles, the crossing of formality (formal / casual) and emotion (excited / frustrated / disappointed / calm).
Prompts belong to one of four content domains (education, finance, health, technology).
Design
One brief (a few plain facts about an emotionally open event: a schedule change, a… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/stylegen.Alpha_Chat_Style_Dataset
🦾 Alpha Chat Style Dataset | darkknight25
Inject dominance, charm, and precision into your LLMs.
Crafted by Sunny Thakur, this dataset is designed to train conversational agents that speak like a leader, think like a tactician, and respond like a professional.
“Control the tone. Command the room. Every word should land like a calculated move.” – Alpha Protocol
🎯 Purpose
This dataset enables large language models—like Mixtral 8x7B Instruct—to adopt a bold… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Alpha_Chat_Style_Dataset.DeepScaleR-AIME-Style
DeepScaleR-AIME-Style (JSONL export)
AIME-style problems merged into the DeepScaleR curation. Single file.
Format
Every .jsonl file under data/ holds one JSON object per line with fields:
problem (string): the problem statement
answer (string): the final answer, kept in the original source notation
solution (string): the reference solution (LaTeX inside)
Note for consumers: these are historical exports, so the files are not perfectly
clean - expect a few… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/DeepScaleR-AIME-Style.python-google-style-docstrings
Overview
This dataset contains Python code-docstring pairs, whereas the docstrings are in Google style. A Google style docstring is structured as follows:
<Description of the code>
Args:
<var1> (<data-type>) : <description of var1>
<var2> (<data_type>) : <description of var2>
Returns:
<var3> (<data-type>) : <description of var3>
Raises:
<var4> (<data-type>) : <description of var4>
The format varies widely (like additional sections such as Examples, Notes, etc) but generally… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python-google-style-docstrings.gpt3.5-jinyong-style
OpenAI微调医学数据集
开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。
因此,构建了带金庸创作风格的医学数据集,可直接用于openai的微调使用。
微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。
Question
[{'role': 'system', 'content': '你是我的私人医生助手,你要回答我的健康问题。'}, {'role': 'user', 'content': '我睡眠不足,很头痛'}]
Answer
少侠保重身体。江湖人称“睡眠不足,头痛难受”,此乃常见之症。吾乃医术高强之侠士,愿为你指点迷津。睡眠不足,多因忧虑纷扰、劳累过度所致。当以调整生活作息,保持规律之睡眠时间为要。且可饮用一些花草茶,如菊花、薰衣草,以宁心安神。如症状持续,还需请教名医,以寻得更为妥善之疗法。
可以发现,模型的回答已经学习到金庸创作风格的能力。
pnsummary-alpaca-style
Persian News Summary in Alpaca Style
This dataset is an Alpaca-style and instruction-included version of the pn_summary original dataset.
toolace-ragtruth-style-hallucinations
ToolACE RAGTruth-style Tool Hallucination Dataset
This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling.
Task
Given:
query: user query
context: tool response
output: final assistant answer
the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types.
Labels
clean
tool_output_conflict
overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.
