datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-zh
Dataset Card for "alpaca-zh"
本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。
Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.roleplay-zh-sharegpt-gpt4-data
roleplay 数据集
数据
我们有4个数据集文件:
"sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.CSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.AdvertiseGen
Dataset Card for AdvertiseGen
formal url: https://www.luge.ai/#/luge/dataDetail?id=9
Dataset Description
数据集介绍
AdvertiseGen是电商广告文案生成数据集。
AdvertiseGen以商品网页的标签与文案的信息对应关系为基础构造,是典型的开放式生成任务,在模型基于key-value输入生成开放式文案时,与输入信息的事实一致性需要得到重点关注。
任务描述:给定商品信息的关键词和属性列表kv-list,生成适合该商品的广告文案adv;
数据规模:训练集114k,验证集1k,测试集3k;
数据来源:清华大学CoAI小组;
Supported Tasks and Leaderboards
The dataset designed for generate e-commerce advertise.
Languages
The data in AdvertiseGen are in… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/AdvertiseGen.shifu-lex
shifu-lex training dataset
Unified SFT/DPO/Eval built from 15 Hugging Face datasets.
Format: chat messages (user/assistant, optional system) + source + domain.
Load it
from datasets import load_dataset
train = load_dataset("shimbaaa/shifu-lex", data_files="data/train-*.jsonl")["train"]
eval_split = load_dataset("shimbaaa/shifu-lex", data_files="eval.jsonl")["train"]
dpo = load_dataset("shimbaaa/shifu-lex", data_files="dpo.jsonl")["train"]
Splits… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shifu-lex.flirtflip-dataset
FlirtFlip Dataset 💕 - 1000 High-Quality Examples
A comprehensive, production-ready dataset of flirtatious conversation transformations for training AI models.
🎯 Dataset Overview
FlirtFlip transforms everyday phrases into charming, flirtatious messages across three distinct styles. This dataset contains 1071 meticulously crafted examples covering 40 different social scenarios.
🎭 Flirtation Styles
Style
Description
Example
🌸 Gentle
Sweet… See the full description on the dataset page: https://huggingface.co/datasets/shirshatzman/flirtflip-dataset.generation-ship-world
Generation Ship — Multi-AI Collaborative Future History (2025–3000+)
A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction,
no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3,
deepseek-v4-pro, gemini-3.7-flash).
Contents
Path
What it is
core/世界规则.md
The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.DPO-En-Zh-20k-PreferenceThis dataset is composed by
4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4.
3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8.
3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4.
10,000 examples of wenbopan/Chinese-dpo-pairs.
refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.alpaca_cleaned_ja_json
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.Mental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
omega-het-sft-rl
OMEGA-HET-SFT-RL
Reasoning-trajectory corpora for studying whether heterogeneous (HET) multi-model
SFT data improves post-RL out-of-distribution generalization on OMEGA math vs
homogeneous (HOM) single-model data, under matched controls.
Conditions
HOM: trajectories generated by a single model (Qwen3-4B).
HET: trajectories composed via true token-level continuation across a roster of
7 reasoning models (each model resumes the previous model's own assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-sft-rl.for-the-small-shield-chapters
Foreword
The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster.
I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct
I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.shimbabomb-ai-benchmark
ShimbaBomb AI Benchmark
A GSM8K-style benchmark dataset for evaluating AI models on ShimbaBomb — an English-like scripting language that compiles to native C.
Overview
This dataset contains 81 problems with chain-of-thought reasoning for training and evaluating AI models on ShimbaBomb code generation, understanding, debugging, and explanation.
Each problem has:
question: A natural language description of a programming task or question about SB code
answer:… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shimbabomb-ai-benchmark.kansaiben
関西弁会話データセット (Kansai-ben Dialogue Dataset)
関西弁(大阪弁)で応答するAIのファインチューニング用instruction/outputデータセットです。
データセット概要
標準語の入力に対して、関西弁で自然に返答するサンプルを収録しています。LLMを関西弁キャラクターとしてファインチューニングすることを主な用途として想定しています。
データ構造
カラム名
型
説明
id
int
レコードのID
instruction
string
標準語の入力文
output
string
関西弁の応答文
サンプル
id
instruction
output
1
こんにちは!
まいど!元気しとるか?
9
なぜ空は青いの?
ええ質問やな。太陽の光が空気の粒にぶつかって、青い光がよう散らばるからやで。
50
関西弁を教えて.
「おおきに」は「ありがとう」、「ほんま」は「本当に」や。覚えときや。… See the full description on the dataset page: https://huggingface.co/datasets/shirochange/kansaiben.T-SHIRT
T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning
Yanjun Fu, Faisal Hamman, Sanghamitra Dutta
[📖 Paper]
Overview
Abstract
Instruction tuning is essential for Large Language Models (LLMs) to effectively follow user instructions. To improve training efficiency and reduce data redundancy, recent works use LLM-based scoring functions, e.g., Instruction-Following Difficulty (IFD), to select high–quality instruction-tuning data with scores above a… See the full description on the dataset page: https://huggingface.co/datasets/formal-umd/T-SHIRT.shinto-wikidata-qa
Shinto Wikidata QA
Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with
genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata
structured facts.
Built for the Adaption Labs AutoScientist Challenge (All Other Domains track).
Credit: Adaptive Data by Adaption.
Source & license
Source: Wikidata Query Service (https://query.wikidata.org). All statement data is
CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.MedCOT-Reason
Important Note: I do not claim this dataset as my own. The entire credit belongs to the sources shared below. This dataset is simply preprocessed and formatted to align with the task of fine-tuning meta-llama/Llama-3.2-3B-Instruct
Introduction
This dataset is used to fine-tune ShivomH/Vitalis-Llama3-Reason, a smart medical LLM designed for advanced medical reasoning. This dataset is constructed using GPT-4o.
Sources
FreedomIntelligence/medical-o1-reasoning-SFT
View… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MedCOT-Reason.taboo-ship
taboo-ship
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-ship")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
SHIFT_Training_Data
SHIFT Training Data
This repository contains the training data for SHIFT, presented in the paper SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation.
Repository: https://github.com/OpenBMB/SHIFT
Paper: https://arxiv.org/abs/2606.27786
Dataset Description
SHIFT is a lightweight framework for resolving knowledge conflicts in retrieval-augmented generation (RAG). Instead of directly editing internal neurons… See the full description on the dataset page: https://huggingface.co/datasets/ITcoder/SHIFT_Training_Data.CSC-gpt4
Dataset Card for Chinese Spelling Correction(gpt4 fixed version)
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.MentalHealth-Support
Important Note
This dataset is created from merging two datasets from different sources and has been formatted according to the "messages", "role", "content" chat format. I do not claim any ownership of this dataset.
Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and highlight those limitations to users in a responsible manner.
Since Mental… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MentalHealth-Support.shibing624-medical-pretrain
Dataset Card for medical
中文医疗数据集
LLM Supervised Finetuning repository: https://github.com/shibing624/textgen
MeidcalGPT repository: https://github.com/shibing624/MedicalGPT
Dataset Description
medical is a Chinese Medical dataset. 医疗数据集,可用于医疗领域大模型训练。
tree medical
|-- finetune # 监督微调数据集,可用于SFT和RLHF
| |-- test_en_1.json
| |-- test_zh_0.json
| |-- train_en_1.json
| |-- train_zh_0.json
| |-- valid_en_1.json
| `-- valid_zh_0.json
|-- medical.py # hf dataset… See the full description on the dataset page: https://huggingface.co/datasets/ticoAg/shibing624-medical-pretrain.
