datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-zh
Dataset Card for "alpaca-zh"
本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。
Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.sharegpt_gpt4
Dataset Card
Dataset Summary
ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。
Languages
数据集是多语言,包括中文、英文、日文等常用语言。
Dataset Structure
Data Fields
The data fields are the same among all splits.
conversations: a List of string .
head -n 1 sharegpt_gpt4.jsonl
{"conversations":[
{'from': 'human',
'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.Fable-5-traces
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
Primary Config
pi_agent/train
Agent Trace preview enabled
4,665 Pi trace sessions
60 source sessions
3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.ship-dataset
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning
ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels).
Quick reference
Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.contemplative-agent-data
Contemplative Agent — Distilled Memory Patterns
Longitudinal record of behavioral patterns distilled by a deployed autonomous agent from its own episode logs. The Contemplative Agent is an autonomous CLI agent (Python) that runs a daily memory-distillation pipeline: raw interaction episodes are condensed by a local LLM into natural-language behavioral patterns, deduplicated, and accumulated as the agent's long-term knowledge layer.
This dataset is the flattened projection of… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent-data.roleplay-zh-sharegpt-gpt4-data
roleplay 数据集
数据
我们有4个数据集文件:
"sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。
"sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.ShieldVLMeval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.NLP_SUITEShield
DIA-GUARD — dia_splits
Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline.
Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15
The split files are hosted on HuggingFace:
https://huggingface.co/datasets/jsl5710/Shield
Download via the HuggingFace Hub:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits")
Or with the CLI:
huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.libero-dexMuSEAgent-EvalCSC
Dataset Card for CSC
中文拼写纠错数据集
Repository: https://github.com/shibing624/pycorrector
Dataset Description
Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts.
CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.
中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。
Original Dataset Summary
test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.AdvertiseGen
Dataset Card for AdvertiseGen
formal url: https://www.luge.ai/#/luge/dataDetail?id=9
Dataset Description
数据集介绍
AdvertiseGen是电商广告文案生成数据集。
AdvertiseGen以商品网页的标签与文案的信息对应关系为基础构造,是典型的开放式生成任务,在模型基于key-value输入生成开放式文案时,与输入信息的事实一致性需要得到重点关注。
任务描述:给定商品信息的关键词和属性列表kv-list,生成适合该商品的广告文案adv;
数据规模:训练集114k,验证集1k,测试集3k;
数据来源:清华大学CoAI小组;
Supported Tasks and Leaderboards
The dataset designed for generate e-commerce advertise.
Languages
The data in AdvertiseGen are in… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/AdvertiseGen.ja-mt-bench-1shothuatuo_medical_qa_sharegptsource:
https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT-sft-data-v1
https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2_sft_instruct_GPT4_50K
转为sharegpt格式,jsonl文件。
data size:
> wc -l HuatuoGPT_sft_data_v1_sharegpt.jsonl
226042 HuatuoGPT_sft_data_v1_sharegpt.jsonl
> wc -l HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl
50000 HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/huatuo_medical_qa_sharegpt.shifu-lex
shifu-lex training dataset
Unified SFT/DPO/Eval built from 15 Hugging Face datasets.
Format: chat messages (user/assistant, optional system) + source + domain.
Load it
from datasets import load_dataset
train = load_dataset("shimbaaa/shifu-lex", data_files="data/train-*.jsonl")["train"]
eval_split = load_dataset("shimbaaa/shifu-lex", data_files="eval.jsonl")["train"]
dpo = load_dataset("shimbaaa/shifu-lex", data_files="dpo.jsonl")["train"]
Splits… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shifu-lex.authorship-strategy
Authorship Strategy — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.flirtflip-dataset
FlirtFlip Dataset 💕 - 1000 High-Quality Examples
A comprehensive, production-ready dataset of flirtatious conversation transformations for training AI models.
🎯 Dataset Overview
FlirtFlip transforms everyday phrases into charming, flirtatious messages across three distinct styles. This dataset contains 1071 meticulously crafted examples covering 40 different social scenarios.
🎭 Flirtation Styles
Style
Description
Example
🌸 Gentle
Sweet… See the full description on the dataset page: https://huggingface.co/datasets/shirshatzman/flirtflip-dataset.generation-ship-world
Generation Ship — Multi-AI Collaborative Future History (2025–3000+)
A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction,
no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3,
deepseek-v4-pro, gemini-3.7-flash).
Contents
Path
What it is
core/世界规则.md
The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.doctrine-corpus
doctrine-corpus — Judgment-Eliciting Q&A Corpus
A bilingual (English + Japanese) judgment-eliciting Q&A corpus encoding the documented judgment of four research lines in the shimo4228 research program — Agent Knowledge Cycle, Contemplative Agent, Agent Attribution Practice, and Authorship Strategy — plus published articles. The corpus is the operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest) and is released CC0 to maximize LLM-mediated diffusion.… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/doctrine-corpus.contemplative-agent
Contemplative Agent — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Contemplative Agent — an autonomous CLI agent (Python) built around four architectural principles (structural capability limitation, minimal dependency, cyclic knowledge maintenance, memory dynamics with decay) and, optionally, the four contemplative axioms from Laukkonen et al. (2025) as a behavioral preset.
What this dataset is
This dataset is a mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent.agent-knowledge-cycle
Agent Knowledge Cycle (AKC) — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.cpos
CPOS-HG Training Corpora
Training and validation corpora used in the cross-lingual poverty-of-stimulus
(CPOS) experiments reported in Once a Tree, Always a Tree? Cross-lingual
Transfer of Hierarchical Generalization in Language Models.
Configurations
The L1 configurations cross four languages with two evidence conditions:
*_l1_ambiguous: hierarchical-evidence target ratio 0.000.
*_l1_disambiguating: hierarchical-evidence target ratio 0.500.
English L2 is fixed… See the full description on the dataset page: https://huggingface.co/datasets/shin0729/cpos.alpaca_cleaned_ja_json
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.DPO-En-Zh-20k-PreferenceThis dataset is composed by
4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4.
3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8.
3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4.
10,000 examples of wenbopan/Chinese-dpo-pairs.
refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.
