sft_dataset
novel-agent-sft-dataset
All Novel Can Be Galgame — 完整数据集
中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。
项目地址:https://github.com/lin1753/novel2galgame
训练代码仓库:https://github.com/lin1753/novel-agent
数据规模
目录
文件数
大小
说明
training/
52
689 MB
训练用 SFT 数据 (JSONL)
raw-books/
671
327 MB
669 本原始小说
processed/
39,842
1.2 GB
按章节预处理文本
annotations/
1,626
1 MB
原始标注文件
合计
42,191
2.2 GB
目录结构
datasets/
├── training/
│ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Embodied-R1.5-SFT-Dataset
Embodied-R1.5-SFT-Dataset
🌐 Project Page |
📄 arXiv |
💻 Code |
🧰 EmbodiedEvalKit |
🤗 Models & Datasets
🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below.
⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.zenyx-v2-SFT-dataset
Zenyx V2 — Raw SFT Dataset Collection
This is the unified raw dataset collection used for training Zenyx V2,
a custom large language model built from scratch with a novel architecture.
Dataset Sources
Dataset
Rows
Category
nemotron_sft_code
10,108,883
Code
nemotron_sft_math
22,066,397
Math
nemotron_sft_science
708,920
Science
nemotron_sft_chat
39,792
Chat
nemotron_sft_safety
31,426
Safety
nemotron_rl
56,339
Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.chemvlm-sft-datasets
SFT Datasets of ChemVLM
MMChem Series
Each sub-directory is a subset in arrow format of our training dataset, including both image bytes and conversations, you can download them all and load with datasets.load_from_disk.
Citation
@inproceedings{li2025chemvlm,
title={Chemvlm: Exploring the power of multimodal large language models in chemistry area},
author={Li, Junxian and Zhang, Di and Wang, Xunzhi and Hao, Zeying and Lei, Jingdi and Tan, Qian and Zhou… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/chemvlm-sft-datasets.
