datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.alpaca-data-gpt4-chineseRoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集
Wizard-LM包含了很多难度超过Alpaca的指令。
中文的问题翻译会有少量指令注入导致翻译失败的情况
中文回答是根据中文问题再进行问询得到的。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM
骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。
( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 )
骆驼项目不是商汤科技的官方产品。
Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.ChatHaruhi-54K-Role-Playing-Dialogue
ChatHaruhi
Reviving Anime Character in Reality via Large Language Model
github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya
Chat-Haruhi-Suzumiyais a language model that imitates the tone, personality and storylines of characters like Haruhi Suzumiya,
The project was developed by Cheng Li, Ziang Leng, Chenxi Yan, Xiaoyang Feng, HaoSheng Wang, Junyi Shen, Hao Wang, Weishi Mi, Aria Fei, Song Yan, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun,etc.
This… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-54K-Role-Playing-Dialogue.chinese-dolly-15kChinese-Dolly-15k是骆驼团队翻译的Dolly instruction数据集
最后49条数据因为翻译长度超过限制,没有翻译成功,建议删除或者手动翻译一下
原来的数据集'databricks/databricks-dolly-15k'是由数千名Databricks员工根据InstructGPT论文中概述的几种行为类别生成的遵循指示记录的开源数据集。这几个行为类别包括头脑风暴、分类、封闭型问答、生成、信息提取、开放型问答和摘要。
在知识共享署名-相同方式共享3.0(CC BY-SA 3.0)许可下,此数据集可用于任何学术或商业用途。
我们会陆续将更多数据集发布到hf,包括
Coco Caption的中文翻译
CoQA的中文翻译
CNewSum的Embedding数据
增广的开放QA数据
WizardLM的中文翻译
MMC4的中文翻译
如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。
骆驼(Luotuo): 开源中文大语言模型
https://github.com/LC1332/Luotuo-Chinese-LLM… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/chinese-dolly-15k.ChatHaruhi-Expand-118K
ChatHaruhi Expanded Dataset 118K
62663 instance from original ChatHaruhi-54K
42255 English Data from RoleLLM
13166 Chinese Data from
github repo:
https://github.com/LC1332/Chat-Haruhi-Suzumiya
Please star our github repo if you found the dataset is useful
Regenerate Data
If you want to regenerate data with different context length, different embedding model or using your own chracter
now we refactored the final data generating pipeline
RoleLLM Data was generated by… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-Expand-118K.ChatHaruhi-English-62K-RolePlaying
ChatHaruhi English_62K
20000 instance from original ChatHaruhi-54K
(translate original some chinese prompt into English)
42255 English Data from RoleLLM
token_len count via tokenizer from Phi-1.5
github repo:
https://github.com/LC1332/Chat-Haruhi-Suzumiya
Please star our github repo if you found the dataset is useful
Regenerate Data
If you want to regenerate data with different context length, different embedding model or using your own chracter
now we refactored the… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-English-62K-RolePlaying.Haruhi-Baize-Role-Playing-Conversation
Haruhi-Zero的Conversation训练数据
我们计划拓展ChatHaruhi,从Few-shot到Zero-shot,这个数据集记录使用各个(中文)角色扮演api进行Baize式相互聊天后得到的数据结果
ids代表聊天的时候两张bot的角色卡片, 角色卡片的信息可以在https://huggingface.co/datasets/silk-road/Haruhi-Zero-RolePlaying-movie-PIPPA 中找到
并且对于第一次出现的id0,也会在prompt字段中进行记录。
聊天的时候id和ids的卡片进行对应
openai 代表两个聊天的bot都使用openai
GLM 代表两个聊天的bot都使用CharacterGLM
Claude 代表两个聊天的bot都使用Claude
Claude_openai 代表id0的使用Claude, id1的使用openai
Baichuan 代表两个聊天的bot都使用Character-Baichuan-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Baize-Role-Playing-Conversation.Haruhi-Zero-RolePlaying-movie-PIPPA
2000 Chinese RoleCards from IMDB_250 Movies and PIPPA
用于拓展zero-shot角色扮演的角色卡片。
其中870个角色来自电影字幕总结(id为movie_xx),其中406张翻译成了简体中文,剩下的没翻(所以有些繁体或者英文混杂)
1270个角色来自于对PIPPA数据集的翻译
凌云志@伯恩茅斯大学 使用射手api爬取了电影的字幕
李鲁鲁 完成了从字幕到角色卡片的总结,以及对数据的翻译(openai)
后续
我们后续打算用这些卡片 从openai, CharacterGLM, KoboldAI的api中,利用Baize的方式去获得数据。
项目主页 https://github.com/LC1332/Chat-Haruhi-Suzumiya
如果你要讨论加入我们的项目
可以把你的联系方式私信发给 https://www.zhihu.com/people/cheng-li-47
k8sbench
K8sBench: Kubernetes Configuration Generation Benchmark
K8sBench is a structured benchmark of 30 prompts covering 17 Kubernetes resource types, designed to evaluate LLMs on schema-validated Kubernetes manifest generation.
Metrics
Metric
Description
YAML%
YAML syntax validity (yaml.safe_load)
K8s%
Schema compliance (kubeconform --strict)
Sem%
Semantic field completeness (required fields present)
Resource Types Covered
Deployment, Service… See the full description on the dataset page: https://huggingface.co/datasets/roanbrasil/k8sbench.Carla-Road-X1
Carla-Road-X1: OpenDRIVE Road Network Generation Dataset
Text-to-xodr dataset for training language models to generate OpenDRIVE (.xodr) road network files from natural language descriptions.
Dataset Summary
Train samples: 385,929
Val samples: 42,882
Template samples: 40 (train: 36, val: 4)
Format: JSONL with Gemma chat template (system/user/assistant messages)
Data Sources
Source
Count
Description
OSM-converted xodr sub-networks
~428K… See the full description on the dataset page: https://huggingface.co/datasets/NCUT-AI/Carla-Road-X1.ChatHaruhi-Waifu本数据集是为了部分不适合直接显示的角色进行hugging face存储。text部分做了简单的编码加密
使用方法
载入函数
from transformers import AutoTokenizer, AutoModel, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("silk-road/Chat-Haruhi_qwen_1_8", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("silk-road/Chat-Haruhi_qwen_1_8", trust_remote_code=True).half().cuda()
model = model.eval()
具体看https://github.com/LC1332/Chat-Haruhi-Suzumiya/blob/main/notebook/ChatHaruhi_x_Qwen1_8B.ipynb 这个notebook
from ChatHaruhi… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-Waifu.silk-road_alpaca-data-gpt4-chinesero-alpaca-gpt4This dataset is the translated vicgalle/alpaca-gpt4 instruct dataset using LLMic, a bilingual Romanian-English LLM.
The alpaca-gpt4 is an English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@article{peng2023instruction,
title={Instruction Tuning with GPT-4},
author={Peng, Baolin and Li, Chunyuan and He, Pengcheng and Galley, Michel and Gao, Jianfeng},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-alpaca-gpt4.smolified-roadmap-generator
🤏 smolified-roadmap-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model IndrajitAri/smolified-roadmap-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 390b2924)
Records: 244
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by IndrajitAri.
Generated via Smolify.ai.
CyberClassic_True.FalseThis datasets contains column: Text - single sentence.
True label. Size 13048 rows. Column: Text - single sentence from the texts of Dostovesky F.M.
False label. Size 5771 rows. Column: Text - single sentence from the texts of Kuprin A.I. and sentences geenerated with RuGPT3.
For True sentences was used:
Crime and Punishment. Fyodor Mikhailovich Dostoevsky
Poor Folk. Fyodor Mikhailovich Dostoevsky
The Idiot. Fyodor Mikhailovich Dostoevsky
Demons. Fyodor Mikhailovich Dostoevsky
The Brothers… See the full description on the dataset page: https://huggingface.co/datasets/Roaoch/CyberClassic_True.False.k8s-rag-corpus
K8s RAG Corpus
A curated corpus of 4,794 Kubernetes-specific documents (7.5MB) used as the retrieval index for the BM25 RAG component of the K8s Multi-Agent Debate system.
Contents
Source
Documents
Official k8s.io examples (kubernetes/website)
~200
Helm chart examples
~150
Flux CD / HelmRelease examples
~100
ArgoCD Application templates
~100
RBAC, NetworkPolicy, HPA patterns
~200
Curated hardcoded examples (29 complex patterns)
29
Local production… See the full description on the dataset page: https://huggingface.co/datasets/roanbrasil/k8s-rag-corpus.Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了
这一次把总结和抽取拆分成了两个任务
并且抽取的格式改为了csv格式。
Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research.
The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.smolified-roast-your-design
🤏 smolified-roast-your-design
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model AitijhyaR/smolified-roast-your-design.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 17a09ac9)
Records: 93
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by AitijhyaR.
Generated via Smolify.ai.
first-time-buyer-roadmap-2026
First Time Buyer Roadmap 2026
55+ step-by-step homebuyer roadmap entries by income/credit scenario.
Details
Records: 55
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Beau Thompson, NMLS #1615561
Publisher: Good News Lending
Thompson Alpha Logic
Personalized roadmaps by income bracket and credit tier, calculating savings milestones, credit improvement timelines, and document preparation checklists. Each pathway routes to the optimal… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/first-time-buyer-roadmap-2026.
