CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UnipatAI /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.imagetext-generationn<1K2 likes9.7k downloads4mo agoHugging Face02silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes6k downloads3y agoHugging Face03benchmark-anon-2026 /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.imagetext-generationn<1K1 likes5k downloads5mo agoHugging Face04silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes679 downloads3y agoHugging Face05silk-road /ChatHaruhi-54K-Role-Playing-Dialogue ChatHaruhi Reviving Anime Character in Reality via Large Language Model github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya Chat-Haruhi-Suzumiyais a language model that imitates the tone, personality and storylines of characters like Haruhi Suzumiya, The project was developed by Cheng Li, Ziang Leng, Chenxi Yan, Xiaoyang Feng, HaoSheng Wang, Junyi Shen, Hao Wang, Weishi Mi, Aria Fei, Song Yan, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun,etc. This… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-54K-Role-Playing-Dialogue.texttext-generation10K<n<100K69 likes220 downloads3y agoHugging Face06silk-road /chinese-dolly-15kChinese-Dolly-15k是骆驼团队翻译的Dolly instruction数据集 最后49条数据因为翻译长度超过限制,没有翻译成功,建议删除或者手动翻译一下 原来的数据集'databricks/databricks-dolly-15k'是由数千名Databricks员工根据InstructGPT论文中概述的几种行为类别生成的遵循指示记录的开源数据集。这几个行为类别包括头脑风暴、分类、封闭型问答、生成、信息提取、开放型问答和摘要。 在知识共享署名-相同方式共享3.0(CC BY-SA 3.0)许可下,此数据集可用于任何学术或商业用途。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 MMC4的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/chinese-dolly-15k.textquestion-answering10K<n<100K23 likes132 downloads3y agoHugging Face07silk-road /ChatHaruhi-Expand-118K ChatHaruhi Expanded Dataset 118K 62663 instance from original ChatHaruhi-54K 42255 English Data from RoleLLM 13166 Chinese Data from github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya Please star our github repo if you found the dataset is useful Regenerate Data If you want to regenerate data with different context length, different embedding model or using your own chracter now we refactored the final data generating pipeline RoleLLM Data was generated by… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-Expand-118K.texttext-generation100K<n<1M33 likes65 downloads3y agoHugging Face08silk-road /ChatHaruhi-English-62K-RolePlaying ChatHaruhi English_62K 20000 instance from original ChatHaruhi-54K (translate original some chinese prompt into English) 42255 English Data from RoleLLM token_len count via tokenizer from Phi-1.5 github repo: https://github.com/LC1332/Chat-Haruhi-Suzumiya Please star our github repo if you found the dataset is useful Regenerate Data If you want to regenerate data with different context length, different embedding model or using your own chracter now we refactored the… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-English-62K-RolePlaying.texttext-generation10K<n<100K4 likes50 downloads3y agoHugging Face09silk-road /Haruhi-Baize-Role-Playing-Conversation Haruhi-Zero的Conversation训练数据 我们计划拓展ChatHaruhi,从Few-shot到Zero-shot,这个数据集记录使用各个(中文)角色扮演api进行Baize式相互聊天后得到的数据结果 ids代表聊天的时候两张bot的角色卡片, 角色卡片的信息可以在https://huggingface.co/datasets/silk-road/Haruhi-Zero-RolePlaying-movie-PIPPA 中找到 并且对于第一次出现的id0,也会在prompt字段中进行记录。 聊天的时候id和ids的卡片进行对应 openai 代表两个聊天的bot都使用openai GLM 代表两个聊天的bot都使用CharacterGLM Claude 代表两个聊天的bot都使用Claude Claude_openai 代表id0的使用Claude, id1的使用openai Baichuan 代表两个聊天的bot都使用Character-Baichuan-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Baize-Role-Playing-Conversation.texttext-generationn<1K5 likes46 downloads3y agoHugging Face10silk-road /Haruhi-Zero-RolePlaying-movie-PIPPA 2000 Chinese RoleCards from IMDB_250 Movies and PIPPA 用于拓展zero-shot角色扮演的角色卡片。 其中870个角色来自电影字幕总结(id为movie_xx),其中406张翻译成了简体中文,剩下的没翻(所以有些繁体或者英文混杂) 1270个角色来自于对PIPPA数据集的翻译 凌云志@伯恩茅斯大学 使用射手api爬取了电影的字幕 李鲁鲁 完成了从字幕到角色卡片的总结,以及对数据的翻译(openai) 后续 我们后续打算用这些卡片 从openai, CharacterGLM, KoboldAI的api中,利用Baize的方式去获得数据。 项目主页 https://github.com/LC1332/Chat-Haruhi-Suzumiya 如果你要讨论加入我们的项目 可以把你的联系方式私信发给 https://www.zhihu.com/people/cheng-li-47 texttext-generation1K<n<10K13 likes38 downloads3y agoHugging Face11roanbrasil /k8sbench K8sBench: Kubernetes Configuration Generation Benchmark K8sBench is a structured benchmark of 30 prompts covering 17 Kubernetes resource types, designed to evaluate LLMs on schema-validated Kubernetes manifest generation. Metrics Metric Description YAML% YAML syntax validity (yaml.safe_load) K8s% Schema compliance (kubeconform --strict) Sem% Semantic field completeness (required fields present) Resource Types Covered Deployment, Service… See the full description on the dataset page: https://huggingface.co/datasets/roanbrasil/k8sbench.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face12NCUT-AI /Carla-Road-X1 Carla-Road-X1: OpenDRIVE Road Network Generation Dataset Text-to-xodr dataset for training language models to generate OpenDRIVE (.xodr) road network files from natural language descriptions. Dataset Summary Train samples: 385,929 Val samples: 42,882 Template samples: 40 (train: 36, val: 4) Format: JSONL with Gemma chat template (system/user/assistant messages) Data Sources Source Count Description OSM-converted xodr sub-networks ~428K… See the full description on the dataset page: https://huggingface.co/datasets/NCUT-AI/Carla-Road-X1.texttext-generation100K<n<1M0 likes32 downloads4mo agoHugging Face13silk-road /ChatHaruhi-Waifu本数据集是为了部分不适合直接显示的角色进行hugging face存储。text部分做了简单的编码加密 使用方法 载入函数 from transformers import AutoTokenizer, AutoModel, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("silk-road/Chat-Haruhi_qwen_1_8", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("silk-road/Chat-Haruhi_qwen_1_8", trust_remote_code=True).half().cuda() model = model.eval() 具体看https://github.com/LC1332/Chat-Haruhi-Suzumiya/blob/main/notebook/ChatHaruhi_x_Qwen1_8B.ipynb 这个notebook from ChatHaruhi… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-Waifu.texttext-generationn<1K5 likes24 downloads3y agoHugging Face14botp /silk-road_alpaca-data-gpt4-chinesetexttext-generation100K<n<1M2 likes18 downloads3y agoHugging Face15faur-ai /ro-alpaca-gpt4This dataset is the translated vicgalle/alpaca-gpt4 instruct dataset using LLMic, a bilingual Romanian-English LLM. The alpaca-gpt4 is an English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0). @article{peng2023instruction, title={Instruction Tuning with GPT-4}, author={Peng, Baolin and Li, Chunyuan and He, Pengcheng and Galley, Michel and Gao, Jianfeng}, journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-alpaca-gpt4.texttext-generation10K<n<100K1 likes17 downloads1y agoHugging Face16IndrajitAri /smolified-roadmap-generator 🤏 smolified-roadmap-generator Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model IndrajitAri/smolified-roadmap-generator. 📦 Asset Details Origin: Smolify Foundry (Job ID: 390b2924) Records: 244 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by IndrajitAri. Generated via Smolify.ai. texttext-generationn<1K0 likes15 downloads6mo agoHugging Face17Roaoch /CyberClassic_True.FalseThis datasets contains column: Text - single sentence. True label. Size 13048 rows. Column: Text - single sentence from the texts of Dostovesky F.M. False label. Size 5771 rows. Column: Text - single sentence from the texts of Kuprin A.I. and sentences geenerated with RuGPT3. For True sentences was used: Crime and Punishment. Fyodor Mikhailovich Dostoevsky Poor Folk. Fyodor Mikhailovich Dostoevsky The Idiot. Fyodor Mikhailovich Dostoevsky Demons. Fyodor Mikhailovich Dostoevsky The Brothers… See the full description on the dataset page: https://huggingface.co/datasets/Roaoch/CyberClassic_True.False.texttext-classification10K<n<100K0 likes14 downloads2y agoHugging Face18roanbrasil /k8s-rag-corpus K8s RAG Corpus A curated corpus of 4,794 Kubernetes-specific documents (7.5MB) used as the retrieval index for the BM25 RAG component of the K8s Multi-Agent Debate system. Contents Source Documents Official k8s.io examples (kubernetes/website) ~200 Helm chart examples ~150 Flux CD / HelmRelease examples ~100 ArgoCD Application templates ~100 RBAC, NetworkPolicy, HPA patterns ~200 Curated hardcoded examples (29 complex patterns) 29 Local production… See the full description on the dataset page: https://huggingface.co/datasets/roanbrasil/k8s-rag-corpus.texttext-generation100K<n<1M0 likes12 downloads5mo agoHugging Face19silk-road /Haruhi-Dialogue-Speaker-Extract-And-Summary之前的 silk-road/Haruhi-Dialogue-Speaker-Extract 要求模型输出json格式,并且采取了CoT策略,感觉有一些难了 这一次把总结和抽取拆分成了两个任务 并且抽取的格式改为了csv格式。 texttext-generationn<1K0 likes11 downloads3y agoHugging Face20Fleur-roar /Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_TextsHere presented a partially synthesized dataset, developed utilizing the GPT-4 model, for the purpose of NLG, particulary for the task of hierarchical generation of longer texts from short summaries. The creation of this dataset was undertaken as a component of my thesis paper. It incorporates excerpts from prominent British and American novels, from which plots, summaries, and metadata have been derived using GPT-4 API to facilitate extensive future research. The metadata included in the… See the full description on the dataset page: https://huggingface.co/datasets/Fleur-roar/Thesis_Development_of_a_Complex_of_Neural_Networks_for_Linked_Generation_of_Large_Texts.texttext-generation1K<n<10K1 likes6 downloads2y agoHugging Face21AitijhyaR /smolified-roast-your-design 🤏 smolified-roast-your-design Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model AitijhyaR/smolified-roast-your-design. 📦 Asset Details Origin: Smolify Foundry (Job ID: 17a09ac9) Records: 93 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by AitijhyaR. Generated via Smolify.ai. texttext-generationn<1K0 likes5 downloads6mo agoHugging Face22Good-News-Lending /first-time-buyer-roadmap-2026 First Time Buyer Roadmap 2026 55+ step-by-step homebuyer roadmap entries by income/credit scenario. Details Records: 55 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Beau Thompson, NMLS #1615561 Publisher: Good News Lending Thompson Alpha Logic Personalized roadmaps by income bracket and credit tier, calculating savings milestones, credit improvement timelines, and document preparation checklists. Each pathway routes to the optimal… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/first-time-buyer-roadmap-2026.textquestion-answeringn<1K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.