CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shibing624 /alpaca-zh Dataset Card for "alpaca-zh" 本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。 Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.texttext-generation10K<n<100K145 likes7k downloads3y agoHugging Face02shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face03shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face04CarsonnnNN /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M1 likes1.3k downloads7mo agoHugging Face05FreedomIntelligence /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M10 likes1.2k downloads1y agoHugging Face06shijunhao /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.tabulartext-generation1K<n<10K1 likes644 downloads3mo agoHugging Face07shibing624 /roleplay-zh-sharegpt-gpt4-data roleplay 数据集 数据 我们有4个数据集文件: "sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.texttext-generation1K<n<10K73 likes514 downloads2y agoHugging Face08shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes470 downloads3mo agoHugging Face09FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes430 downloads1y agoHugging Face10shibing624 /CSC Dataset Card for CSC 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings. 中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。 Original Dataset Summary test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.texttext-generation100K<n<1M37 likes222 downloads3y agoHugging Face11shibing624 /AdvertiseGen Dataset Card for AdvertiseGen formal url: https://www.luge.ai/#/luge/dataDetail?id=9 Dataset Description 数据集介绍 AdvertiseGen是电商广告文案生成数据集。 AdvertiseGen以商品网页的标签与文案的信息对应关系为基础构造,是典型的开放式生成任务,在模型基于key-value输入生成开放式文案时,与输入信息的事实一致性需要得到重点关注。 任务描述:给定商品信息的关键词和属性列表kv-list,生成适合该商品的广告文案adv; 数据规模:训练集114k,验证集1k,测试集3k; 数据来源:清华大学CoAI小组; Supported Tasks and Leaderboards The dataset designed for generate e-commerce advertise. Languages The data in AdvertiseGen are in… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/AdvertiseGen.texttext-generation100K<n<1M29 likes216 downloads3y agoHugging Face12shimbaaa /shifu-lex shifu-lex training dataset Unified SFT/DPO/Eval built from 15 Hugging Face datasets. Format: chat messages (user/assistant, optional system) + source + domain. Load it from datasets import load_dataset train = load_dataset("shimbaaa/shifu-lex", data_files="data/train-*.jsonl")["train"] eval_split = load_dataset("shimbaaa/shifu-lex", data_files="eval.jsonl")["train"] dpo = load_dataset("shimbaaa/shifu-lex", data_files="dpo.jsonl")["train"] Splits… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shifu-lex.texttext-generation100K<n<1M1 likes187 downloads8d agoHugging Face13shirshatzman /flirtflip-dataset FlirtFlip Dataset 💕 - 1000 High-Quality Examples A comprehensive, production-ready dataset of flirtatious conversation transformations for training AI models. 🎯 Dataset Overview FlirtFlip transforms everyday phrases into charming, flirtatious messages across three distinct styles. This dataset contains 1071 meticulously crafted examples covering 40 different social scenarios. 🎭 Flirtation Styles Style Description Example 🌸 Gentle Sweet… See the full description on the dataset page: https://huggingface.co/datasets/shirshatzman/flirtflip-dataset.texttext-generation1K<n<10K3 likes144 downloads1y agoHugging Face14dahongge /generation-ship-world Generation Ship — Multi-AI Collaborative Future History (2025–3000+) A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction, no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3, deepseek-v4-pro, gemini-3.7-flash). Contents Path What it is core/世界规则.md The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.texttext-generationn<1K0 likes132 downloads1mo agoHugging Face15shibing624 /DPO-En-Zh-20k-PreferenceThis dataset is composed by 4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4. 3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8. 3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4. 10,000 examples of wenbopan/Chinese-dpo-pairs. refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.texttext-generation10K<n<100K18 likes113 downloads2y agoHugging Face16shi3z /alpaca_cleaned_ja_json Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.texttext-generation100K<n<1M13 likes111 downloads3y agoHugging Face17CarsonnnNN /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M0 likes108 downloads7mo agoHugging Face18ShivomH /Mental-Health-Conversations Dataset Card This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic". Source jerryjalapeno/nart-100k-synthetic texttext-generation10K<n<100K4 likes106 downloads1y agoHugging Face19shizhuo2 /omega-het-sft-rl OMEGA-HET-SFT-RL Reasoning-trajectory corpora for studying whether heterogeneous (HET) multi-model SFT data improves post-RL out-of-distribution generalization on OMEGA math vs homogeneous (HOM) single-model data, under matched controls. Conditions HOM: trajectories generated by a single model (Qwen3-4B). HET: trajectories composed via true token-level continuation across a roster of 7 reasoning models (each model resumes the previous model's own assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-sft-rl.texttext-generation100K<n<1M0 likes94 downloads4mo agoHugging Face20wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes93 downloads1mo agoHugging Face21shimbaaa /shimbabomb-ai-benchmark ShimbaBomb AI Benchmark A GSM8K-style benchmark dataset for evaluating AI models on ShimbaBomb — an English-like scripting language that compiles to native C. Overview This dataset contains 81 problems with chain-of-thought reasoning for training and evaluating AI models on ShimbaBomb code generation, understanding, debugging, and explanation. Each problem has: question: A natural language description of a programming task or question about SB code answer:… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shimbabomb-ai-benchmark.texttext-generationn<1K1 likes92 downloads27d agoHugging Face22shirochange /kansaiben 関西弁会話データセット (Kansai-ben Dialogue Dataset) 関西弁(大阪弁)で応答するAIのファインチューニング用instruction/outputデータセットです。 データセット概要 標準語の入力に対して、関西弁で自然に返答するサンプルを収録しています。LLMを関西弁キャラクターとしてファインチューニングすることを主な用途として想定しています。 データ構造 カラム名 型 説明 id int レコードのID instruction string 標準語の入力文 output string 関西弁の応答文 サンプル id instruction output 1 こんにちは! まいど!元気しとるか? 9 なぜ空は青いの? ええ質問やな。太陽の光が空気の粒にぶつかって、青い光がよう散らばるからやで。 50 関西弁を教えて. 「おおきに」は「ありがとう」、「ほんま」は「本当に」や。覚えときや。… See the full description on the dataset page: https://huggingface.co/datasets/shirochange/kansaiben.texttext-generation1K<n<10K4 likes90 downloads6mo agoHugging Face23formal-umd /T-SHIRT T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning Yanjun Fu, Faisal Hamman, Sanghamitra Dutta [📖 Paper] Overview Abstract Instruction tuning is essential for Large Language Models (LLMs) to effectively follow user instructions. To improve training efficiency and reduce data redundancy, recent works use LLM-based scoring functions, e.g., Instruction-Following Difficulty (IFD), to select high–quality instruction-tuning data with scores above a… See the full description on the dataset page: https://huggingface.co/datasets/formal-umd/T-SHIRT.texttext-generation10K<n<100K0 likes66 downloads10mo agoHugging Face24EmmaLeonhart /shinto-wikidata-qa Shinto Wikidata QA Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata structured facts. Built for the Adaption Labs AutoScientist Challenge (All Other Domains track). Credit: Adaptive Data by Adaption. Source & license Source: Wikidata Query Service (https://query.wikidata.org). All statement data is CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.textquestion-answering100K<n<1M0 likes60 downloads3mo agoHugging Face25ShivomH /MedCOT-Reason Important Note: I do not claim this dataset as my own. The entire credit belongs to the sources shared below. This dataset is simply preprocessed and formatted to align with the task of fine-tuning meta-llama/Llama-3.2-3B-Instruct Introduction This dataset is used to fine-tune ShivomH/Vitalis-Llama3-Reason, a smart medical LLM designed for advanced medical reasoning. This dataset is constructed using GPT-4o. Sources FreedomIntelligence/medical-o1-reasoning-SFT View… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MedCOT-Reason.textquestion-answering10K<n<100K1 likes55 downloads1y agoHugging Face26bcywinski /taboo-ship taboo-ship This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/taboo-ship") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generationn<1K0 likes49 downloads1y agoHugging Face27ITcoder /SHIFT_Training_Data SHIFT Training Data This repository contains the training data for SHIFT, presented in the paper SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation. Repository: https://github.com/OpenBMB/SHIFT Paper: https://arxiv.org/abs/2606.27786 Dataset Description SHIFT is a lightweight framework for resolving knowledge conflicts in retrieval-augmented generation (RAG). Instead of directly editing internal neurons… See the full description on the dataset page: https://huggingface.co/datasets/ITcoder/SHIFT_Training_Data.texttext-generation10K<n<100K1 likes49 downloads3mo agoHugging Face28shibing624 /CSC-gpt4 Dataset Card for Chinese Spelling Correction(gpt4 fixed version) 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings.… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC-gpt4.texttext-generation1K<n<10K4 likes46 downloads2y agoHugging Face29ShivomH /MentalHealth-Support Important Note This dataset is created from merging two datasets from different sources and has been formatted according to the "messages", "role", "content" chat format. I do not claim any ownership of this dataset. Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and highlight those limitations to users in a responsible manner. Since Mental… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MentalHealth-Support.texttext-generation10K<n<100K2 likes38 downloads1y agoHugging Face30ticoAg /shibing624-medical-pretrain Dataset Card for medical 中文医疗数据集 LLM Supervised Finetuning repository: https://github.com/shibing624/textgen MeidcalGPT repository: https://github.com/shibing624/MedicalGPT Dataset Description medical is a Chinese Medical dataset. 医疗数据集,可用于医疗领域大模型训练。 tree medical |-- finetune # 监督微调数据集,可用于SFT和RLHF | |-- test_en_1.json | |-- test_zh_0.json | |-- train_en_1.json | |-- train_zh_0.json | |-- valid_en_1.json | `-- valid_zh_0.json |-- medical.py # hf dataset… See the full description on the dataset page: https://huggingface.co/datasets/ticoAg/shibing624-medical-pretrain.texttext-generation100K<n<1M12 likes33 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.