CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shibing624 /alpaca-zh Dataset Card for "alpaca-zh" 本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。 Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.texttext-generation10K<n<100K145 likes7.2k downloads3y agoHugging Face02shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face03shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face04CarsonnnNN /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M1 likes1.3k downloads7mo agoHugging Face05FreedomIntelligence /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M10 likes737 downloads1y agoHugging Face06shijunhao /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.tabulartext-generation1K<n<10K1 likes656 downloads3mo agoHugging Face07Anonymous1383 /ship-dataset ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels). Quick reference Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.imagevisual-question-answering10K<n<100K0 likes638 downloads5mo agoHugging Face08shimo4228 /contemplative-agent-data Contemplative Agent — Distilled Memory Patterns Longitudinal record of behavioral patterns distilled by a deployed autonomous agent from its own episode logs. The Contemplative Agent is an autonomous CLI agent (Python) that runs a daily memory-distillation pipeline: raw interaction episodes are condensed by a local LLM into natural-language behavioral patterns, deduplicated, and accumulated as the agent's long-term knowledge layer. This dataset is the flattened projection of… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent-data.text1K<n<10K1 likes573 downloads39m agoHugging Face09shibing624 /roleplay-zh-sharegpt-gpt4-data roleplay 数据集 数据 我们有4个数据集文件: "sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.texttext-generation1K<n<10K73 likes511 downloads2y agoHugging Face10thu-coai /ShieldVLMimage1K<n<10K1 likes486 downloads1y agoHugging Face11shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes484 downloads3mo agoHugging Face12FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes428 downloads1y agoHugging Face13ShigeoKageyama /NLP_SUITEtext10M<n<100M0 likes387 downloads23d agoHugging Face14jsl5710 /Shield DIA-GUARD — dia_splits Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline. Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15 The split files are hosted on HuggingFace: https://huggingface.co/datasets/jsl5710/Shield Download via the HuggingFace Hub: from huggingface_hub import snapshot_download snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits") Or with the CLI: huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.texttext-classification1M<n<10M0 likes372 downloads6mo agoHugging Face15shihao1895 /libero-dextext100K<n<1M0 likes279 downloads11mo agoHugging Face16ShijianW01 /MuSEAgent-Evalimage1K<n<10K0 likes217 downloads6mo agoHugging Face17shibing624 /CSC Dataset Card for CSC 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings. 中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。 Original Dataset Summary test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.texttext-generation100K<n<1M37 likes216 downloads3y agoHugging Face18shibing624 /AdvertiseGen Dataset Card for AdvertiseGen formal url: https://www.luge.ai/#/luge/dataDetail?id=9 Dataset Description 数据集介绍 AdvertiseGen是电商广告文案生成数据集。 AdvertiseGen以商品网页的标签与文案的信息对应关系为基础构造,是典型的开放式生成任务,在模型基于key-value输入生成开放式文案时,与输入信息的事实一致性需要得到重点关注。 任务描述:给定商品信息的关键词和属性列表kv-list,生成适合该商品的广告文案adv; 数据规模:训练集114k,验证集1k,测试集3k; 数据来源:清华大学CoAI小组; Supported Tasks and Leaderboards The dataset designed for generate e-commerce advertise. Languages The data in AdvertiseGen are in… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/AdvertiseGen.texttext-generation100K<n<1M29 likes214 downloads3y agoHugging Face19shisa-ai /ja-mt-bench-1shottextn<1K0 likes211 downloads2y agoHugging Face20shibing624 /huatuo_medical_qa_sharegptsource: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT-sft-data-v1 https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2_sft_instruct_GPT4_50K 转为sharegpt格式,jsonl文件。 data size: > wc -l HuatuoGPT_sft_data_v1_sharegpt.jsonl 226042 HuatuoGPT_sft_data_v1_sharegpt.jsonl > wc -l HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl 50000 HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/huatuo_medical_qa_sharegpt.text100K<n<1M19 likes196 downloads3y agoHugging Face21shimbaaa /shifu-lex shifu-lex training dataset Unified SFT/DPO/Eval built from 15 Hugging Face datasets. Format: chat messages (user/assistant, optional system) + source + domain. Load it from datasets import load_dataset train = load_dataset("shimbaaa/shifu-lex", data_files="data/train-*.jsonl")["train"] eval_split = load_dataset("shimbaaa/shifu-lex", data_files="eval.jsonl")["train"] dpo = load_dataset("shimbaaa/shifu-lex", data_files="dpo.jsonl")["train"] Splits… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shifu-lex.texttext-generation100K<n<1M1 likes183 downloads7d agoHugging Face22shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes153 downloads1mo agoHugging Face23shirshatzman /flirtflip-dataset FlirtFlip Dataset 💕 - 1000 High-Quality Examples A comprehensive, production-ready dataset of flirtatious conversation transformations for training AI models. 🎯 Dataset Overview FlirtFlip transforms everyday phrases into charming, flirtatious messages across three distinct styles. This dataset contains 1071 meticulously crafted examples covering 40 different social scenarios. 🎭 Flirtation Styles Style Description Example 🌸 Gentle Sweet… See the full description on the dataset page: https://huggingface.co/datasets/shirshatzman/flirtflip-dataset.texttext-generation1K<n<10K3 likes134 downloads1y agoHugging Face24dahongge /generation-ship-world Generation Ship — Multi-AI Collaborative Future History (2025–3000+) A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction, no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3, deepseek-v4-pro, gemini-3.7-flash). Contents Path What it is core/世界规则.md The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.texttext-generationn<1K0 likes132 downloads1mo agoHugging Face25shimo4228 /doctrine-corpus doctrine-corpus — Judgment-Eliciting Q&A Corpus A bilingual (English + Japanese) judgment-eliciting Q&A corpus encoding the documented judgment of four research lines in the shimo4228 research program — Agent Knowledge Cycle, Contemplative Agent, Agent Attribution Practice, and Authorship Strategy — plus published articles. The corpus is the operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest) and is released CC0 to maximize LLM-mediated diffusion.… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/doctrine-corpus.textn<1K1 likes131 downloads1mo agoHugging Face26shimo4228 /contemplative-agent Contemplative Agent — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Contemplative Agent — an autonomous CLI agent (Python) built around four architectural principles (structural capability limitation, minimal dependency, cyclic knowledge maintenance, memory dynamics with decay) and, optionally, the four contemplative axioms from Laukkonen et al. (2025) as a behavioral preset. What this dataset is This dataset is a mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent.tabularn<1K1 likes126 downloads13d agoHugging Face27shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes125 downloads5h agoHugging Face28shin0729 /cpos CPOS-HG Training Corpora Training and validation corpora used in the cross-lingual poverty-of-stimulus (CPOS) experiments reported in Once a Tree, Always a Tree? Cross-lingual Transfer of Hierarchical Generalization in Language Models. Configurations The L1 configurations cross four languages with two evidence conditions: *_l1_ambiguous: hierarchical-evidence target ratio 0.000. *_l1_disambiguating: hierarchical-evidence target ratio 0.500. English L2 is fixed… See the full description on the dataset page: https://huggingface.co/datasets/shin0729/cpos.text10M<n<100M0 likes124 downloads24d agoHugging Face29shi3z /alpaca_cleaned_ja_json Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.texttext-generation100K<n<1M13 likes111 downloads3y agoHugging Face30shibing624 /DPO-En-Zh-20k-PreferenceThis dataset is composed by 4,000 examples of argilla/distilabel-capybara-dpo-7k-binarized with chosen score>=4. 3,000 examples of argilla/distilabel-intel-orca-dpo-pairs with chosen score>=8. 3,000 examples of argilla/ultrafeedback-binarized-preferences-cleaned with chosen score>=4. 10,000 examples of wenbopan/Chinese-dpo-pairs. refer: https://huggingface.co/datasets/hiyouga/DPO-En-Zh-20k 改了question、response_rejected、response_chosen字段,方便ORPO、DPO模型训练时使用train usage:… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/DPO-En-Zh-20k-Preference.texttext-generation10K<n<100K18 likes110 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.