CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shibing624 /alpaca-zh Dataset Card for "alpaca-zh" 本数据集是参考Alpaca方法基于GPT4得到的self-instruct数据,约5万条。 Dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM It is the chinese dataset from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM/blob/main/data/alpaca_gpt4_data_zh.json Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only non-commercial use) and models trained using the dataset should not… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/alpaca-zh.texttext-generation10K<n<100K145 likes7.2k downloads3y agoHugging Face02shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face03shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face04CarsonnnNN /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M1 likes1.3k downloads7mo agoHugging Face05shijunhao /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.tabulartext-generation1K<n<10K1 likes692 downloads3mo agoHugging Face06FreedomIntelligence /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M10 likes673 downloads1y agoHugging Face07Anonymous1383 /ship-dataset ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on parametrically-generated ship structural drawings. Six commercial ship types × nine drawing-grounded sub-tasks × deterministic ground truth derived directly from the generator's input dictionary (no human annotation, no rule-citation labels). Quick reference Total candidates: 6{,}450 across 6 ship types (Tanker, VLCC, BULKC, CNTR, LNGC… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous1383/ship-dataset.imagevisual-question-answering10K<n<100K0 likes638 downloads5mo agoHugging Face08shimo4228 /contemplative-agent-data Contemplative Agent — Distilled Memory Patterns Longitudinal record of behavioral patterns distilled by a deployed autonomous agent from its own episode logs. The Contemplative Agent is an autonomous CLI agent (Python) that runs a daily memory-distillation pipeline: raw interaction episodes are condensed by a local LLM into natural-language behavioral patterns, deduplicated, and accumulated as the agent's long-term knowledge layer. This dataset is the flattened projection of… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent-data.text1K<n<10K1 likes587 downloads7h agoHugging Face09shibing624 /roleplay-zh-sharegpt-gpt4-data roleplay 数据集 数据 我们有4个数据集文件: "sharegpt_formatted_data-evol-gpt4.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-evol-male-gpt35.jsonl" 来自 bai-roleplay/evol-character-entire 将其转换为sharegpt格式。 "sharegpt_formatted_data-roleplay-chat-1k.jsonl" 来自 Minami-su/roleplay_multiturn_chat_1k_zh_v0.1 将其转换为sharegpt格式。… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/roleplay-zh-sharegpt-gpt4-data.texttext-generation1K<n<10K73 likes521 downloads2y agoHugging Face10thu-coai /ShieldVLMimage1K<n<10K1 likes486 downloads1y agoHugging Face11shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes483 downloads3mo agoHugging Face12FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes420 downloads1y agoHugging Face13jsl5710 /Shield DIA-GUARD — dia_splits Canonical train/val/test splits for the DIA-GUARD safety-guard training pipeline. Generated on 2026-03-30 | Seed: 42 | Ratios: 70 / 15 / 15 The split files are hosted on HuggingFace: https://huggingface.co/datasets/jsl5710/Shield Download via the HuggingFace Hub: from huggingface_hub import snapshot_download snapshot_download(repo_id="jsl5710/Shield", repo_type="dataset", local_dir="dataset/dia_splits") Or with the CLI: huggingface-cli download… See the full description on the dataset page: https://huggingface.co/datasets/jsl5710/Shield.texttext-classification1M<n<10M0 likes387 downloads6mo agoHugging Face14ShigeoKageyama /NLP_SUITEtext10M<n<100M0 likes387 downloads22d agoHugging Face15Shiki42 /piperx-workpiece-storage-0909-62ep-raw piperx-workpiece-storage-0909-62ep-raw 61 retained manually collected episodes recorded with EvoMind on 2026-09-09 (UTC+8). Task: put the copper screw into the left box and the black sleeves into the right box. LeRobot v3.0; 30 FPS; 53,347 frames; 1,778.2333 seconds. Robot: bi_piperx_follower. Three original 640x480 RGB video views: left_wrist, right_wrist, right_environment_1. Merged in chronological session order. Original episode 26 (19 frames) was removed on 2026-09-15.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-workpiece-storage-0909-62ep-raw.tabularroboticsn<1K0 likes338 downloads10d agoHugging Face16shihao1895 /libero-dextext100K<n<1M0 likes277 downloads11mo agoHugging Face17shibing624 /AdvertiseGen Dataset Card for AdvertiseGen formal url: https://www.luge.ai/#/luge/dataDetail?id=9 Dataset Description 数据集介绍 AdvertiseGen是电商广告文案生成数据集。 AdvertiseGen以商品网页的标签与文案的信息对应关系为基础构造,是典型的开放式生成任务,在模型基于key-value输入生成开放式文案时,与输入信息的事实一致性需要得到重点关注。 任务描述:给定商品信息的关键词和属性列表kv-list,生成适合该商品的广告文案adv; 数据规模:训练集114k,验证集1k,测试集3k; 数据来源:清华大学CoAI小组; Supported Tasks and Leaderboards The dataset designed for generate e-commerce advertise. Languages The data in AdvertiseGen are in… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/AdvertiseGen.texttext-generation100K<n<1M29 likes232 downloads3y agoHugging Face18ShijianW01 /MuSEAgent-Evalimage1K<n<10K0 likes217 downloads6mo agoHugging Face19shibing624 /CSC Dataset Card for CSC 中文拼写纠错数据集 Repository: https://github.com/shibing624/pycorrector Dataset Description Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. CSC is challenging since many Chinese characters are visually or phonologically similar but with quite different semantic meanings. 中文拼写纠错数据集,共27万条,是通过原始SIGHAN13、14、15年数据集和Wang271k数据集合并整理后得到,json格式,带错误字符位置信息。 Original Dataset Summary test.json 和… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/CSC.texttext-generation100K<n<1M37 likes214 downloads3y agoHugging Face20shisa-ai /ja-mt-bench-1shottextn<1K0 likes208 downloads2y agoHugging Face21shibing624 /huatuo_medical_qa_sharegptsource: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT-sft-data-v1 https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2_sft_instruct_GPT4_50K 转为sharegpt格式,jsonl文件。 data size: > wc -l HuatuoGPT_sft_data_v1_sharegpt.jsonl 226042 HuatuoGPT_sft_data_v1_sharegpt.jsonl > wc -l HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl 50000 HuatuoGPT2_sft_instruct_GPT4_sharegpt.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/huatuo_medical_qa_sharegpt.text100K<n<1M19 likes195 downloads3y agoHugging Face22shimbaaa /shifu-lex shifu-lex training dataset Unified SFT/DPO/Eval built from 15 Hugging Face datasets. Format: chat messages (user/assistant, optional system) + source + domain. Load it from datasets import load_dataset train = load_dataset("shimbaaa/shifu-lex", data_files="data/train-*.jsonl")["train"] eval_split = load_dataset("shimbaaa/shifu-lex", data_files="eval.jsonl")["train"] dpo = load_dataset("shimbaaa/shifu-lex", data_files="dpo.jsonl")["train"] Splits… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shifu-lex.texttext-generation100K<n<1M1 likes179 downloads6d agoHugging Face23shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes162 downloads1mo agoHugging Face24shimo4228 /doctrine-corpus doctrine-corpus — Judgment-Eliciting Q&A Corpus A bilingual (English + Japanese) judgment-eliciting Q&A corpus encoding the documented judgment of four research lines in the shimo4228 research program — Agent Knowledge Cycle, Contemplative Agent, Agent Attribution Practice, and Authorship Strategy — plus published articles. The corpus is the operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest) and is released CC0 to maximize LLM-mediated diffusion.… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/doctrine-corpus.textn<1K1 likes135 downloads1mo agoHugging Face25dahongge /generation-ship-world Generation Ship — Multi-AI Collaborative Future History (2025–3000+) A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction, no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3, deepseek-v4-pro, gemini-3.7-flash). Contents Path What it is core/世界规则.md The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.texttext-generationn<1K0 likes132 downloads1mo agoHugging Face26shirshatzman /flirtflip-dataset FlirtFlip Dataset 💕 - 1000 High-Quality Examples A comprehensive, production-ready dataset of flirtatious conversation transformations for training AI models. 🎯 Dataset Overview FlirtFlip transforms everyday phrases into charming, flirtatious messages across three distinct styles. This dataset contains 1071 meticulously crafted examples covering 40 different social scenarios. 🎭 Flirtation Styles Style Description Example 🌸 Gentle Sweet… See the full description on the dataset page: https://huggingface.co/datasets/shirshatzman/flirtflip-dataset.texttext-generation1K<n<10K3 likes130 downloads1y agoHugging Face27shimo4228 /contemplative-agent Contemplative Agent — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Contemplative Agent — an autonomous CLI agent (Python) built around four architectural principles (structural capability limitation, minimal dependency, cyclic knowledge maintenance, memory dynamics with decay) and, optionally, the four contemplative axioms from Laukkonen et al. (2025) as a behavioral preset. What this dataset is This dataset is a mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent.tabularn<1K1 likes127 downloads13d agoHugging Face28shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes126 downloads24d agoHugging Face29shizhuo2 /omega-het-expandA-sft OMEGA-HET-expandA — matched HET-vs-HOM SFT (equal-size) Matched supervised-fine-tuning data for the OMEGA diversity experiment: for each math prompt, reasoning trajectories are sampled two ways and only prompts solved (math-verified correct) in both conditions are kept (matched HOM∩HET = 3,219 prompts), so HET and HOM are directly comparable. HET (heterogeneous): true token-level continuation across a 3×32B roster (Qwen3-32B + DeepSeek-R1-Distill-Qwen-32B +… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-expandA-sft.text100K<n<1M0 likes125 downloads3mo agoHugging Face30shin0729 /cpos CPOS-HG Training Corpora Training and validation corpora used in the cross-lingual poverty-of-stimulus (CPOS) experiments reported in Once a Tree, Always a Tree? Cross-lingual Transfer of Hierarchical Generalization in Language Models. Configurations The L1 configurations cross four languages with two evidence conditions: *_l1_ambiguous: hierarchical-evidence target ratio 0.000. *_l1_disambiguating: hierarchical-evidence target ratio 0.500. English L2 is fixed… See the full description on the dataset page: https://huggingface.co/datasets/shin0729/cpos.text10M<n<100M0 likes124 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.