CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Chenyu-Zhou /OR-Space OR-Space A full-lifecycle workspace benchmark for industrial optimization agents. OR-Space evaluates whether language-model agents can work reliably with operations research problems represented as executable, multi-file workspaces. Rather than presenting a self-contained mathematical prompt, each task distributes evidence across business requirements, structured data, source code, execution logs, and solver records. The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.textquestion-answeringn<1K4 likes633 downloads2mo agoHugging Face02Yang-Zhou /DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill A high-quality Chain-of-Thought (CoT) dataset generated using Qwen/Qwen3-235B-A22B-Thinking-2507 with rejection sampling on BytedTsinghua-SIA/DAPO-Math-17k. This dataset is ideal for SFT distillation training to improve mathematical reasoning capabilities of models. The dataset format is compatible with LLaMA-Factory for efficient SFT training. Files dapo_distill_boxed.json: Single sampling subset (15,129… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill.text-generation100K<n<1M2 likes392 downloads11mo agoHugging Face03zhourax977 /Mix-ECom Mix-ECom Dataset Data files for Mix-ECom, a benchmark for evaluating LLM-based agents on e-commerce customer service tasks. See the code repository and paper for the agent, evaluation harness, and usage instructions. Contents . ├── eval/ │ ├── eval_gt_a.json # After-sales split (91 samples, text + image) │ ├── eval_gt_d.json # During-sales split (100 samples, text + video) │ └── eval_gt_l.json # Logistics split (108 samples, text) ├── database/… See the full description on the dataset page: https://huggingface.co/datasets/zhourax977/Mix-ECom.text-generation1K<n<10K0 likes234 downloads2mo agoHugging Face04agentlans /ZhonghuaChat Zhonghua Chat Overview Zhonghua Chat is a multilingual dataset featuring questions and answers in: Mandarin (Simplified Chinese and Traditional Chinese) Cantonese English The dataset is designed to support research and development in conversational AI, natural language understanding, and multilingual language modeling for Chinese dialects and writing systems. Composition Rows were sampled in roughly equal proportions from the following high-quality… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ZhonghuaChat.texttext-generation100K<n<1M0 likes79 downloads1y agoHugging Face05BabyLM-community /babylm-zhogated BabyLM Dataset Dataset Description This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm Dataset Summary Language: zho Script: Hani, Hans, Latn Tier: > 100M Byte Premium Factor: 0.935966 Size (MB): 518.85 Expected Size (MB): 508.23 Number of Documents: 203,891 Total Tokens: 137,835,046 Tokenizer: Qwen/Qwen3-0.6B Tokens Per Category child-available-speech: 7,403,441 tokens child-books: 15… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.texttext-generation100K<n<1M1 likes76 downloads8mo agoHugging Face06zhongyi-zhou /toolgrad-500 🛠️ ToolGrad-500: A Tool-Use Dataset Generated by ToolGrad (ACL 2026 Findings) ToolGrad-500 is a synthetic tool-use dataset targeting single-turn parallel function calling scenarios. The dataset is generated using the ToolGrad framework (ACL 2026 Findings), leveraging textual gradients to optimize tool selection and calling. 📊 Dataset Structure This dataset consists of formatted single-turn conversations suitable for standard conversational trainers.… See the full description on the dataset page: https://huggingface.co/datasets/zhongyi-zhou/toolgrad-500.texttext-generationn<1K2 likes75 downloads4mo agoHugging Face07zhongqy /RMCBench RMCBench Benchmarking Large Language Models’ Resistance to Malicious Code Generation Prompts ██████╗ ███╗ ███╗ ██████╗██████╗ ███████╗███╗ ██╗ ██████╗██╗ ██╗ ██╔══██╗████╗ ████║██╔════╝██╔══██╗██╔════╝████╗ ██║██╔════╝██║ ██║ ██████╔╝██╔████╔██║██║ ██████╔╝█████╗ ██╔██╗ ██║██║ ███████║ ██╔══██╗██║╚██╔╝██║██║ ██╔══██╗██╔══╝ ██║╚██╗██║██║ ██╔══██║ ██║ ██║██║ ╚═╝ ██║╚██████╗██████╔╝███████╗██║ ╚████║╚██████╗██║ ██║ ╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝╚═════╝… See the full description on the dataset page: https://huggingface.co/datasets/zhongqy/RMCBench.tabulartext-generationn<1K3 likes70 downloads1y agoHugging Face08CMLM /ZhongJing-OMNI ZhongJing-OMNI: The First Multimodal Benchmark for Evaluating Traditional Chinese Medicine ZhongJing-OMNI is the first multimodal benchmark dataset designed to evaluate Traditional Chinese Medicine (TCM) knowledge in large language models. This dataset provides a diverse array of questions and multimodal data, combining visual and textual information to assess the model’s ability to reason through complex TCM diagnostic and therapeutic scenarios. The unique combination of TCM… See the full description on the dataset page: https://huggingface.co/datasets/CMLM/ZhongJing-OMNI.imagequestion-answeringn<1K2 likes68 downloads2y agoHugging Face09zhongweixie /inplace-ttt-results In-Place TTT - Experimental Results This dataset contains experimental results and training logs for the In-Place Test-Time Training project. 📦 Dataset Contents 1. Experimental Results (results_20260904.tar.gz) Size: 191MB (compressed from 1.7GB) Files: 1,999 files Contents: Language model evaluation results RULER benchmark outputs ProLong training metrics Various test configurations 2. WandB Training Logs (wandb_20260904.tar.gz)… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/inplace-ttt-results.text-generationn<1K0 likes61 downloads19d agoHugging Face10jared-zhou /TQA-Distill-R1 TQA-Distill-R1 TQA-Distill-R1 is a distilled dataset designed for training and evaluating large language models (LLMs) on Table Question Answering (TQA) tasks. It was created by prompting a local deployment of DeepSeek R1 using question-table pairs sourced from two public datasets: TQA-HiTab and TQA-WTQ.The dataset focuses on multi-step reasoning, aggregation, and table understanding abilities. 📑 Dataset Summary Each sample contains: A table (structured in JSON… See the full description on the dataset page: https://huggingface.co/datasets/jared-zhou/TQA-Distill-R1.texttable-question-answering10K<n<100K2 likes41 downloads1y agoHugging Face11zeaver /multifactor_squad1.1_zhou MultiFactor-HotpotQA-SuppFacts The MultiFactor datasets -- SQuAD1.1-Zhou Split [1] in EMNLP 2023 Findings: Improving Question Generation with Multi-level Content Planning. 1. Dataset Details 1.1 Dataset Description SQuAD1.1-Zhou Split [1, 2] in EMNLP 2023 Findings: Improving Question Generation with Multi-level Content Planning. Based on the dataset in [2], we add the p_hrase, n_phrase and full answer attributes for every dataset instance. The full answer… See the full description on the dataset page: https://huggingface.co/datasets/zeaver/multifactor_squad1.1_zhou.text-generation10K<n<100K0 likes38 downloads3y agoHugging Face12ZhongshengWang /Alpaca-pubmed-summarizationThis data set is a lightweight fine-tuned data format version of the Llama2 large language model for Stanford Alpaca. You can click here to view. cite original code @inproceedings{cohan-etal-2018-discourse, title = "A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents", author = "Cohan, Arman and Dernoncourt, Franck and Kim, Doo Soon and Bui, Trung and Kim, Seokhwan and Chang, Walter and Goharian, Nazli", booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-pubmed-summarization.textsummarization100K<n<1M4 likes30 downloads3y agoHugging Face13wobure /zhongyi-books zhongyi-books 简介 中医典籍,源远流长 本数据集包含 21 本中文书籍,以 Markdown 格式存储。 数据来源 来源网站: 劝学网 (https://www.quanxue.cn) 文件结构 中医典籍_books/ ├── _十二经脉_动画图.md ├── 中医伤科按摩学.md ├── 中医养生学.md ├── 中医眼科备读.md ├── 中药基本理论知识.md ├── 伤寒论.md ├── 千金翼方.md ├── 千金要方.md ├── 本草纲目.md ├── 神农本草经.md ├── 穴位.md ├── 经络.md ├── 肘后备急方.md ├── 腧穴.md ├── 自我调养巧治病.md ├── 邹孟城三十年临证经验集.md ├── 金匮要略.md ├── 针灸甲乙经.md ├── 黄帝八十一难经.md ├── 黄帝内经_灵枢.md ├── 黄帝内经_素问.md 使用方法 from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/wobure/zhongyi-books.text-generation1K<n<10K0 likes29 downloads1y agoHugging Face14bdx33 /ted-talks-in-chinese-zhongwen TED中文 Podcast 聚焦华语地区的创意,本节目从上万个TED和TEDx演讲中,为您精选中文演讲,以及少量中文配音的经典英语演讲。演讲人包括科技和人文专家、关心当下与未来的思考者、关注挑战与探索的实践者。英雄不论出处,谁有创意谁讲。让这些演讲成为一把把钥匙,开启你的好奇心,升级你的行动力。 Focusing on creativity within the Chinese-speaking world, this program curates Chinese-language talks from tens of thousands of TED and TEDx presentations, along with a select few English classics dubbed into Chinese. Our speakers span technology and humanities experts, thinkers engaged with the present and future, and practitioners… See the full description on the dataset page: https://huggingface.co/datasets/bdx33/ted-talks-in-chinese-zhongwen.texttext-generationn<1K1 likes28 downloads1y agoHugging Face15ZhongshengWang /Alpaca-cnn-dailymail Data Summary Data set Alpaca-cnn-dailymail is a data set version format changed by ccdv/cnn_dailymail to meet Alpaca fine-tuning Llama2. Only versions 3.0.0 and 2.0.0 were used for merging and as a key data set for the summary extraction task. Licensing Information The Alpaca-cnn-dailymail dataset version 1.0.0 is released under the Apache-2.0 License. Citation Information @inproceedings{see-etal-2017-get, title = "Get To The Point: Summarization with… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-cnn-dailymail.textsummarization100K<n<1M0 likes19 downloads3y agoHugging Face16hhhwmws /zhouzhiruo支持ChatHaruhi2 的周芷若数据,可以使用如下方式调用 from chatharuhi import ChatHaruhi chatbot = ChatHaruhi( role_from_hf = 'hhhwmws/zhouzhiruo', \ llm = 'openai') response = chatbot.chat(role='张无忌', text = '周芷若!') print(response) 上传者: 米唯实 更具体的信息,见 ChatHaruhi 欢迎加入我们的 众筹角色创建项目 Citation引用 Please cite the repo if you use the data or code in this repo. @misc{li2023chatharuhi, title={ChatHaruhi: Reviving Anime Character in Reality via Large Language Model}… See the full description on the dataset page: https://huggingface.co/datasets/hhhwmws/zhouzhiruo.texttext-generationn<1K0 likes12 downloads3y agoHugging Face17zhoutop1 /gsm8k Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/zhoutop1/gsm8k.texttext-generation10K<n<100K0 likes11 downloads3mo agoHugging Face18zhoushuuki /ChinaNB1text-generation1K<n<10K0 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.