datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OR-Space
OR-Space
A full-lifecycle workspace benchmark for industrial optimization agents.
OR-Space evaluates whether language-model agents can work reliably with
operations research problems represented as executable, multi-file workspaces.
Rather than presenting a self-contained mathematical prompt, each task
distributes evidence across business requirements, structured data, source
code, execution logs, and solver records.
The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill
A high-quality Chain-of-Thought (CoT) dataset generated using Qwen/Qwen3-235B-A22B-Thinking-2507 with rejection sampling on BytedTsinghua-SIA/DAPO-Math-17k. This dataset is ideal for SFT distillation training to improve mathematical reasoning capabilities of models.
The dataset format is compatible with LLaMA-Factory for efficient SFT training.
Files
dapo_distill_boxed.json: Single sampling subset (15,129… See the full description on the dataset page: https://huggingface.co/datasets/Yang-Zhou/DAPO-Math-17k-Qwen3-235B-A22B-Thinking-2507-rejection-distill.Mix-ECom
Mix-ECom Dataset
Data files for Mix-ECom, a benchmark for evaluating LLM-based agents on e-commerce customer service tasks. See the code repository and paper for the agent, evaluation harness, and usage instructions.
Contents
.
├── eval/
│ ├── eval_gt_a.json # After-sales split (91 samples, text + image)
│ ├── eval_gt_d.json # During-sales split (100 samples, text + video)
│ └── eval_gt_l.json # Logistics split (108 samples, text)
├── database/… See the full description on the dataset page: https://huggingface.co/datasets/zhourax977/Mix-ECom.ZhonghuaChat
Zhonghua Chat
Overview
Zhonghua Chat is a multilingual dataset featuring questions and answers in:
Mandarin (Simplified Chinese and Traditional Chinese)
Cantonese
English
The dataset is designed to support research and development in conversational AI, natural language understanding, and multilingual language modeling for Chinese dialects and writing systems.
Composition
Rows were sampled in roughly equal proportions from the following high-quality… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ZhonghuaChat.babylm-zho
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: zho
Script: Hani, Hans, Latn
Tier: > 100M
Byte Premium Factor: 0.935966
Size (MB): 518.85
Expected Size (MB): 508.23
Number of Documents: 203,891
Total Tokens: 137,835,046
Tokenizer: Qwen/Qwen3-0.6B
Tokens Per Category
child-available-speech: 7,403,441 tokens
child-books: 15… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-zho.toolgrad-500
🛠️ ToolGrad-500: A Tool-Use Dataset Generated by ToolGrad (ACL 2026 Findings)
ToolGrad-500 is a synthetic tool-use dataset targeting single-turn parallel function calling scenarios. The dataset is generated using the ToolGrad framework (ACL 2026 Findings), leveraging textual gradients to optimize tool selection and calling.
📊 Dataset Structure
This dataset consists of formatted single-turn conversations suitable for standard conversational trainers.… See the full description on the dataset page: https://huggingface.co/datasets/zhongyi-zhou/toolgrad-500.RMCBench
RMCBench
Benchmarking Large Language Models’ Resistance to Malicious Code Generation Prompts
██████╗ ███╗ ███╗ ██████╗██████╗ ███████╗███╗ ██╗ ██████╗██╗ ██╗
██╔══██╗████╗ ████║██╔════╝██╔══██╗██╔════╝████╗ ██║██╔════╝██║ ██║
██████╔╝██╔████╔██║██║ ██████╔╝█████╗ ██╔██╗ ██║██║ ███████║
██╔══██╗██║╚██╔╝██║██║ ██╔══██╗██╔══╝ ██║╚██╗██║██║ ██╔══██║
██║ ██║██║ ╚═╝ ██║╚██████╗██████╔╝███████╗██║ ╚████║╚██████╗██║ ██║
╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝╚═════╝… See the full description on the dataset page: https://huggingface.co/datasets/zhongqy/RMCBench.ZhongJing-OMNI
ZhongJing-OMNI: The First Multimodal Benchmark for Evaluating Traditional Chinese Medicine
ZhongJing-OMNI is the first multimodal benchmark dataset designed to evaluate Traditional Chinese Medicine (TCM) knowledge in large language models. This dataset provides a diverse array of questions and multimodal data, combining visual and textual information to assess the model’s ability to reason through complex TCM diagnostic and therapeutic scenarios. The unique combination of TCM… See the full description on the dataset page: https://huggingface.co/datasets/CMLM/ZhongJing-OMNI.inplace-ttt-results
In-Place TTT - Experimental Results
This dataset contains experimental results and training logs for the In-Place Test-Time Training project.
📦 Dataset Contents
1. Experimental Results (results_20260904.tar.gz)
Size: 191MB (compressed from 1.7GB)
Files: 1,999 files
Contents:
Language model evaluation results
RULER benchmark outputs
ProLong training metrics
Various test configurations
2. WandB Training Logs (wandb_20260904.tar.gz)… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/inplace-ttt-results.TQA-Distill-R1
TQA-Distill-R1
TQA-Distill-R1 is a distilled dataset designed for training and evaluating large language models (LLMs) on Table Question Answering (TQA) tasks. It was created by prompting a local deployment of DeepSeek R1 using question-table pairs sourced from two public datasets: TQA-HiTab and TQA-WTQ.The dataset focuses on multi-step reasoning, aggregation, and table understanding abilities.
📑 Dataset Summary
Each sample contains:
A table (structured in JSON… See the full description on the dataset page: https://huggingface.co/datasets/jared-zhou/TQA-Distill-R1.multifactor_squad1.1_zhou
MultiFactor-HotpotQA-SuppFacts
The MultiFactor datasets -- SQuAD1.1-Zhou Split [1] in EMNLP 2023 Findings: Improving Question Generation with Multi-level Content Planning.
1. Dataset Details
1.1 Dataset Description
SQuAD1.1-Zhou Split [1, 2] in EMNLP 2023 Findings: Improving Question Generation with Multi-level Content Planning.
Based on the dataset in [2], we add the p_hrase, n_phrase and full answer attributes for every dataset instance.
The full answer… See the full description on the dataset page: https://huggingface.co/datasets/zeaver/multifactor_squad1.1_zhou.Alpaca-pubmed-summarizationThis data set is a lightweight fine-tuned data format version of the Llama2 large language model for Stanford Alpaca. You can click here to view.
cite original code
@inproceedings{cohan-etal-2018-discourse,
title = "A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents",
author = "Cohan, Arman and
Dernoncourt, Franck and
Kim, Doo Soon and
Bui, Trung and
Kim, Seokhwan and
Chang, Walter and
Goharian, Nazli",
booktitle = "Proceedings… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-pubmed-summarization.zhongyi-books
zhongyi-books
简介
中医典籍,源远流长
本数据集包含 21 本中文书籍,以 Markdown 格式存储。
数据来源
来源网站: 劝学网 (https://www.quanxue.cn)
文件结构
中医典籍_books/
├── _十二经脉_动画图.md
├── 中医伤科按摩学.md
├── 中医养生学.md
├── 中医眼科备读.md
├── 中药基本理论知识.md
├── 伤寒论.md
├── 千金翼方.md
├── 千金要方.md
├── 本草纲目.md
├── 神农本草经.md
├── 穴位.md
├── 经络.md
├── 肘后备急方.md
├── 腧穴.md
├── 自我调养巧治病.md
├── 邹孟城三十年临证经验集.md
├── 金匮要略.md
├── 针灸甲乙经.md
├── 黄帝八十一难经.md
├── 黄帝内经_灵枢.md
├── 黄帝内经_素问.md
使用方法
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/wobure/zhongyi-books.ted-talks-in-chinese-zhongwen
TED中文 Podcast
聚焦华语地区的创意,本节目从上万个TED和TEDx演讲中,为您精选中文演讲,以及少量中文配音的经典英语演讲。演讲人包括科技和人文专家、关心当下与未来的思考者、关注挑战与探索的实践者。英雄不论出处,谁有创意谁讲。让这些演讲成为一把把钥匙,开启你的好奇心,升级你的行动力。
Focusing on creativity within the Chinese-speaking world, this program curates Chinese-language talks from tens of thousands of TED and TEDx presentations, along with a select few English classics dubbed into Chinese. Our speakers span technology and humanities experts, thinkers engaged with the present and future, and practitioners… See the full description on the dataset page: https://huggingface.co/datasets/bdx33/ted-talks-in-chinese-zhongwen.Alpaca-cnn-dailymail
Data Summary
Data set Alpaca-cnn-dailymail is a data set version format changed by ccdv/cnn_dailymail to meet Alpaca fine-tuning Llama2. Only versions 3.0.0 and 2.0.0 were used for merging and as a key data set for the summary extraction task.
Licensing Information
The Alpaca-cnn-dailymail dataset version 1.0.0 is released under the Apache-2.0 License.
Citation Information
@inproceedings{see-etal-2017-get,
title = "Get To The Point: Summarization with… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-cnn-dailymail.zhouzhiruo支持ChatHaruhi2 的周芷若数据,可以使用如下方式调用
from chatharuhi import ChatHaruhi
chatbot = ChatHaruhi( role_from_hf = 'hhhwmws/zhouzhiruo', \
llm = 'openai')
response = chatbot.chat(role='张无忌', text = '周芷若!')
print(response)
上传者: 米唯实
更具体的信息,见 ChatHaruhi
欢迎加入我们的 众筹角色创建项目
Citation引用
Please cite the repo if you use the data or code in this repo.
@misc{li2023chatharuhi,
title={ChatHaruhi: Reviving Anime Character in Reality via Large Language Model}… See the full description on the dataset page: https://huggingface.co/datasets/hhhwmws/zhouzhiruo.gsm8k
Dataset Card for GSM8K
Dataset Summary
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.
These problems take between 2 and 8 steps to solve.
Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/zhoutop1/gsm8k.ChinaNB1
