datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese-classical-corpus
Chinese Classical Corpus
🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件
🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样)
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
为什么做这个
中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。
Tokenizer 层面 —— 真成立
7 个主流 tokenizer 横评(tokenizer_study):
DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.chinese-classical-corpus
Chinese Classical Corpus
🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite.
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
Quick Start
from datasets import load_dataset
# 源语料 (12,005 条章节级记录, 17.2M 字)
corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train")
# 古译今 / 今译古 双向指令数据 (1,924,378 条)
translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.classic-eda-c-trajectories
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 200 rounds.
Every model turn is one row, including the ones that went nowhere.
This is a partial snapshot. 130 of 1… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.test-subset-classic-eda
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.chinese-classical-bench
Chinese Classical Bench
中国古典语言能力评测基准 — 6 个任务 × 100 题 = 600 道,覆盖翻译、断句、字义、典故、续写填空、现代→文言压缩。
📊 在线排行榜: 🤗 Space — chinese-classical-bench-leaderboard
🔗 评测代码 & runner: github.com/gujilab/chinese-classical-bench — eval runner(OpenAI 兼容端点)、打分器、排行榜聚合脚本
📦 配套语料集: gujilab/chinese-classical-corpus (CC0 公有领域) — 题目均从该语料抽样生成
为什么做这个
中文(尤其文言文)常被说成"高密度优势"。这套基础设施(bench + corpus + 4 个论点实证实验)想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。
Tokenizer 层面(实证)
7 个主流 tokenizer 横评(详见下方… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-bench.pmjay-classifier-sft
PM-JAY Health Benefit Package Classifier
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Medical specialty + procedure → PM-JAY HBP code, package name, and rate
Why download this
Automate PM-JAY / Ayushman Bharat claim processing. Map procedures to Health Benefit Packages for pre-authorization and reimbursement.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pmjay-classifier-sft.classic-eda
Classic EDA - Period Software Task Specifications
1020 specifications for software that plausibly could have been written between
1985 and 1996, mined from two in-era archives and shaped as coding tasks with
machine-checkable requirements.
Each record names a program, describes it in a paragraph, and states 3-12 atomic
requirements plus an explicit interface contract (argv, stdin, stdout, exit
codes) so a grader can test an implementation.
Why period software… See the full description on the dataset page: https://huggingface.co/datasets/gdiamos/classic-eda.classical-chinese-poetry-benchmark-70
English Readme see below
(README由Claude 3.5 Sonnet生成)
中国古诗词大模型评测基准
简介
这是一个专门用于评测大语言模型在中国古诗词理解和生成方面能力的基准测试集。该基准包含了一个多样化的测试数据集和完整的评测框架,可用于系统性地评估和比较不同模型在古诗词领域的表现。
数据集说明
数据集(poetry_benchmark.jsonl)包含70个测试样本,涵盖以下维度:
题型分布:
对联补全
诗句填空
诗词识别
提示词补全
首尾互补
难度等级:
简单(easy)
中等(medium)
困难(hard)
朝代覆盖:
先秦至近现代
包括唐、宋、元、明、清等重要朝代
评测维度
评测框架从以下维度对模型进行全面评估:
整体准确率
不同题型的表现
不同难度等级的表现
不同朝代诗词的掌握程度
评测结果
模型
blank_filling
couplet
find_poetry… See the full description on the dataset page: https://huggingface.co/datasets/happyme531/classical-chinese-poetry-benchmark-70.insurance-classifier-sft
Insurance Coverage Classifier (Stark Law DHS)
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
CPT/HCPCS codes → Stark Law DHS classification + compliance notes
Why download this
Compliance automation for physician self-referral rules. Identify which services are Designated Health Services under Stark Law Section… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/insurance-classifier-sft.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.it-support-l1-ticket-classification
IT Support L1 Multilingual Dataset
Dataset Summary
IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping.
This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.stack-v2-sparse-classes-75kplus
Stack v2 Sparse Python Classes 75kplus
This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments.
Splits
train.jsonl: 74829
val.jsonl: 500
test.jsonl: 500
all.jsonl: 75829
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.stack-v2-sparse-classes-36k
Stack v2 Sparse Python Classes 36k
This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 35,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.classical-chinese-punctuation
Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset)
This is a commercial dataset. A free 200-record sample is provided below
(sample.jsonl); the full 5.3M-pair corpus is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — built from public-domain classical works.
The task
Restore punctuation and sentence segmentation (句读) to unpunctuated Classical
Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.classical-chinese-variant-collation
Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset)
This is a commercial dataset. A free 50-work preview sample is provided
below (sample.jsonl, texts truncated); the full set with complete aligned
texts is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — both members are public-domain classical works.
What this is
A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.
