datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
curio-rewrite-non-edu-dataset
Curio Rewrite — Non-Educational
Portuguese web texts (non-educational subset, sampled from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Companion to the educational subset; used to train the Curio rewrite models.
Config
Prompt style
Rows
easy
Simple vocabulary, child-friendly paraphrase
22,237,886
medium
Moderate paraphrase
14,698,285
hard
Sophisticated paraphrase
18,576,570
qa
Reformatted as question/answer
18,664,285
Fields… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-non-edu-dataset.buddhist-classics-vol1-121.2.3.5.6.7.8.11.12等卷2025年7月-11月制作的g2.0翻译本中有约1%弱数量,整段脱译
的情况,为这种情况做了专门的程序,进行了电子校勘和补译。
零散校勘(0.n比例)升级,自20251122起不再更替整个版本群的网盘,只上传到
https://huggingface.co/ospx1u 存档。 最新版本有可能不在长期链接表,只
在 https://huggingface.co/ospx1u 请自行查询,一般情况1-12卷的频繁升级全表
在 https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12/tree/main
这意味着1-13卷的具体内容的长期下载链接,已经不是最新和最准确的内容,
只是聊备一格,特别是变更了数据仓库和转入校勘和平衡双译本时期以后,
已经是可有可无的数据陈迹,但实际差别有限的,所以仍然会有基于internxt的更新,
和过去数据的存留,再次重复,最新数据都在 https://huggingface.co/ospx1u
Buddhist… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12.buddhist-classics-vol1-121.2.3.5.6.7.8.11.12等卷2025年7月-11月制作的g2.0翻译本中有约1%弱数量,整段脱译
的情况,为这种情况做了专门的程序,进行了电子校勘和补译。
零散校勘(0.n比例)升级,自20251122起不再更替整个版本群的网盘,只上传到
https://huggingface.co/ospx1u 存档。 最新版本有可能不在长期链接表,只
在 https://huggingface.co/ospx1u 请自行查询,一般情况1-12卷的频繁升级全表
在 https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12/tree/main
这意味着1-13卷的具体内容的长期下载链接,已经不是最新和最准确的内容,
只是聊备一格,特别是变更了数据仓库和转入校勘和平衡双译本时期以后,
已经是可有可无的数据陈迹,但实际差别有限的,所以仍然会有基于internxt的更新,
和过去数据的存留,再次重复,最新数据都在 https://huggingface.co/ospx1u
Buddhist… See the full description on the dataset page: https://huggingface.co/datasets/renjiezhang/buddhist-classics-vol1-12.lichess_classic_20006,643,902 chess games from the Lichess Open Database (https://database.lichess.org/#standard_games) that meet the following criteria:
At least one player with ELO>=2,000
Rated Classical game mode
Normal termination
Result of 0-1 or 1-0 (no ties)
chinese-classical-corpus
Chinese Classical Corpus
🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件
🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样)
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
为什么做这个
中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。
Tokenizer 层面 —— 真成立
7 个主流 tokenizer 横评(tokenizer_study):
DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.classical-greek
Classical Greek Corpus
Ancient and classical Greek (grc) text segments drawn from the open scholarly
corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the
classical/secular comparand within the NuBerea corpus estate, alongside its biblical,
Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon
(Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic
and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.ru-classic
Russian Classical Literature — Corpus for Language Models (English and Russian)
English
A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age."
Contents
866 MB of clean text.
61 authors:
Classical Prose of the 19th Century (37 authors)
Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.buddhist-classics-vol13-english
license: cc-by-4.0
language:
- bo # Tibetan (source, if applicable)
- en # English (translations)
multilinguality:
- translation
task_categories:
- translation
pretty_name: 佛典AI译丛第十三卷:English Translation Collection of Buddhist Classics AI Series Version 1.0
size_categories:
- 1.7GB
tags:
- buddhism
- tibetan-buddhism
- english-translation
- ai-generated
- northern-buddhism
- kangyur
- tengyur
license: cc-by-4.0
English Translation… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol13-english.curio-rewrite-edu-dataset
Curio Rewrite — Educational
Portuguese web texts (educational subset, filtered from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Used to train the Curio rewrite models.
Each config holds the same source documents with a different rewrite style:
Config
Prompt style
Rows
easy
Simple vocabulary, child-friendly paraphrase
7,777,128
medium
Moderate paraphrase
7,777,128
hard
Sophisticated paraphrase
7,777,128
qa
Reformatted as question/answer
7,777… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-edu-dataset.chinese-classical-corpus
Chinese Classical Corpus
🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite.
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
Quick Start
from datasets import load_dataset
# 源语料 (12,005 条章节级记录, 17.2M 字)
corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train")
# 古译今 / 今译古 双向指令数据 (1,924,378 条)
translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format.
Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese
Detailed Dataset Description
This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.greathangpt-classical-chinese
GreatHanGPT 古汉语数据集
数据集描述
这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。
数据来源
来源
内容
链接
chinese-poetry
唐诗宋词、楚辞、诗经、四书五经
GitHub
Werneror/Poetry
先秦到清末诗词,按朝代分
GitHub
CBETA
大正藏佛经
GitHub
数据规模
指标
数值
总记录数
2,400,939
总字符数
450,496,972
估计token数
~300M
时代分布
时代
记录数
字符数
占比
先秦
1,376
14,493,846
3.2%
汉魏
16,450
14,348,817
3.2%
隋唐
729,162
122,265,102
27.1%
两宋
728,569
97,907,260
21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.sanskrit_classicThis dataset combines some of the classical Sanskrit texts.bookcorpus-gutenberg-classics
BookCorpus + Gutenberg Classics Training Corpus
Large-scale training corpus combining BookCorpus fiction, Project Gutenberg 19th-century literature (PG-19), and curated classical philosophy texts. Cleaned, deduplicated, and organized into curriculum phases for character-level language model training.
Dataset Description
This corpus is the primary training dataset for the Julia SLM project, combining three major text sources into a unified, cleaned training set with… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/bookcorpus-gutenberg-classics.classical-tcm-canon
Classical Chinese Medicine Canon — 中医经典文本数据集 (v1)
A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text
digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon),
难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics —
assembled from public-domain source works. Useful for LLM training, RAG, and search
over classical Chinese medicine / 中医药 literature.
Summary
115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。
一篇样例如下:
古训《增广贤文》
昔时贤文,诲汝谆谆,集韵增文,多见多闻。
观今宜鉴古,无古不成今。
知己知彼,将心比心。
酒逢知己饮,诗向会人吟。
相识满天下,知心能几人。
相逢好似初相识,到老终无怨恨心。
近水知鱼性,近山识鸟音。
易涨易退山溪水,易反易覆小人心。
运去金成铁,时来铁似金,读书须用意,一字值千金。
classic-eda-c-trajectories
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.test-subset-classic-eda
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.classic-eda
Classic EDA - Period Software Task Specifications
1020 specifications for software that plausibly could have been written between
1985 and 1996, mined from two in-era archives and shaped as coding tasks with
machine-checkable requirements.
Each record names a program, describes it in a paragraph, and states 3-12 atomic
requirements plus an explicit interface contract (argv, stdin, stdout, exit
codes) so a grader can test an implementation.
Why period software… See the full description on the dataset page: https://huggingface.co/datasets/gdiamos/classic-eda.Appreciation-of-Chinese-Classical-Poetry
Appreciation of Chinese Classical Poetry
Chinese classical poetry with paired metadata and five-aspect literary analyses for PoetryMTEB / MTEB-style evaluation and computational poetics research.
Poems are drawn from expert appreciation volumes (mainly Shanghai Lexicographical Publishing House dictionaries). The released analysis fields are LLM distillations (DeepSeek-V3.1) of those expert appreciation texts into five free-text facets. The original long-form appreciation prose… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/Appreciation-of-Chinese-Classical-Poetry.chinese-classical-bench
Chinese Classical Bench
中国古典语言能力评测基准 — 6 个任务 × 100 题 = 600 道,覆盖翻译、断句、字义、典故、续写填空、现代→文言压缩。
📊 在线排行榜: 🤗 Space — chinese-classical-bench-leaderboard
🔗 评测代码 & runner: github.com/gujilab/chinese-classical-bench — eval runner(OpenAI 兼容端点)、打分器、排行榜聚合脚本
📦 配套语料集: gujilab/chinese-classical-corpus (CC0 公有领域) — 题目均从该语料抽样生成
为什么做这个
中文(尤其文言文)常被说成"高密度优势"。这套基础设施(bench + corpus + 4 个论点实证实验)想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。
Tokenizer 层面(实证)
7 个主流 tokenizer 横评(详见下方… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-bench.classical-chinese-poetry-benchmark-70
English Readme see below
(README由Claude 3.5 Sonnet生成)
中国古诗词大模型评测基准
简介
这是一个专门用于评测大语言模型在中国古诗词理解和生成方面能力的基准测试集。该基准包含了一个多样化的测试数据集和完整的评测框架,可用于系统性地评估和比较不同模型在古诗词领域的表现。
数据集说明
数据集(poetry_benchmark.jsonl)包含70个测试样本,涵盖以下维度:
题型分布:
对联补全
诗句填空
诗词识别
提示词补全
首尾互补
难度等级:
简单(easy)
中等(medium)
困难(hard)
朝代覆盖:
先秦至近现代
包括唐、宋、元、明、清等重要朝代
评测维度
评测框架从以下维度对模型进行全面评估:
整体准确率
不同题型的表现
不同难度等级的表现
不同朝代诗词的掌握程度
评测结果
模型
blank_filling
couplet
find_poetry… See the full description on the dataset page: https://huggingface.co/datasets/happyme531/classical-chinese-poetry-benchmark-70.Classical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.classical_armenian_pd
Classical Armenian Public Domain Literature
This dataset consists of 102 Classical Armenian texts in the public domain, which were collected from the Eastern Armenian National Corpus.
A list of the works is provided below.
Full list of works
List of Works
Աբովյան Խաչատուր՝ Առաջին սերը (First Love by Khachatur Abovian)
Աբովյան Խաչատուր՝ Պարապ վախտի խաղալիք (Idle Time Toy by Khachatur Abovian)
Աբովյան Խաչատուր՝ Թուրքի աղջիկը (The Turkish Girl by Khachatur Abovian)… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/classical_armenian_pd.classical-chinese-punctuation
Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset)
This is a commercial dataset. A free 200-record sample is provided below
(sample.jsonl); the full 5.3M-pair corpus is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — built from public-domain classical works.
The task
Restore punctuation and sentence segmentation (句读) to unpunctuated Classical
Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.classical-chinese-variant-collation
Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset)
This is a commercial dataset. A free 50-work preview sample is provided
below (sample.jsonl, texts truncated); the full set with complete aligned
texts is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — both members are public-domain classical works.
What this is
A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。
一篇样例如下:
古训《增广贤文》
昔时贤文,诲汝谆谆,集韵增文,多见多闻。
观今宜鉴古,无古不成今。
知己知彼,将心比心。
酒逢知己饮,诗向会人吟。
相识满天下,知心能几人。
相逢好似初相识,到老终无怨恨心。
近水知鱼性,近山识鸟音。
易涨易退山溪水,易反易覆小人心。
运去金成铁,时来铁似金,读书须用意,一字值千金。
