datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.curio-rewrite-non-edu-dataset
Curio Rewrite — Non-Educational
Portuguese web texts (non-educational subset, sampled from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Companion to the educational subset; used to train the Curio rewrite models.
Config
Prompt style
Rows
easy
Simple vocabulary, child-friendly paraphrase
22,237,886
medium
Moderate paraphrase
14,698,285
hard
Sophisticated paraphrase
18,576,570
qa
Reformatted as question/answer
18,664,285
Fields… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-non-edu-dataset.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/ygonet/midi-classical-music.midi-classical-music-toio-json
MIDI Classical Music
drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット
data format
[
{
"track_name": "ALBENIZ: Aragon Op 47/6",
"priority": 1,
"notes": [
{
"note_number": 77,
"start_time_ms": 0,
"duration_units": 26
},
{
},
},
{
"track_name": "apurdam@pcug.org.au",
"priority": 2,
"notes": [
{
"note_number": 53,
"start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.ClassiCC-PT
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.ClassicalPoetryRetrieval
Classical Poetry Retrieval
BEIR-style multi-aspect classical Chinese poetry retrieval for PoetryMTEB / MTEB.
Chinese queries retrieve classical poems along four aspects
(emotion / intent / theme / thought), with graded relevance (score ∈ {0,1,2,3}).
Item
Description
Dataset version
1.2.0
Hub repo
PoetryMTEB/ClassicalPoetryRetrieval
Task
Retrieval (graded, score ∈ {0,1,2,3})
Language
Classical Chinese / Chinese (zh)
Aspects
emotion · intent · theme · thought… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/ClassicalPoetryRetrieval.classical-greek
Classical Greek Corpus
Ancient and classical Greek (grc) text segments drawn from the open scholarly
corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the
classical/secular comparand within the NuBerea corpus estate, alongside its biblical,
Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon
(Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic
and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.chinese-classical-corpus
Chinese Classical Corpus
🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件
🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样)
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
为什么做这个
中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。
Tokenizer 层面 —— 真成立
7 个主流 tokenizer 横评(tokenizer_study):
DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/elizawhitfield/midi-classical-music.ru-classic
Russian Classical Literature — Corpus for Language Models (English and Russian)
English
A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age."
Contents
866 MB of clean text.
61 authors:
Classical Prose of the 19th Century (37 authors)
Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.sft_classiccurio-rewrite-edu-dataset
Curio Rewrite — Educational
Portuguese web texts (educational subset, filtered from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Used to train the Curio rewrite models.
Each config holds the same source documents with a different rewrite style:
Config
Prompt style
Rows
easy
Simple vocabulary, child-friendly paraphrase
7,777,128
medium
Moderate paraphrase
7,777,128
hard
Sophisticated paraphrase
7,777,128
qa
Reformatted as question/answer
7,777… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-edu-dataset.Rasaif-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period.
Content Details
Contained within this dataset are English translations of the following texts, sourced from the Rasaif website:
A Muslim Manual of War
Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.Burmese-Classics-OCR-RAW
Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project)
Overview
A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio.
This project fills a critical gap in Burmese-language resources:
Scarcity of public-domain Burmese text.
High technical and financial barriers to corpus building.
Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format.
Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese
Detailed Dataset Description
This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.philosophy-classics-structured
Classical Decision Frameworks — Philosophy Dataset
Structured public domain philosophical texts focused on decision-making, leadership,
and organizational ethics. All content is in the public domain.
Content
Works from classical philosophy structured for AI analysis:
Stoic decision principles (Marcus Aurelius, Epictetus, Seneca)
Political philosophy (Machiavelli, Aristotle)
Virtue ethics (Aristotle, Plato)
Sources
All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.chinese-classical-corpus
Chinese Classical Corpus
🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite.
中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。
全部 CC0 公有领域,可商用、可改用、无附加限制。
Quick Start
from datasets import load_dataset
# 源语料 (12,005 条章节级记录, 17.2M 字)
corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train")
# 古译今 / 今译古 双向指令数据 (1,924,378 条)
translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.greathangpt-classical-chinese
GreatHanGPT 古汉语数据集
数据集描述
这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。
数据来源
来源
内容
链接
chinese-poetry
唐诗宋词、楚辞、诗经、四书五经
GitHub
Werneror/Poetry
先秦到清末诗词,按朝代分
GitHub
CBETA
大正藏佛经
GitHub
数据规模
指标
数值
总记录数
2,400,939
总字符数
450,496,972
估计token数
~300M
时代分布
时代
记录数
字符数
占比
先秦
1,376
14,493,846
3.2%
汉魏
16,450
14,348,817
3.2%
隋唐
729,162
122,265,102
27.1%
两宋
728,569
97,907,260
21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.anomalies_classiclatin-classical-intertextuality-corpus
Latin Classical Authors Corpus
This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature.
Related Datasets
This corpus is part of the Latin Jerome Intertextuality collection:
Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.Chinese_modern_classical
Dataset Card for "Chinese_modern_classical"
数据来自于NiuTrans/Classical-Modern: 非常全的文言文(古文)-现代文平行语料 (github.com)。
由于原始数据中部分古文没有译文,所以本数据集的数据仅包括了双语数据 。
latin-classical-intertextuality-labels
Latin Jerome Intertextuality Labels
This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.latin-classical-intertextuality-queries
Latin Classical Intertextuality Queries
This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature.
Related Datasets
This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.classical-tcm-canon
Classical Chinese Medicine Canon — 中医经典文本数据集 (v1)
A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text
digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon),
难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics —
assembled from public-domain source works. Useful for LLM training, RAG, and search
over classical Chinese medicine / 中医药 literature.
Summary
115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.bookcorpus-gutenberg-classics
BookCorpus + Gutenberg Classics Training Corpus
Large-scale training corpus combining BookCorpus fiction, Project Gutenberg 19th-century literature (PG-19), and curated classical philosophy texts. Cleaned, deduplicated, and organized into curriculum phases for character-level language model training.
Dataset Description
This corpus is the primary training dataset for the Julia SLM project, combining three major text sources into a unified, cleaned training set with… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/bookcorpus-gutenberg-classics.maestro-classicalfrench-classic-conversationsParsed from https://huggingface.co/datasets/Volko76/french-classic-books
classical-humanities-corpusChinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。
一篇样例如下:
古训《增广贤文》
昔时贤文,诲汝谆谆,集韵增文,多见多闻。
观今宜鉴古,无古不成今。
知己知彼,将心比心。
酒逢知己饮,诗向会人吟。
相识满天下,知心能几人。
相逢好似初相识,到老终无怨恨心。
近水知鱼性,近山识鸟音。
易涨易退山溪水,易反易覆小人心。
运去金成铁,时来铁似金,读书须用意,一字值千金。
classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.
