datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClassiCC-PT
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.classical-greek
Classical Greek Corpus
Ancient and classical Greek (grc) text segments drawn from the open scholarly
corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the
classical/secular comparand within the NuBerea corpus estate, alongside its biblical,
Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon
(Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic
and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.sft_classicBurmese-Classics-OCR-RAW
Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project)
Overview
A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio.
This project fills a critical gap in Burmese-language resources:
Scarcity of public-domain Burmese text.
High technical and financial barriers to corpus building.
Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.greathangpt-classical-chinese
GreatHanGPT 古汉语数据集
数据集描述
这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。
数据来源
来源
内容
链接
chinese-poetry
唐诗宋词、楚辞、诗经、四书五经
GitHub
Werneror/Poetry
先秦到清末诗词,按朝代分
GitHub
CBETA
大正藏佛经
GitHub
数据规模
指标
数值
总记录数
2,400,939
总字符数
450,496,972
估计token数
~300M
时代分布
时代
记录数
字符数
占比
先秦
1,376
14,493,846
3.2%
汉魏
16,450
14,348,817
3.2%
隋唐
729,162
122,265,102
27.1%
两宋
728,569
97,907,260
21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.latin-classical-intertextuality-labels
Latin Jerome Intertextuality Labels
This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.classical-tcm-canon
Classical Chinese Medicine Canon — 中医经典文本数据集 (v1)
A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text
digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon),
难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics —
assembled from public-domain source works. Useful for LLM training, RAG, and search
over classical Chinese medicine / 中医药 literature.
Summary
115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.assistments-classic-recommender
assistments
Processed dataset for the LLM as Recommender project.
classical-grasp
Classical grasp (n=2500 envelope + 1 kHz pulse)
Friction cone on a 4×4 pad. Local slip if |τ| > μ N. Micro: outer ring, inner stuck (or shear within 10% of the cone). Macro: inner slip or |v_slip| > 0.005 m/s. Reflex ramps F ← F + scale·dt, clamp 45 N. Law: evaluate_grasp_dynamics in src/physics/dexterous.rs. Reflex twin: ztp_dexterous_evaluate_grasp.
Clock
Envelope rows are summaries of 1 kHz loops. Pulse is 100 steps at dt = 0.001 s.
Envelope rates… See the full description on the dataset page: https://huggingface.co/datasets/spiderpilot89/classical-grasp.classic-eda-c-trajectories
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.test-subset-classic-eda
nano-rl trajectories
Agent trajectories from nano-rl: an LLM is asked to specify, test and then
implement a small C program, which is compiled and executed inside a confined
sandbox and scored against tests the model wrote before it saw its own program.
When the program fails, the model is shown the build log and the failing cases
and asked to repair it, for up to 100 rounds.
Every model turn is one row, including the ones that went nowhere.
What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.ClassicNovelsmedia-metadata-classical-composers
TigreGotico/media-metadata-classical-composers
Rich entity dataset scraped by metadatarr
scraper classical_composers.
Rows: 15,587
Fields
composer_id
name
country
life
birth
death
period
image_url
url
bio
radio_id
notable
must_know
n_recordings
n_performers
n_albums
n_works_listed
n_albums_listed
Source
Generated by scrapers/classical_composers.py. See the metadatarr repo for the full
pipeline and scraper source code.
openart-portraits-classical
OpenArt — Portraits & the Classical Figure
openart-portraits-classical is the portraits classical subject collection of the OpenArt family
of open, public-domain art datasets: 28,011 works (13,868 paintings/illustrations · 13,970
photographed objects · 173 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
The human figure and portraiture across the full range of media — painted and drawn portraits alongside photographic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-portraits-classical.Classic
📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese
📖 Overview
ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering.
This corpus was created as part of a study on continued pretraining for adapting… See the full description on the dataset page: https://huggingface.co/datasets/Vwegba/Classic.qwen3-classic-rl-distlclassical-swedish-allusions-gold
Classical-Swedish Allusions: Gold Benchmark
A small curated reference set of cross-lingual allusions from classical Greek and Latin into Swedish literature. Used as a held-out evaluation set for the Ericu950/classical-swedish-citations-v2 project.
Purpose
This is not training data. It is a fixed, public reference standard. Our cross-lingual allusion-detection system will aspire to recover these. The set was archived publicly before evaluating any system against it.
The… See the full description on the dataset page: https://huggingface.co/datasets/Ericu950/classical-swedish-allusions-gold.sft_classic_numinaenglish-classics-parallel-samples
Booklern English classics: parallel samples
Paragraph-aligned opening passages of public-domain English classics with a
translation into Spanish, Japanese, Brazilian Portuguese, Russian, Chinese, published by Booklern, a
bilingual book reader for learning English through real books. Each book is
read on Booklern with a sentence-by-sentence translation under the English,
read-aloud audio, a dictionary and vocabulary tools; the rows here are the
same opening paragraphs that appear… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.africa-rwanda-eicv7-classic-public-work-11112876
EICV7: Classic public work | Africa (Rwanda Data Sharing Platform - NISR)
15,054 rows - 1 Africa country/area - 2023-10-16-2024-10-15 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 15,054 rows from Rwanda Data Sharing Platform - NISR, covering EICV7: Classic public work. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-eicv7-classic-public-work-11112876.classical-chinese-variant-collation
Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset)
This is a commercial dataset. A free 50-work preview sample is provided
below (sample.jsonl, texts truncated); the full set with complete aligned
texts is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — both members are public-domain classical works.
What this is
A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.french-classic-books-v2africa-rwanda-eicv7-vup-classic-public-work-c660cd2c
EICV7 (VUP): Classic public work | Africa (Rwanda Data Sharing Platform - NISR)
925 rows - 1 Africa country/area - 2023-10-16-2024-10-15 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 925 rows from Rwanda Data Sharing Platform - NISR, covering EICV7 (VUP): Classic public work. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-eicv7-vup-classic-public-work-c660cd2c.classic_benchmark-v1
CLASSic Benchmark (v1)
Version v1 of the CLASSic Benchmark. Uploaded on June 02, 2025.
Please see Github for more information and model leaderboards.
Note: This is a filtered subset of the dataset originally published in the CLASSIC Benchmark ICLR 2025 Workshop paper which was cleared for public release.
Usage
from datasets import load_dataset
# Load dataset subsets
ds_messages = load_dataset('Miking98/classic_benchmark-v1', 'messages')
ds_workflows =… See the full description on the dataset page: https://huggingface.co/datasets/Miking98/classic_benchmark-v1.music_queries_classical
🎵 MusicQueries Dataset - ClassicalComposers
MusicQueries is a synthetic natural language dataset focused on music-related utterances designed for media playback scenarios. Every sentence in this dataset is a playback request — a query that should result in a media search followed by music playback.
This dataset is ideal for training and evaluating models in intent classification, named entity recognition (NER), and retrieval-based music assistants.
This repository contains the… See the full description on the dataset page: https://huggingface.co/datasets/Jarbas/music_queries_classical.midi-classical-music-toio-json-audit
MIDI Classical Music toio JSON — aggregate audit
This metadata-only audit describes ayousanz/midi-classical-music-toio-json at
revision d07a0210bb7cff7757b9d941b131df10a752eb8c. It contains no MIDI files,
converted song JSON, filenames, or recovered source payloads.
Of 4,796 source MIDI files, 4,712 have conversions. All 84 missing conversions were
invalid under strict SMF parsing; 28 could nevertheless be recovered as RIFF/RMID or
MacBinary containers. Across the 4,712 valid… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json-audit.
