CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01drengskapur /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.text1K<n<10K19 likes9.6k downloads2y agoHugging Face02ClassiCC-Corpus /curio-rewrite-non-edu-dataset Curio Rewrite — Non-Educational Portuguese web texts (non-educational subset, sampled from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Companion to the educational subset; used to train the Curio rewrite models. Config Prompt style Rows easy Simple vocabulary, child-friendly paraphrase 22,237,886 medium Moderate paraphrase 14,698,285 hard Sophisticated paraphrase 18,576,570 qa Reformatted as question/answer 18,664,285 Fields… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-non-edu-dataset.texttext-generation10M<n<100M0 likes1.4k downloads5mo agoHugging Face03ygonet /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/ygonet/midi-classical-music.text1K<n<10K0 likes494 downloads2mo agoHugging Face04ayousanz /midi-classical-music-toio-json MIDI Classical Music drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット data format [ { "track_name": "ALBENIZ: Aragon Op 47/6", "priority": 1, "notes": [ { "note_number": 77, "start_time_ms": 0, "duration_units": 26 }, { }, }, { "track_name": "apurdam@pcug.org.au", "priority": 2, "notes": [ { "note_number": 53, "start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.text10K<n<100K2 likes492 downloads2y agoHugging Face05ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes476 downloads8mo agoHugging Face06PoetryMTEB /ClassicalPoetryRetrieval Classical Poetry Retrieval BEIR-style multi-aspect classical Chinese poetry retrieval for PoetryMTEB / MTEB. Chinese queries retrieve classical poems along four aspects (emotion / intent / theme / thought), with graded relevance (score ∈ {0,1,2,3}). Item Description Dataset version 1.2.0 Hub repo PoetryMTEB/ClassicalPoetryRetrieval Task Retrieval (graded, score ∈ {0,1,2,3}) Language Classical Chinese / Chinese (zh) Aspects emotion · intent · theme · thought… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/ClassicalPoetryRetrieval.texttext-retrieval1M<n<10M0 likes461 downloads16d agoHugging Face07NuBerea /classical-greekgated Classical Greek Corpus Ancient and classical Greek (grc) text segments drawn from the open scholarly corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the classical/secular comparand within the NuBerea corpus estate, alongside its biblical, Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon (Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.tabulartext-generation10M<n<100M0 likes441 downloads11d agoHugging Face08gujilab /chinese-classical-corpus Chinese Classical Corpus 🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件 🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样) 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面 —— 真成立 7 个主流 tokenizer 横评(tokenizer_study): DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.texttext-generation1M<n<10M1 likes412 downloads4mo agoHugging Face09elizawhitfield /midi-classical-music MIDI Classical Music This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers. The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others. The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions. The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/elizawhitfield/midi-classical-music.text1K<n<10K4 likes357 downloads9mo agoHugging Face10Imperius /ru-classic Russian Classical Literature — Corpus for Language Models (English and Russian) English A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age." Contents 866 MB of clean text. 61 authors: Classical Prose of the 19th Century (37 authors) Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.texttext-generation1M<n<10M1 likes295 downloads5mo agoHugging Face11formalmathatepfl /sft_classictabular1M<n<10M0 likes287 downloads1mo agoHugging Face12ClassiCC-Corpus /curio-rewrite-edu-dataset Curio Rewrite — Educational Portuguese web texts (educational subset, filtered from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Used to train the Curio rewrite models. Each config holds the same source documents with a different rewrite style: Config Prompt style Rows easy Simple vocabulary, child-friendly paraphrase 7,777,128 medium Moderate paraphrase 7,777,128 hard Sophisticated paraphrase 7,777,128 qa Reformatted as question/answer 7,777… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-edu-dataset.texttext-generation10M<n<100M0 likes249 downloads5mo agoHugging Face13ImruQays /Rasaif-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period. Content Details Contained within this dataset are English translations of the following texts, sourced from the Rasaif website: A Muslim Manual of War Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes207 downloads3y agoHugging Face14minthanthtoo-cs /Burmese-Classics-OCR-RAW Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project) Overview A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio. This project fills a critical gap in Burmese-language resources: Scarcity of public-domain Burmese text. High technical and financial barriers to corpus building. Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.tabular1M<n<10M1 likes172 downloads1y agoHugging Face15bobboyms /portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format. Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese Detailed Dataset Description This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.texttext-generation10K<n<100K0 likes170 downloads1y agoHugging Face16gmahia /philosophy-classics-structured Classical Decision Frameworks — Philosophy Dataset Structured public domain philosophical texts focused on decision-making, leadership, and organizational ethics. All content is in the public domain. Content Works from classical philosophy structured for AI analysis: Stoic decision principles (Marcus Aurelius, Epictetus, Seneca) Political philosophy (Machiavelli, Aristotle) Virtue ethics (Aristotle, Plato) Sources All works published before 1928… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/philosophy-classics-structured.texttext-classificationn<1K0 likes165 downloads2mo agoHugging Face17flowerone /chinese-classical-corpus Chinese Classical Corpus 🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite. 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 Quick Start from datasets import load_dataset # 源语料 (12,005 条章节级记录, 17.2M 字) corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train") # 古译今 / 今译古 双向指令数据 (1,924,378 条) translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.texttext-generation1M<n<10M0 likes157 downloads5mo agoHugging Face18kenpusney /greathangpt-classical-chinese GreatHanGPT 古汉语数据集 数据集描述 这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。 数据来源 来源 内容 链接 chinese-poetry 唐诗宋词、楚辞、诗经、四书五经 GitHub Werneror/Poetry 先秦到清末诗词,按朝代分 GitHub CBETA 大正藏佛经 GitHub 数据规模 指标 数值 总记录数 2,400,939 总字符数 450,496,972 估计token数 ~300M 时代分布 时代 记录数 字符数 占比 先秦 1,376 14,493,846 3.2% 汉魏 16,450 14,348,817 3.2% 隋唐 729,162 122,265,102 27.1% 两宋 728,569 97,907,260 21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.tabulartext-generation1M<n<10M0 likes157 downloads3mo agoHugging Face19acroitoru /anomalies_classicimage10K<n<100K0 likes143 downloads8mo agoHugging Face20julian-schelb /latin-classical-intertextuality-corpus Latin Classical Authors Corpus This dataset contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature. Related Datasets This corpus is part of the Latin Jerome Intertextuality collection: Queries: latin-classical-intertextuality-queries - The whole works of… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-corpus.texttext-retrieval10K<n<100K4 likes137 downloads1mo agoHugging Face21xmj2002 /Chinese_modern_classical Dataset Card for "Chinese_modern_classical" 数据来自于NiuTrans/Classical-Modern: 非常全的文言文(古文)-现代文平行语料 (github.com)。 由于原始数据中部分古文没有译文,所以本数据集的数据仅包括了双语数据 。 texttranslation100K<n<1M48 likes133 downloads3y agoHugging Face22julian-schelb /latin-classical-intertextuality-labels Latin Jerome Intertextuality Labels This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.tabulartext-retrieval1K<n<10K2 likes121 downloads1mo agoHugging Face23julian-schelb /latin-classical-intertextuality-queries Latin Classical Intertextuality Queries This dataset contains query texts used for finding intertextual relationships with classical Latin authors. It comprises the whole works of Hieronymus (Jerome) and Lactantius, which are searched against a corpus of classical Latin literature. It is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Lactantius and classical Latin literature. Related Datasets This queries… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-queries.texttext-retrieval10K<n<100K2 likes113 downloads1mo agoHugging Face24wangekxy /classical-tcm-canon Classical Chinese Medicine Canon — 中医经典文本数据集 (v1) A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon), 难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics — assembled from public-domain source works. Useful for LLM training, RAG, and search over classical Chinese medicine / 中医药 literature. Summary 115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.tabulartext-generationn<1K0 likes113 downloads3mo agoHugging Face25LisaMegaWatts /bookcorpus-gutenberg-classics BookCorpus + Gutenberg Classics Training Corpus Large-scale training corpus combining BookCorpus fiction, Project Gutenberg 19th-century literature (PG-19), and curated classical philosophy texts. Cleaned, deduplicated, and organized into curriculum phases for character-level language model training. Dataset Description This corpus is the primary training dataset for the Julia SLM project, combining three major text sources into a unified, cleaned training set with… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/bookcorpus-gutenberg-classics.texttext-generation10M<n<100M0 likes112 downloads7mo agoHugging Face26xenon111 /maestro-classicalaudion<1K0 likes105 downloads4mo agoHugging Face27Volko76 /french-classic-conversationsParsed from https://huggingface.co/datasets/Volko76/french-classic-books text10K<n<100K0 likes103 downloads10mo agoHugging Face28LisaMegaWatts /classical-humanities-corpustext100K<n<1M0 likes103 downloads7mo agoHugging Face29Mxode /Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。 一篇样例如下: 古训《增广贤文》 昔时贤文,诲汝谆谆,集韵增文,多见多闻。 观今宜鉴古,无古不成今。 知己知彼,将心比心。 酒逢知己饮,诗向会人吟。 相识满天下,知心能几人。 相逢好似初相识,到老终无怨恨心。 近水知鱼性,近山识鸟音。 易涨易退山溪水,易反易覆小人心。 运去金成铁,时来铁似金,读书须用意,一字值千金。 texttext-generation100K<n<1M5 likes102 downloads1y agoHugging Face30systemslibrarian /classical-cipher-corpus Classical Cipher Corpus A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers. Part of the Cipher Detective AI project: 🕵️ Space: systemslibrarian/cipher-detective-ai 📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo) 🤖 Model: systemslibrarian/cipher-detective-classifier Intended use Teach classical cryptanalysis. Benchmark educational cipher-family detectors. Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.tabulartext-classification10K<n<100K0 likes102 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.