CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClassiCC-Corpus /ClassiCC-PT 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting English-trained… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/ClassiCC-PT.tabular10M<n<100M15 likes578 downloads8mo agoHugging Face02NuBerea /classical-greekgated Classical Greek Corpus Ancient and classical Greek (grc) text segments drawn from the open scholarly corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the classical/secular comparand within the NuBerea corpus estate, alongside its biblical, Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon (Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.tabulartext-generation10M<n<100M0 likes440 downloads11d agoHugging Face03formalmathatepfl /sft_classictabular1M<n<10M0 likes291 downloads1mo agoHugging Face04minthanthtoo-cs /Burmese-Classics-OCR-RAW Burmese (Myanmar) Books Dataset – Burmese Classics OCR (Daily Rolling Project) Overview A daily rolling dataset of Burmese books, built for AI, OCR, and NLP research.Each entry includes OCR text with metadata: title, author, page index, and Burmese character ratio. This project fills a critical gap in Burmese-language resources: Scarcity of public-domain Burmese text. High technical and financial barriers to corpus building. Enables incremental, open access for… See the full description on the dataset page: https://huggingface.co/datasets/minthanthtoo-cs/Burmese-Classics-OCR-RAW.tabular1M<n<10M1 likes179 downloads1y agoHugging Face05kenpusney /greathangpt-classical-chinese GreatHanGPT 古汉语数据集 数据集描述 这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。 数据来源 来源 内容 链接 chinese-poetry 唐诗宋词、楚辞、诗经、四书五经 GitHub Werneror/Poetry 先秦到清末诗词,按朝代分 GitHub CBETA 大正藏佛经 GitHub 数据规模 指标 数值 总记录数 2,400,939 总字符数 450,496,972 估计token数 ~300M 时代分布 时代 记录数 字符数 占比 先秦 1,376 14,493,846 3.2% 汉魏 16,450 14,348,817 3.2% 隋唐 729,162 122,265,102 27.1% 两宋 728,569 97,907,260 21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.tabulartext-generation1M<n<10M0 likes159 downloads3mo agoHugging Face06julian-schelb /latin-classical-intertextuality-labels Latin Jerome Intertextuality Labels This dataset contains known intertextual relationships between Jerome (Hieronymus) texts and classical Latin authors. These are manually verified or scholarly-identified cases of intertextuality, serving as ground truth labels for training and evaluating intertextuality detection systems. Each link carries item-level provenance identifying which prior publication (if any) it was adopted from -- see provenance_dataset under Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/julian-schelb/latin-classical-intertextuality-labels.tabulartext-retrieval1K<n<10K2 likes113 downloads1mo agoHugging Face07wangekxy /classical-tcm-canon Classical Chinese Medicine Canon — 中医经典文本数据集 (v1) A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon), 难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics — assembled from public-domain source works. Useful for LLM training, RAG, and search over classical Chinese medicine / 中医药 literature. Summary 115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.tabulartext-generationn<1K0 likes113 downloads3mo agoHugging Face08systemslibrarian /classical-cipher-corpus Classical Cipher Corpus A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers. Part of the Cipher Detective AI project: 🕵️ Space: systemslibrarian/cipher-detective-ai 📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo) 🤖 Model: systemslibrarian/cipher-detective-classifier Intended use Teach classical cryptanalysis. Benchmark educational cipher-family detectors. Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.tabulartext-classification10K<n<100K0 likes106 downloads5mo agoHugging Face09OloriBern /assistments-classic-recommender assistments Processed dataset for the LLM as Recommender project. tabular100K<n<1M0 likes94 downloads1y agoHugging Face10spiderpilot89 /classical-grasp Classical grasp (n=2500 envelope + 1 kHz pulse) Friction cone on a 4×4 pad. Local slip if |τ| > μ N. Micro: outer ring, inner stuck (or shear within 10% of the cone). Macro: inner slip or |v_slip| > 0.005 m/s. Reflex ramps F ← F + scale·dt, clamp 45 N. Law: evaluate_grasp_dynamics in src/physics/dexterous.rs. Reflex twin: ztp_dexterous_evaluate_grasp. Clock Envelope rows are summaries of 1 kHz loops. Pulse is 100 steps at dt = 0.001 s. Envelope rates… See the full description on the dataset page: https://huggingface.co/datasets/spiderpilot89/classical-grasp.tabular1K<n<10K0 likes71 downloads16d agoHugging Face11Sudnya /classic-eda-c-trajectoriesgated nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 100 rounds. Every model turn is one row, including the ones that went nowhere. What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.tabulartext-generation100K<n<1M0 likes65 downloads6d agoHugging Face12Sudnya /test-subset-classic-eda nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 100 rounds. Every model turn is one row, including the ones that went nowhere. What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.tabulartext-generation1K<n<10K0 likes65 downloads2d agoHugging Face13hugfaceguy0001 /ClassicNovelstabular10K<n<100K0 likes64 downloads2y agoHugging Face14LeData /media-metadata-classical-composers TigreGotico/media-metadata-classical-composers Rich entity dataset scraped by metadatarr scraper classical_composers. Rows: 15,587 Fields composer_id name country life birth death period image_url url bio radio_id notable must_know n_recordings n_performers n_albums n_works_listed n_albums_listed Source Generated by scrapers/classical_composers.py. See the metadatarr repo for the full pipeline and scraper source code. tabular10K<n<100K0 likes61 downloads3mo agoHugging Face15jaddai /openart-portraits-classical OpenArt — Portraits & the Classical Figure openart-portraits-classical is the portraits classical subject collection of the OpenArt family of open, public-domain art datasets: 28,011 works (13,868 paintings/illustrations · 13,970 photographed objects · 173 unclassified), each paired with a structured VLM caption plus medium, attribution and inscription metadata. The human figure and portraiture across the full range of media — painted and drawn portraits alongside photographic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-portraits-classical.tabularimage-to-text10K<n<100K1 likes51 downloads4mo agoHugging Face16Vwegba /Classic 📚 ClassiCC-PT: Classified Common Crawl Corpus for Portuguese 📖 Overview ClassiCC-PT (Classified Common Crawl – Portuguese) is a large-scale web corpus containing ~120B Portuguese tokens extracted from Common Crawl snapshots. It is specifically curated for training large language models in Portuguese, with a focus on data quality, language specificity, and targeted filtering. This corpus was created as part of a study on continued pretraining for adapting… See the full description on the dataset page: https://huggingface.co/datasets/Vwegba/Classic.tabular10M<n<100M0 likes50 downloads3mo agoHugging Face17formalmathatepfl /qwen3-classic-rl-distltabular10K<n<100K0 likes47 downloads13d agoHugging Face18Ericu950 /classical-swedish-allusions-gold Classical-Swedish Allusions: Gold Benchmark A small curated reference set of cross-lingual allusions from classical Greek and Latin into Swedish literature. Used as a held-out evaluation set for the Ericu950/classical-swedish-citations-v2 project. Purpose This is not training data. It is a fixed, public reference standard. Our cross-lingual allusion-detection system will aspire to recover these. The set was archived publicly before evaluating any system against it. The… See the full description on the dataset page: https://huggingface.co/datasets/Ericu950/classical-swedish-allusions-gold.tabularsentence-similarityn<1K0 likes41 downloads4mo agoHugging Face19formalmathatepfl /sft_classic_numinatabular10K<n<100K0 likes41 downloads1mo agoHugging Face20ssergaroo /english-classics-parallel-samples Booklern English classics: parallel samples Paragraph-aligned opening passages of public-domain English classics with a translation into Spanish, Japanese, Brazilian Portuguese, Russian, Chinese, published by Booklern, a bilingual book reader for learning English through real books. Each book is read on Booklern with a sentence-by-sentence translation under the English, read-aloud audio, a dictionary and vocabulary tools; the rows here are the same opening paragraphs that appear… See the full description on the dataset page: https://huggingface.co/datasets/ssergaroo/english-classics-parallel-samples.tabulartranslation1K<n<10K0 likes24 downloads2d agoHugging Face21electricsheepafrica /africa-rwanda-eicv7-classic-public-work-11112876 EICV7: Classic public work | Africa (Rwanda Data Sharing Platform - NISR) 15,054 rows - 1 Africa country/area - 2023-10-16-2024-10-15 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 15,054 rows from Rwanda Data Sharing Platform - NISR, covering EICV7: Classic public work. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-eicv7-classic-public-work-11112876.tabulartabular-classification10K<n<100K0 likes22 downloads1mo agoHugging Face22wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes17 downloads3mo agoHugging Face23Volko76 /french-classic-books-v2tabularn<1K0 likes12 downloads10mo agoHugging Face24electricsheepafrica /africa-rwanda-eicv7-vup-classic-public-work-c660cd2c EICV7 (VUP): Classic public work | Africa (Rwanda Data Sharing Platform - NISR) 925 rows - 1 Africa country/area - 2023-10-16-2024-10-15 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 925 rows from Rwanda Data Sharing Platform - NISR, covering EICV7 (VUP): Classic public work. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-rwanda-eicv7-vup-classic-public-work-c660cd2c.tabulartabular-classificationn<1K0 likes9 downloads1mo agoHugging Face25Miking98 /classic_benchmark-v1gated CLASSic Benchmark (v1) Version v1 of the CLASSic Benchmark. Uploaded on June 02, 2025. Please see Github for more information and model leaderboards. Note: This is a filtered subset of the dataset originally published in the CLASSIC Benchmark ICLR 2025 Workshop paper which was cleared for public release. Usage from datasets import load_dataset # Load dataset subsets ds_messages = load_dataset('Miking98/classic_benchmark-v1', 'messages') ds_workflows =… See the full description on the dataset page: https://huggingface.co/datasets/Miking98/classic_benchmark-v1.tabulartext-classification1K<n<10K1 likes8 downloads1y agoHugging Face26Jarbas /music_queries_classical 🎵 MusicQueries Dataset - ClassicalComposers MusicQueries is a synthetic natural language dataset focused on music-related utterances designed for media playback scenarios. Every sentence in this dataset is a playback request — a query that should result in a media search followed by music playback. This dataset is ideal for training and evaluating models in intent classification, named entity recognition (NER), and retrieval-based music assistants. This repository contains the… See the full description on the dataset page: https://huggingface.co/datasets/Jarbas/music_queries_classical.tabulartext-classification10K<n<100K0 likes7 downloads1y agoHugging Face27ayousanz /midi-classical-music-toio-json-auditgated MIDI Classical Music toio JSON — aggregate audit This metadata-only audit describes ayousanz/midi-classical-music-toio-json at revision d07a0210bb7cff7757b9d941b131df10a752eb8c. It contains no MIDI files, converted song JSON, filenames, or recovered source payloads. Of 4,796 source MIDI files, 4,712 have conversions. All 84 missing conversions were invalid under strict SMF parsing; 28 could nevertheless be recovered as RIFF/RMID or MacBinary containers. Across the 4,712 valid… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json-audit.tabularn<1K0 likes5 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.