CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ClassiCC-Corpus /curio-rewrite-non-edu-dataset Curio Rewrite — Non-Educational Portuguese web texts (non-educational subset, sampled from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Companion to the educational subset; used to train the Curio rewrite models. Config Prompt style Rows easy Simple vocabulary, child-friendly paraphrase 22,237,886 medium Moderate paraphrase 14,698,285 hard Sophisticated paraphrase 18,576,570 qa Reformatted as question/answer 18,664,285 Fields… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-non-edu-dataset.texttext-generation10M<n<100M0 likes1.4k downloads5mo agoHugging Face02ospx1u /buddhist-classics-vol1-121.2.3.5.6.7.8.11.12等卷2025年7月-11月制作的g2.0翻译本中有约1%弱数量,整段脱译 的情况,为这种情况做了专门的程序,进行了电子校勘和补译。 零散校勘(0.n比例)升级,自20251122起不再更替整个版本群的网盘,只上传到 https://huggingface.co/ospx1u 存档。 最新版本有可能不在长期链接表,只 在 https://huggingface.co/ospx1u 请自行查询,一般情况1-12卷的频繁升级全表 在 https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12/tree/main 这意味着1-13卷的具体内容的长期下载链接,已经不是最新和最准确的内容, 只是聊备一格,特别是变更了数据仓库和转入校勘和平衡双译本时期以后, 已经是可有可无的数据陈迹,但实际差别有限的,所以仍然会有基于internxt的更新, 和过去数据的存留,再次重复,最新数据都在 https://huggingface.co/ospx1u Buddhist… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12.translation1 likes1.1k downloads8mo agoHugging Face03renjiezhang /buddhist-classics-vol1-121.2.3.5.6.7.8.11.12等卷2025年7月-11月制作的g2.0翻译本中有约1%弱数量,整段脱译 的情况,为这种情况做了专门的程序,进行了电子校勘和补译。 零散校勘(0.n比例)升级,自20251122起不再更替整个版本群的网盘,只上传到 https://huggingface.co/ospx1u 存档。 最新版本有可能不在长期链接表,只 在 https://huggingface.co/ospx1u 请自行查询,一般情况1-12卷的频繁升级全表 在 https://huggingface.co/datasets/ospx1u/buddhist-classics-vol1-12/tree/main 这意味着1-13卷的具体内容的长期下载链接,已经不是最新和最准确的内容, 只是聊备一格,特别是变更了数据仓库和转入校勘和平衡双译本时期以后, 已经是可有可无的数据陈迹,但实际差别有限的,所以仍然会有基于internxt的更新, 和过去数据的存留,再次重复,最新数据都在 https://huggingface.co/ospx1u Buddhist… See the full description on the dataset page: https://huggingface.co/datasets/renjiezhang/buddhist-classics-vol1-12.translation0 likes594 downloads9mo agoHugging Face04vllg /lichess_classic_20006,643,902 chess games from the Lichess Open Database (https://database.lichess.org/#standard_games) that meet the following criteria: At least one player with ELO>=2,000 Rated Classical game mode Normal termination Result of 0-1 or 1-0 (no ties) text-generation1M<n<10M0 likes480 downloads3y agoHugging Face05gujilab /chinese-classical-corpus Chinese Classical Corpus 🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件 🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样) 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面 —— 真成立 7 个主流 tokenizer 横评(tokenizer_study): DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.texttext-generation1M<n<10M1 likes454 downloads4mo agoHugging Face06NuBerea /classical-greekgated Classical Greek Corpus Ancient and classical Greek (grc) text segments drawn from the open scholarly corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the classical/secular comparand within the NuBerea corpus estate, alongside its biblical, Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon (Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.tabulartext-generation10M<n<100M0 likes440 downloads11d agoHugging Face07Imperius /ru-classic Russian Classical Literature — Corpus for Language Models (English and Russian) English A clean text corpus of Russian classical literature from the 19th to the early 20th centuries, collected from lib.ru and subjected to several iterations of cleaning. Suitable for pre-training language models on the style of the Russian prose "Golden Age." Contents 866 MB of clean text. 61 authors: Classical Prose of the 19th Century (37 authors) Chekhov, Tolstoy… See the full description on the dataset page: https://huggingface.co/datasets/Imperius/ru-classic.texttext-generation1M<n<10M1 likes299 downloads5mo agoHugging Face08ospx1u /buddhist-classics-vol13-english license: cc-by-4.0 language: - bo # Tibetan (source, if applicable) - en # English (translations) multilinguality: - translation task_categories: - translation pretty_name: 佛典AI译丛第十三卷:English Translation Collection of Buddhist Classics AI Series Version 1.0 size_categories: - 1.7GB tags: - buddhism - tibetan-buddhism - english-translation - ai-generated - northern-buddhism - kangyur - tengyur license: cc-by-4.0 English Translation… See the full description on the dataset page: https://huggingface.co/datasets/ospx1u/buddhist-classics-vol13-english.translation0 likes262 downloads6mo agoHugging Face09ClassiCC-Corpus /curio-rewrite-edu-dataset Curio Rewrite — Educational Portuguese web texts (educational subset, filtered from ClassiCC) rewritten by Qwen2.5-7B-Instruct under four prompt styles. Used to train the Curio rewrite models. Each config holds the same source documents with a different rewrite style: Config Prompt style Rows easy Simple vocabulary, child-friendly paraphrase 7,777,128 medium Moderate paraphrase 7,777,128 hard Sophisticated paraphrase 7,777,128 qa Reformatted as question/answer 7,777… See the full description on the dataset page: https://huggingface.co/datasets/ClassiCC-Corpus/curio-rewrite-edu-dataset.texttext-generation10M<n<100M0 likes247 downloads5mo agoHugging Face10flowerone /chinese-classical-corpus Chinese Classical Corpus 🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite. 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 Quick Start from datasets import load_dataset # 源语料 (12,005 条章节级记录, 17.2M 字) corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train") # 古译今 / 今译古 双向指令数据 (1,924,378 条) translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.texttext-generation1M<n<10M0 likes208 downloads5mo agoHugging Face11bobboyms /portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format. Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese Detailed Dataset Description This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.texttext-generation10K<n<100K0 likes170 downloads1y agoHugging Face12kenpusney /greathangpt-classical-chinese GreatHanGPT 古汉语数据集 数据集描述 这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。 数据来源 来源 内容 链接 chinese-poetry 唐诗宋词、楚辞、诗经、四书五经 GitHub Werneror/Poetry 先秦到清末诗词,按朝代分 GitHub CBETA 大正藏佛经 GitHub 数据规模 指标 数值 总记录数 2,400,939 总字符数 450,496,972 估计token数 ~300M 时代分布 时代 记录数 字符数 占比 先秦 1,376 14,493,846 3.2% 汉魏 16,450 14,348,817 3.2% 隋唐 729,162 122,265,102 27.1% 两宋 728,569 97,907,260 21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.tabulartext-generation1M<n<10M0 likes159 downloads3mo agoHugging Face13surajp /sanskrit_classicThis dataset combines some of the classical Sanskrit texts.text-generation100K<n<1M4 likes153 downloads3y agoHugging Face14LisaMegaWatts /bookcorpus-gutenberg-classics BookCorpus + Gutenberg Classics Training Corpus Large-scale training corpus combining BookCorpus fiction, Project Gutenberg 19th-century literature (PG-19), and curated classical philosophy texts. Cleaned, deduplicated, and organized into curriculum phases for character-level language model training. Dataset Description This corpus is the primary training dataset for the Julia SLM project, combining three major text sources into a unified, cleaned training set with… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/bookcorpus-gutenberg-classics.texttext-generation10M<n<100M0 likes115 downloads7mo agoHugging Face15wangekxy /classical-tcm-canon Classical Chinese Medicine Canon — 中医经典文本数据集 (v1) A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon), 难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics — assembled from public-domain source works. Useful for LLM training, RAG, and search over classical Chinese medicine / 中医药 literature. Summary 115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.tabulartext-generationn<1K0 likes113 downloads3mo agoHugging Face16Mxode /Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。 一篇样例如下: 古训《增广贤文》 昔时贤文,诲汝谆谆,集韵增文,多见多闻。 观今宜鉴古,无古不成今。 知己知彼,将心比心。 酒逢知己饮,诗向会人吟。 相识满天下,知心能几人。 相逢好似初相识,到老终无怨恨心。 近水知鱼性,近山识鸟音。 易涨易退山溪水,易反易覆小人心。 运去金成铁,时来铁似金,读书须用意,一字值千金。 texttext-generation100K<n<1M5 likes97 downloads1y agoHugging Face17Sudnya /classic-eda-c-trajectoriesgated nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 200 rounds. Every model turn is one row, including the ones that went nowhere. This is a partial snapshot. 130 of 1… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.tabulartext-generation1K<n<10K0 likes65 downloads1h agoHugging Face18Sudnya /test-subset-classic-eda nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 100 rounds. Every model turn is one row, including the ones that went nowhere. What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.tabulartext-generation1K<n<10K0 likes65 downloads2d agoHugging Face19gdiamos /classic-eda Classic EDA - Period Software Task Specifications 1020 specifications for software that plausibly could have been written between 1985 and 1996, mined from two in-era archives and shaped as coding tasks with machine-checkable requirements. Each record names a program, describes it in a paragraph, and states 3-12 atomic requirements plus an explicit interface contract (argv, stdin, stdout, exit codes) so a grader can test an implementation. Why period software… See the full description on the dataset page: https://huggingface.co/datasets/gdiamos/classic-eda.texttext-generation1K<n<10K0 likes54 downloads26d agoHugging Face20PoetryMTEB /Appreciation-of-Chinese-Classical-Poetry Appreciation of Chinese Classical Poetry Chinese classical poetry with paired metadata and five-aspect literary analyses for PoetryMTEB / MTEB-style evaluation and computational poetics research. Poems are drawn from expert appreciation volumes (mainly Shanghai Lexicographical Publishing House dictionaries). The released analysis fields are LLM distillations (DeepSeek-V3.1) of those expert appreciation texts into five free-text facets. The original long-form appreciation prose… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/Appreciation-of-Chinese-Classical-Poetry.texttext-classification10K<n<100K0 likes53 downloads1mo agoHugging Face21gujilab /chinese-classical-bench Chinese Classical Bench 中国古典语言能力评测基准 — 6 个任务 × 100 题 = 600 道,覆盖翻译、断句、字义、典故、续写填空、现代→文言压缩。 📊 在线排行榜: 🤗 Space — chinese-classical-bench-leaderboard 🔗 评测代码 & runner: github.com/gujilab/chinese-classical-bench — eval runner(OpenAI 兼容端点)、打分器、排行榜聚合脚本 📦 配套语料集: gujilab/chinese-classical-corpus (CC0 公有领域) — 题目均从该语料抽样生成 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。这套基础设施(bench + corpus + 4 个论点实证实验)想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面(实证) 7 个主流 tokenizer 横评(详见下方… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-bench.texttext-generationn<1K0 likes53 downloads4mo agoHugging Face22happyme531 /classical-chinese-poetry-benchmark-70 English Readme see below (README由Claude 3.5 Sonnet生成) 中国古诗词大模型评测基准 简介 这是一个专门用于评测大语言模型在中国古诗词理解和生成方面能力的基准测试集。该基准包含了一个多样化的测试数据集和完整的评测框架,可用于系统性地评估和比较不同模型在古诗词领域的表现。 数据集说明 数据集(poetry_benchmark.jsonl)包含70个测试样本,涵盖以下维度: 题型分布: 对联补全 诗句填空 诗词识别 提示词补全 首尾互补 难度等级: 简单(easy) 中等(medium) 困难(hard) 朝代覆盖: 先秦至近现代 包括唐、宋、元、明、清等重要朝代 评测维度 评测框架从以下维度对模型进行全面评估: 整体准确率 不同题型的表现 不同难度等级的表现 不同朝代诗词的掌握程度 评测结果 模型 blank_filling couplet find_poetry… See the full description on the dataset page: https://huggingface.co/datasets/happyme531/classical-chinese-poetry-benchmark-70.texttext-generationn<1K4 likes49 downloads2y agoHugging Face23ChaoticEconomist /Classical-Mechanics-Equations-Dataset_SFT-or-LoRA Classical Mechanics Equations Dataset (SFT / LoRA Ready) A structured dataset of 64 classical mechanics equations from Newtonian, Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning rows across three task types: equation explanation, Q&A, and derivation. Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and equation understanding tasks. Overview Property Value Domain Classical Mechanics (Physics) Total rows 448 Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face24catherinearnett /classical_armenian_pd Classical Armenian Public Domain Literature This dataset consists of 102 Classical Armenian texts in the public domain, which were collected from the Eastern Armenian National Corpus. A list of the works is provided below. Full list of works List of Works Աբովյան Խաչատուր՝ Առաջին սերը (First Love by Khachatur Abovian) Աբովյան Խաչատուր՝ Պարապ վախտի խաղալիք (Idle Time Toy by Khachatur Abovian) Աբովյան Խաչատուր՝ Թուրքի աղջիկը (The Turkish Girl by Khachatur Abovian)… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/classical_armenian_pd.texttext-generationn<1K0 likes26 downloads6mo agoHugging Face25wangekxy /classical-chinese-punctuation Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset) This is a commercial dataset. A free 200-record sample is provided below (sample.jsonl); the full 5.3M-pair corpus is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — built from public-domain classical works. The task Restore punctuation and sentence segmentation (句读) to unpunctuated Classical Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.texttext-generationn<1K0 likes24 downloads3mo agoHugging Face26wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes17 downloads3mo agoHugging Face27sunorme /Chinese-Classics-Partial偶然找到的 200 多篇古籍相关的纯 txt 文件,简单洗了一下,去除了部分噪声和空白行。 一篇样例如下: 古训《增广贤文》 昔时贤文,诲汝谆谆,集韵增文,多见多闻。 观今宜鉴古,无古不成今。 知己知彼,将心比心。 酒逢知己饮,诗向会人吟。 相识满天下,知心能几人。 相逢好似初相识,到老终无怨恨心。 近水知鱼性,近山识鸟音。 易涨易退山溪水,易反易覆小人心。 运去金成铁,时来铁似金,读书须用意,一字值千金。 texttext-generation100K<n<1M0 likes16 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.