CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gujilab /chinese-classical-corpus Chinese Classical Corpus 🔗 源码 & 构建脚本: github.com/gujilab/chinese-classical-corpus — 完整抽取 pipeline、14 个 Python 脚本、验证套件 🎯 配套评测基准: gujilab/chinese-classical-bench — 500 道题 × 5 任务,测 LLM 古典文献能力(题目均从本语料抽样) 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。本语料集 + 配套评测想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面 —— 真成立 7 个主流 tokenizer 横评(tokenizer_study): DeepSeek-V3 /… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-corpus.texttext-generation1M<n<10M1 likes452 downloads4mo agoHugging Face02flowerone /chinese-classical-corpus Chinese Classical Corpus 🔗 Source code & build scripts: github.com/zi6me/chinese-classical-corpus — full extraction pipeline, 14 Python scripts, validation suite. 中国古典文献结构化语料集,含完整十三经 + 说文解字 + 资治通鉴 + 二十四史前 15 部,以及 197 万条古译今/今译古/断句指令对。 全部 CC0 公有领域,可商用、可改用、无附加限制。 Quick Start from datasets import load_dataset # 源语料 (12,005 条章节级记录, 17.2M 字) corpus = load_dataset("dzxr/chinese-classical-corpus", "corpus", split="train") # 古译今 / 今译古 双向指令数据 (1,924,378 条) translate =… See the full description on the dataset page: https://huggingface.co/datasets/flowerone/chinese-classical-corpus.texttext-generation1M<n<10M0 likes228 downloads5mo agoHugging Face03vhands /audio-event-classification-post-public audio-event-classification-post-public Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning. Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.textaudio-classification100K<n<1M1 likes192 downloads3mo agoHugging Face04CJJones /Wikipedia_RAG_QA_Classification 🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training 📊 Dataset Description This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning. 🖥️ Demo Interface: Discord Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.tabulartext-generation100K<n<1M1 likes113 downloads6mo agoHugging Face05Sudnya /classic-eda-c-trajectoriesgated nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 200 rounds. Every model turn is one row, including the ones that went nowhere. This is a partial snapshot. 130 of 1… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/classic-eda-c-trajectories.tabulartext-generation1K<n<10K0 likes73 downloads20h agoHugging Face06Sudnya /test-subset-classic-eda nano-rl trajectories Agent trajectories from nano-rl: an LLM is asked to specify, test and then implement a small C program, which is compiled and executed inside a confined sandbox and scored against tests the model wrote before it saw its own program. When the program fails, the model is shown the build log and the failing cases and asked to repair it, for up to 100 rounds. Every model turn is one row, including the ones that went nowhere. What is in it… See the full description on the dataset page: https://huggingface.co/datasets/Sudnya/test-subset-classic-eda.tabulartext-generation1K<n<10K0 likes72 downloads2d agoHugging Face07gujilab /chinese-classical-bench Chinese Classical Bench 中国古典语言能力评测基准 — 6 个任务 × 100 题 = 600 道,覆盖翻译、断句、字义、典故、续写填空、现代→文言压缩。 📊 在线排行榜: 🤗 Space — chinese-classical-bench-leaderboard 🔗 评测代码 & runner: github.com/gujilab/chinese-classical-bench — eval runner(OpenAI 兼容端点)、打分器、排行榜聚合脚本 📦 配套语料集: gujilab/chinese-classical-corpus (CC0 公有领域) — 题目均从该语料抽样生成 为什么做这个 中文(尤其文言文)常被说成"高密度优势"。这套基础设施(bench + corpus + 4 个论点实证实验)想把这个论点变成可验证的数字 —— 包括它在哪些场景成立、在哪些场景不成立。 Tokenizer 层面(实证) 7 个主流 tokenizer 横评(详见下方… See the full description on the dataset page: https://huggingface.co/datasets/gujilab/chinese-classical-bench.texttext-generationn<1K0 likes54 downloads4mo agoHugging Face08AmareshHebbar /pmjay-classifier-sft PM-JAY Health Benefit Package Classifier Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Medical specialty + procedure → PM-JAY HBP code, package name, and rate Why download this Automate PM-JAY / Ayushman Bharat claim processing. Map procedures to Health Benefit Packages for pre-authorization and reimbursement.… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pmjay-classifier-sft.texttext-generation10K<n<100K0 likes54 downloads3mo agoHugging Face09gdiamos /classic-eda Classic EDA - Period Software Task Specifications 1020 specifications for software that plausibly could have been written between 1985 and 1996, mined from two in-era archives and shaped as coding tasks with machine-checkable requirements. Each record names a program, describes it in a paragraph, and states 3-12 atomic requirements plus an explicit interface contract (argv, stdin, stdout, exit codes) so a grader can test an implementation. Why period software… See the full description on the dataset page: https://huggingface.co/datasets/gdiamos/classic-eda.texttext-generation1K<n<10K0 likes54 downloads26d agoHugging Face10happyme531 /classical-chinese-poetry-benchmark-70 English Readme see below (README由Claude 3.5 Sonnet生成) 中国古诗词大模型评测基准 简介 这是一个专门用于评测大语言模型在中国古诗词理解和生成方面能力的基准测试集。该基准包含了一个多样化的测试数据集和完整的评测框架,可用于系统性地评估和比较不同模型在古诗词领域的表现。 数据集说明 数据集(poetry_benchmark.jsonl)包含70个测试样本,涵盖以下维度: 题型分布: 对联补全 诗句填空 诗词识别 提示词补全 首尾互补 难度等级: 简单(easy) 中等(medium) 困难(hard) 朝代覆盖: 先秦至近现代 包括唐、宋、元、明、清等重要朝代 评测维度 评测框架从以下维度对模型进行全面评估: 整体准确率 不同题型的表现 不同难度等级的表现 不同朝代诗词的掌握程度 评测结果 模型 blank_filling couplet find_poetry… See the full description on the dataset page: https://huggingface.co/datasets/happyme531/classical-chinese-poetry-benchmark-70.texttext-generationn<1K4 likes52 downloads2y agoHugging Face11AmareshHebbar /insurance-classifier-sft Insurance Coverage Classifier (Stark Law DHS) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does CPT/HCPCS codes → Stark Law DHS classification + compliance notes Why download this Compliance automation for physician self-referral rules. Identify which services are Designated Health Services under Stark Law Section… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/insurance-classifier-sft.texttext-generation1K<n<10K0 likes50 downloads3mo agoHugging Face12hybrid-diff-ar /stack-v2-sparse-classes-10k Stack v2 Sparse Python Classes 10k This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 9,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.tabulartext-generation10K<n<100K0 likes41 downloads5mo agoHugging Face13w1z4rd3k /it-support-l1-ticket-classification IT Support L1 Multilingual Dataset Dataset Summary IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping. This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.texttext-classificationn<1K0 likes37 downloads5mo agoHugging Face14hybrid-diff-ar /stack-v2-sparse-classes-75kplus Stack v2 Sparse Python Classes 75kplus This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments. Splits train.jsonl: 74829 val.jsonl: 500 test.jsonl: 500 all.jsonl: 75829 Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.tabulartext-generation10K<n<100K0 likes33 downloads5mo agoHugging Face15hybrid-diff-ar /stack-v2-sparse-classes-36k Stack v2 Sparse Python Classes 36k This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 35,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.tabulartext-generation10K<n<100K0 likes27 downloads5mo agoHugging Face16wangekxy /classical-chinese-punctuation Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset) This is a commercial dataset. A free 200-record sample is provided below (sample.jsonl); the full 5.3M-pair corpus is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — built from public-domain classical works. The task Restore punctuation and sentence segmentation (句读) to unpunctuated Classical Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.texttext-generationn<1K0 likes25 downloads3mo agoHugging Face17wangekxy /classical-chinese-variant-collation Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset) This is a commercial dataset. A free 50-work preview sample is provided below (sample.jsonl, texts truncated); the full set with complete aligned texts is available upon request. 📧 To license / purchase or request a quote, email wangeksy@gmail.com. ✅ Cleared for commercial use — both members are public-domain classical works. What this is A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.tabulartext-generationn<1K0 likes20 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.