CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.5k downloads2mo agoHugging Face02greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes669 downloads2mo agoHugging Face03NuBerea /classical-greekgated Classical Greek Corpus Ancient and classical Greek (grc) text segments drawn from the open scholarly corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the classical/secular comparand within the NuBerea corpus estate, alongside its biblical, Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon (Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.tabulartext-generation10M<n<100M0 likes439 downloads12d agoHugging Face04AUEB-NLP /greek-bar-bench Dataset Card for GreekBarBench 🇬🇷🏛️⚖️ GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts. This repository hosts two related benchmarks: Benchmark Subsets Task GreekBarBench (GBB) greekbarbench, gbb-jme Free-text legal reasoning with citations, and LLM-judge meta-evaluation GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.tabularquestion-answering1K<n<10K5 likes203 downloads2d agoHugging Face05fffoivos /hplt-greek-ge8-no-mt-clean60-wave4 HPLT Greek GE8 No-MT Clean60 Wave4 A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass. Snapshot Rows: 48728774 Data parquet files: 250 Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60 Quality bins: 8, 9, 10 MT/register filtering: applied before this release Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.tabulartext-generation10M<n<100M0 likes181 downloads4mo agoHugging Face06kenpusney /greathangpt-classical-chinese GreatHanGPT 古汉语数据集 数据集描述 这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。 数据来源 来源 内容 链接 chinese-poetry 唐诗宋词、楚辞、诗经、四书五经 GitHub Werneror/Poetry 先秦到清末诗词,按朝代分 GitHub CBETA 大正藏佛经 GitHub 数据规模 指标 数值 总记录数 2,400,939 总字符数 450,496,972 估计token数 ~300M 时代分布 时代 记录数 字符数 占比 先秦 1,376 14,493,846 3.2% 汉魏 16,450 14,348,817 3.2% 隋唐 729,162 122,265,102 27.1% 两宋 728,569 97,907,260 21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.tabulartext-generation1M<n<10M0 likes161 downloads3mo agoHugging Face07greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face08gretelai /synthetic_multilingual_llm_prompts Image generated by DALL-E. See prompt for more details 📝🌐 Synthetic Multilingual LLM Prompts Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.tabulartext-generation1K<n<10K11 likes128 downloads2y agoHugging Face09yoonholee /poetry-greats-public-domain Poetry Greats Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study. Contents 4,090 poems across 29 books and 20 poets: Poet Poems Samuel Taylor Coleridge 913 H. W. Longfellow 616 Christina Rossetti 459 Emily Dickinson 446 Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.tabulartext-generation1K<n<10K0 likes80 downloads5mo agoHugging Face10tadad /diorisis-ancient-greek Diorisis Ancient Greek Corpus A Hugging Face conversion of Alessandro Vatri and Barbara McGillivray's Diorisis Ancient Greek Corpus for the BigLAM community. Diorisis contains 820 literary texts from Homer through the fifth century CE, with automatic lemma, part-of-speech, and morphological annotations. The conversion combines the original XML headers with the JSON corpus and a checksum-pinned snapshot of the author's public per-file corrections. It retains Beta Code and adds… See the full description on the dataset page: https://huggingface.co/datasets/tadad/diorisis-ancient-greek.tabulartoken-classification100K<n<1M0 likes77 downloads20d agoHugging Face11gregH /OccuBench OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models Dataset Description OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation. Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.tabulartext-generationn<1K5 likes64 downloads5mo agoHugging Face12G-reen /cc-2021-raw cc-2021-raw English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.tabulartext-generation1M<n<10M0 likes36 downloads2mo agoHugging Face13JackHsieh /qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained.tabulartext-generation100K<n<1M0 likes32 downloads1mo agoHugging Face14JackHsieh /qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3 qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3 Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts (SFT: lr 3e-5, batch 256, 4-epoch cosine; this is step 4724, 2.0 epochs of data seen). Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each. Why this checkpoint: the best held-out perplexity on luna thoughts across the lr sweep Sampling: greedy (temperature 0), seed 0, one… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.tabulartext-generation100K<n<1M0 likes30 downloads1mo agoHugging Face15JackHsieh /qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3 qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3 Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts (SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 1440, 0.6 epochs of data seen). Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each. Why this checkpoint: the best measured uplift when its thoughts are scored through the trained luna consumer Sampling: greedy… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.tabulartext-generation100K<n<1M0 likes27 downloads1mo agoHugging Face16JackHsieh /qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained.tabulartext-generation100K<n<1M0 likes24 downloads1mo agoHugging Face17G-reen /cc-2020-raw cc-2020-raw English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.tabulartext-generation1M<n<10M0 likes7 downloads2mo agoHugging Face18fffoivos /glossapi-greek-nanochat-pretraining-datasetgated Glossapi Greek Nanochat Pretraining Dataset This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream. Current Snapshot Total rows: 49474947 Total characters: 248276390721 Included source datasets: 19 Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.tabulartext-generation10M<n<100M0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.