datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kimi-k3-coding-and-debugging-traces
Kimi K3 Coding, Tool Use & Instruction Following Traces
582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables
below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.glm-5.2-coding-and-debugging-traces
GLM 5.2 Agent Traces
207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool-use, and agent trajectories
from GLM 5.2 (glm-5.2). The category and row-share tables
below describe the actual mix seen during training rather than assuming a
particular task domain.
This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.classical-greek
Classical Greek Corpus
Ancient and classical Greek (grc) text segments drawn from the open scholarly
corpora of the Perseus Digital Library and the OpenGreekAndLatin project — the
classical/secular comparand within the NuBerea corpus estate, alongside its biblical,
Second-Temple, and patristic Greek collections. Coverage runs from the archaic canon
(Homer, Hesiod, the tragedians, the historians, Plato, Aristotle) through Hellenistic
and imperial prose (Plutarch, Lucian, Galen)… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/classical-greek.greek-bar-bench
Dataset Card for GreekBarBench 🇬🇷🏛️⚖️
GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts.
This repository hosts two related benchmarks:
Benchmark
Subsets
Task
GreekBarBench (GBB)
greekbarbench, gbb-jme
Free-text legal reasoning with citations, and LLM-judge meta-evaluation
GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.hplt-greek-ge8-no-mt-clean60-wave4
HPLT Greek GE8 No-MT Clean60 Wave4
A standalone release of the filtered Greek HPLT slice used in the GlossAPI Greek pretraining corpus. It contains the full HPLT/ell_Grek_ge8_no_mt_clean60 source after the Wave4 re-cleaning and normalization pass.
Snapshot
Rows: 48728774
Data parquet files: 250
Source dataset value: HPLT/ell_Grek_ge8_no_mt_clean60
Quality bins: 8, 9, 10
MT/register filtering: applied before this release
Cleaner gate: greek_badness_score <= 60 before… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/hplt-greek-ge8-no-mt-clean60-wave4.greathangpt-classical-chinese
GreatHanGPT 古汉语数据集
数据集描述
这是一个用于训练古汉语大语言模型的数据集,包含从先秦到清初(1644-1722)的汉语文献。
数据来源
来源
内容
链接
chinese-poetry
唐诗宋词、楚辞、诗经、四书五经
GitHub
Werneror/Poetry
先秦到清末诗词,按朝代分
GitHub
CBETA
大正藏佛经
GitHub
数据规模
指标
数值
总记录数
2,400,939
总字符数
450,496,972
估计token数
~300M
时代分布
时代
记录数
字符数
占比
先秦
1,376
14,493,846
3.2%
汉魏
16,450
14,348,817
3.2%
隋唐
729,162
122,265,102
27.1%
两宋
728,569
97,907,260
21.7%… See the full description on the dataset page: https://huggingface.co/datasets/kenpusney/greathangpt-classical-chinese.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.poetry-greats-public-domain
Poetry Greats
Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study.
Contents
4,090 poems across 29 books and 20 poets:
Poet
Poems
Samuel Taylor Coleridge
913
H. W. Longfellow
616
Christina Rossetti
459
Emily Dickinson
446
Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.diorisis-ancient-greek
Diorisis Ancient Greek Corpus
A Hugging Face conversion of Alessandro Vatri and Barbara McGillivray's
Diorisis Ancient Greek Corpus for the
BigLAM community. Diorisis contains 820 literary texts from Homer through the fifth century CE,
with automatic lemma, part-of-speech, and morphological annotations.
The conversion combines the original XML headers with the
JSON corpus and a checksum-pinned snapshot of
the author's public per-file corrections. It retains Beta Code and adds… See the full description on the dataset page: https://huggingface.co/datasets/tadad/diorisis-ancient-greek.OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.cc-2021-raw
cc-2021-raw
English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained
qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained.qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3
qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3
Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a
Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts
(SFT: lr 3e-5, batch 256, 4-epoch cosine; this is step 4724, 2.0 epochs of data seen).
Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each.
Why this checkpoint: the best held-out perplexity on luna thoughts across the lr sweep
Sampling: greedy (temperature 0), seed 0, one… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-3e5-4ep-s4724-greedy.k-8.statml-arxiv-qwen3.qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3
qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3
Reasoning about the next 8 Qwen3 tokens of stat.ML arXiv LaTeX, generated by a
Qwen3-4B-Instruct-2507 distilled on gpt-5.6-luna thoughts
(SFT: lr 1e-5, batch 256, 1-epoch cosine; this is step 1440, 0.6 epochs of data seen).
Documents: JackHsieh/statML-arxiv-40M-20M, 4096 Qwen3 tokens each.
Why this checkpoint: the best measured uplift when its thoughts are scored through the trained luna consumer
Sampling: greedy… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained
qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained
Tokenized, tag-wrapped form of JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.
Each thought is wrapped as
<|note|>
This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next.
KEY: <last 8 prefix tokens>
VALUE: <thought>
<|/note|>
and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/qwen3-distill-1e5-1ep-s1440-greedy.k-8.statml-arxiv-qwen3.qwen3-ids.kv-tags-explained.cc-2020-raw
cc-2020-raw
English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.glossapi-greek-nanochat-pretraining-dataset
Glossapi Greek Nanochat Pretraining Dataset
This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream.
Current Snapshot
Total rows: 49474947
Total characters: 248276390721
Included source datasets: 19
Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.
