datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smolgpt-markdown-stories
SmolGPT-Fables Stories
A deterministic, text-only corpus of 96,000 original English
Markdown stories built for SmolGPT-Fables. Every row is one complete supervised
story example with an exact prompt / completion boundary, a requested scene
count from one to six, and plain-language conditioning fields.
No model, API, browser, or network service was used to create this dataset.
Dataset summary
96,000 stories across 96,000 isolated story families
25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.korean-llm-citation-baseline-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018479 (Zenodo record).
Cite as:
@dataset{neogenesis_20018479,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Korean LLM Citation Baseline 2026 (Neo Genesis GEO Measurement)},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018479},
url = {https://doi.org/10.5281/zenodo.20018479}
}
Korean LLM Citation Baseline 2026 (Neo Genesis GEO… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-llm-citation-baseline-2026.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.whylab-gemini-2-5-docker-validation
🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced.
DOI
This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.neo_ara_v2neo_ara_v1pile-neox-uint16-partsTokenized uint16 shard parts for language-model pretraining.
Original source: The Pile / NeoX-style preprocessing.
