CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01neonforestmist /smolgpt-markdown-stories SmolGPT-Fables Stories A deterministic, text-only corpus of 96,000 original English Markdown stories built for SmolGPT-Fables. Every row is one complete supervised story example with an exact prompt / completion boundary, a requested scene count from one to six, and plain-language conditioning fields. No model, API, browser, or network service was used to create this dataset. Dataset summary 96,000 stories across 96,000 isolated story families 25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.tabulartext-generation10K<n<100K0 likes136 downloads2mo agoHugging Face02neogenesislab /korean-llm-citation-baseline-2026 DOI This dataset is citable via DataCite DOI 10.5281/zenodo.20018479 (Zenodo record). Cite as: @dataset{neogenesis_20018479, author = {Heo, Yesol and Neo Genesis Lab}, title = {Korean LLM Citation Baseline 2026 (Neo Genesis GEO Measurement)}, year = 2026, publisher = {Zenodo}, doi = {10.5281/zenodo.20018479}, url = {https://doi.org/10.5281/zenodo.20018479} } Korean LLM Citation Baseline 2026 (Neo Genesis GEO… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-llm-citation-baseline-2026.tabulartext-generationn<1K0 likes49 downloads5mo agoHugging Face03dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes38 downloads8mo agoHugging Face04neogenesislab /whylab-gemini-2-5-docker-validation 🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced. DOI This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.tabularothern<1K0 likes25 downloads4mo agoHugging Face05giseldo /neo_ara_v2tabulartext-generation10K<n<100K1 likes22 downloads1y agoHugging Face06giseldo /neo_ara_v1tabulartext-generation10K<n<100K1 likes6 downloads1y agoHugging Face07leo20240112 /pile-neox-uint16-partsTokenized uint16 shard parts for language-model pretraining. Original source: The Pile / NeoX-style preprocessing. tabulartext-generationn<1K0 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.