CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01touati-kamel /TinyStories-Algerian-Darijatabular10K<n<100K0 likes2.3k downloads21d agoHugging Face02algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes165 downloads8d agoHugging Face03Pondsiders /tinystories-gpt4-instruct tinystories-gpt4-instruct Request→story pairs for supervised fine-tuning of small language models, derived from karpathy/tinystories-gpt4-clean. Each example pairs a natural-language request ("Can you tell me a story about a boy named Tim?") with a TinyStories story that satisfies it. The dataset lives on Hugging Face; the notebook that generates it lives on GitHub. This is not roneneldan/TinyStoriesInstruct. That dataset frames its tasks in a structured format (Words:… See the full description on the dataset page: https://huggingface.co/datasets/Pondsiders/tinystories-gpt4-instruct.tabulartext-generation10K<n<100K0 likes88 downloads25d agoHugging Face04EXOROBOURII /Stanza-TinyStories Dataset Card for Stanza-TinyStories-2 Dataset Summary Stanza-TinyStories-2 is a structurally and morphologically enriched iteration of the TinyStories dataset (Eldan and Li, 2023). This dataset projects the 1D synthetic text generated by large language models into a fully resolved grammatical and topological space. Every sentence in the 2.7-million-story training split and the 21,000-story validation split has been deterministically parsed to extract Universal… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-TinyStories.tabular100K<n<1M0 likes74 downloads5mo agoHugging Face05DaveGabe /TinyStoriesV2_cleaned-voc2048-seq256-overlap25tabular100K<n<1M0 likes56 downloads1y agoHugging Face06HydraLM /TinyStoriesInstruct-standardized Dataset Card for "TinyStoriesInstruct-standardized" More Information needed tabular1M<n<10M0 likes51 downloads3y agoHugging Face07alexliap /tinystories-gr TinyStories-GR A full Modern Greek translation of the TinyStories dataset (~2.1 million short English children's stories), with AI-generated quality scores for each translation. Dataset Description TinyStories-GR was generated by running the entire TinyStories corpus through a two-stage AI pipeline: Translation — each English story was translated to Modern Greek by Google Gemini (gemini-3.1-flash-lite-preview) Evaluation — each translation was independently scored (1–5)… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/tinystories-gr.tabulartranslation1M<n<10M0 likes32 downloads7mo agoHugging Face08DaveGabe /TinyStoriesV2_cleaned-inv-first-voc2048-seq256-overlap25tabular100K<n<1M0 likes28 downloads1y agoHugging Face09ZachW /gpt-oss-20b_tinystories-val1pct-raw openai/gpt-oss-20b — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: openai/gpt-oss-20b Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gpt-oss-20b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes25 downloads5mo agoHugging Face10DaveGabe /TinyStoriesV2_cleanedtabular100K<n<1M0 likes18 downloads6mo agoHugging Face11ZachW /olmo-3-7b-instruct_tinystories-val1pct-raw allenai/OLMo-3-7B-Instruct — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: allenai/OLMo-3-7B-Instruct Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-instruct_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes18 downloads5mo agoHugging Face12ZachW /qwen3-8b_tinystories-val1pct-raw Qwen/Qwen3-8B — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-8B Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face1313point5 /reverse-text-tinystories-hard Reverse Text TinyStories Hard This is the hard-difficulty TinyStories dataset for the reverse-text task. Splits train: 4000 rows test: 500 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories using non-overlapping word windows with cross-dataset prefix checks. Difficulty Rule hard rows only character-count range: 99-160 Notes The reverse answer is not stored because the reverse-text… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-hard.tabulartext-generation1K<n<10K0 likes15 downloads6mo agoHugging Face14ZachW /nanbeige4-3b-thinking-2511_tinystories-val1pct-raw Nanbeige/Nanbeige4-3B-Thinking-2511 — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: Nanbeige/Nanbeige4-3B-Thinking-2511 Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes15 downloads5mo agoHugging Face15ZachW /qwen3-32b_tinystories-val1pct-raw Qwen/Qwen3-32B — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-32B Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes13 downloads5mo agoHugging Face16ZachW /gemma-3-27b-it_tinystories-val1pct-raw google/gemma-3-27b-it — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: google/gemma-3-27b-it Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes12 downloads5mo agoHugging Face17ZachW /gemma-4-31b-it_tinystories-val1pct-raw google/gemma-4-31b-it — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes12 downloads5mo agoHugging Face18ZachW /llama-3.1-8b-instruct_tinystories-val1pct-raw meta-llama/Llama-3.1-8B-Instruct — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: meta-llama/Llama-3.1-8B-Instruct Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes12 downloads5mo agoHugging Face19haf1g /tinystories-darija-1ktabularn<1K0 likes11 downloads1y agoHugging Face20davanstrien /tinystories-atlas-datatabular1M<n<10M0 likes10 downloads7mo agoHugging Face21ZachW /olmo-3-7b-think_tinystories-val1pct-raw allenai/OLMo-3-7B-Think — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: allenai/OLMo-3-7B-Think Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes10 downloads5mo agoHugging Face22avgJo3 /tinystories-strattabular10K<n<100K0 likes9 downloads3mo agoHugging Face23ZachW /qwen3-30b-a3b_tinystories-val1pct-raw Qwen/Qwen3-30B-A3B — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-30B-A3B Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-30b-a3b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes8 downloads5mo agoHugging Face24bigshishiga /tiny-stories-10ktabular10K<n<100K0 likes7 downloads10mo agoHugging Face25codewithdark /tinystories-urdutabular100K<n<1M0 likes5 downloads1y agoHugging Face2613point5 /reverse-text-tinystories-easy-smoke Reverse Text TinyStories Easy Smoke This is a small smoke-test dataset for the reverse-text task. Splits train: 12 rows test: 4 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories by taking non-overlapping word windows and keeping only prompts that fall in the easy character-length bucket. Difficulty Rule All rows in this dataset are easy examples with prompt lengths in the 20-74 character range.… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-easy-smoke.tabulartext-generationn<1K0 likes5 downloads6mo agoHugging Face2713point5 /reverse-text-tinystories-medium Reverse Text TinyStories Medium This is the medium-difficulty TinyStories dataset for the reverse-text task. Splits train: 4000 rows test: 500 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories using non-overlapping word windows with cross-dataset prefix checks. Difficulty Rule medium rows only character-count range: 75-98 Notes The reverse answer is not stored because the… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-medium.tabulartext-generation1K<n<10K0 likes5 downloads6mo agoHugging Face28ZachW /mistral-small-3.2-24b-instruct-2506_tinystories-val1pct-raw mistralai/Mistral-Small-3.2-24B-Instruct-2506 — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes5 downloads5mo agoHugging Face2913point5 /reverse-text-tinystories-easy Reverse Text TinyStories Easy This is the easy-difficulty TinyStories dataset for the reverse-text task. Splits train: 4000 rows test: 500 rows Columns prompt char_count word_count source Source Derived from roneneldan/TinyStories using non-overlapping word windows with cross-dataset prefix checks. Difficulty Rule easy rows only character-count range: 20-74 Notes The reverse answer is not stored because the reverse-text… See the full description on the dataset page: https://huggingface.co/datasets/13point5/reverse-text-tinystories-easy.tabulartext-generation1K<n<10K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.