CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Seelee789 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M0 likes808 downloads2mo agoHugging Face02PocketDoc /Dans-Prosemaxx-Opus-Writingtextn<1K1 likes193 downloads2y agoHugging Face03docketx /us-pro-se US Pro Se — what the courts themselves tell people who have no lawyer Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them. 12,103 documents from 19 state court systems: 2,835 self-help guide pages, 727 instruction documents and 8,541 forms, 119,978,550 characters of text, each… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-pro-se.texttext-retrieval10K<n<100K0 likes89 downloads5d agoHugging Face04PocketDoc /Dans-Prosemaxx-Gutenbergtextn<1K1 likes88 downloads2y agoHugging Face05PocketDoc /Dans-Prosemaxx-Adventuretabularn<1K4 likes74 downloads2y agoHugging Face06open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v12-Prose-DS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.tabular10K<n<100K0 likes50 downloads2y agoHugging Face07zerofata /Anime-AMA-ProseDataset of anime / game characters and VTubers being asked various questions referencing their wiki page. Prompts were generated by GLM 4.6 Scenes were generated by GLM 4.6 Response Plan was generated by GLM 4.6 Initial response was generated by GLM 4.6 Rewrite (if enough overused words were found) was done by Gemma 3 27b. The rewrites are generally lower quality than the response provided by GLM, but they offer some different prose / word choices if preferred. Additionally, some rewrites… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Anime-AMA-Prose.text1K<n<10K7 likes46 downloads11mo agoHugging Face08MichaelAnthony /lemonseed-prose lemonseed-prose LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments). Format JSON Lines (.jsonl), one example per line. Provenance & License Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0. texttext-generation1K<n<10K0 likes45 downloads1mo agoHugging Face09PocketDoc /Dans-Prosemaxx-Gryphe-GPT4o-WritingPromptstextn<1K0 likes43 downloads2y agoHugging Face10open-llm-leaderboard /sometimesanotion__Qwen-14B-ProseStock-v4-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4 Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face11open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v13-Prose-DS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v13-Prose-DS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v13-Prose-DS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face12artindnr /khayyam-challenge-prose-terra Khayyam Challenge - Prose (Terra) AI-generated Persian prose descriptions of the 20 classical poems in the Khayyam Challenge benchmark, produced by the model internally labeled Terra. Split low / medium / long by poem length (low: 10, medium: 7, long: 3). Each record contains everything in the poems repo (id, poet, title, form, verse_count, theme, text, ...) plus a conversion object with the generated prose, so this repo is self-contained -- no join required: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-terra.tabularn<1K1 likes34 downloads2mo agoHugging Face13artindnr /khayyam-challenge-prose-luna Khayyam Challenge - Prose (Luna) AI-generated Persian prose descriptions of the 20 classical poems in the Khayyam Challenge benchmark, produced by the model internally labeled Luna. Split low / medium / long by poem length (low: 10, medium: 7, long: 3). Each record contains everything in the poems repo (id, poet, title, form, verse_count, theme, text, ...) plus a conversion object with the generated prose, so this repo is self-contained -- no join required: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-luna.tabularn<1K1 likes32 downloads2mo agoHugging Face14jdpressman /retro-easy-prose-repair-diffs-v0.1 RetroInstruct Easy Prose Repair Diffs This component of RetroInstruct trains language models to repair prose by outputting a diff that patches its flaws. The dataset is made through backtranslation by running a synthetic corruption pass over prose. I use mostly syntactic corruptions made with traditional programs, which makes them 'easy' compared to more subtle semantic problems that could be introduced by a neural network. The text I backtranslate from was generated by Mixtral… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-easy-prose-repair-diffs-v0.1.text1K<n<10K1 likes19 downloads2y agoHugging Face15khursanirevo /sft-bm-prose khursanirevo/sft-bm-prose Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows). Splits split rows train 221,170 validation 11,635 Stratified 95/5 by source/category (seed=42). Source files data/midtrain/synth_hf_prose.jsonl data/midtrain/synth_bm_50m.jsonl Schema Each row is a JSON object. See the loader script for field details. Provenance Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.tabulartext-generation100K<n<1M0 likes18 downloads3mo agoHugging Face16open-llm-leaderboard /sometimesanotion__lamarck-14b-prose-model_stock-detailsgated Dataset Card for Evaluation run of sometimesanotion/lamarck-14b-prose-model_stock Dataset automatically created during the evaluation run of model sometimesanotion/lamarck-14b-prose-model_stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__lamarck-14b-prose-model_stock-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face17open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v15-Prose-MS-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v15-Prose-MS Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v15-Prose-MS The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v15-Prose-MS-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face18USS-Inferprise /Phi4-Mini-Prose2Tags-4B-Raw-Training-DataRaw data used to train USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B) texttable-to-text100K<n<1M0 likes10 downloads5mo agoHugging Face19LSXPrime /ProseFlow-Actions-v1 ProseFlow-Actions-v1 Dataset Dataset Description ProseFlow-Actions-v1 is a high-quality, diverse dataset of structured instructions designed for fine-tuning language models to act as versatile text-processing assistants. This dataset is the backbone of the local AI engine for the ProseFlow desktop application, a universal, hotkey-driven AI utility. The dataset is composed of 1,805 examples (1742 training, 63 testing) across 88 unique "Actions". Each example is a… See the full description on the dataset page: https://huggingface.co/datasets/LSXPrime/ProseFlow-Actions-v1.text1K<n<10K2 likes9 downloads1y agoHugging Face20open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v3-Prose-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v3-Prose Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v3-Prose The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v3-Prose-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face21open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face22open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v6-Prose-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face23SeveralOptional /gemini-3.1-pro-2048-reasoning-1100xtext1K<n<10K1 likes8 downloads7mo agoHugging Face24open-llm-leaderboard /sometimesanotion__Qwentinuum-14B-v6-Prose-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v6-Prose Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v6-Prose The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v6-Prose-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face25PocketDoc /Dans-Prosemaxx-Cowriter-3-XLgatedtext100K<n<1M0 likes7 downloads2y agoHugging Face26semran1 /c4-pro-booktext100K<n<1M0 likes7 downloads1y agoHugging Face27open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v2-Prose-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v2-Prose Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v2-Prose The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v2-Prose-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face28open-llm-leaderboard /sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01 Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face29open-llm-leaderboard /sometimesanotion__Qwenvergence-14B-v12-Prose-detailsgated Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face30PocketDoc /Dans-Prosemaxx-InstructWriter-Continue-2gatedtext1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.