datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dans-Prosemaxx-Opus-WritingDans-Prosemaxx-GutenbergDans-Prosemaxx-Adventurelemonseed-prose
lemonseed-prose
LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments).
Format
JSON Lines (.jsonl), one example per line.
Provenance & License
Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0.
sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.Anime-AMA-ProseDataset of anime / game characters and VTubers being asked various questions referencing their wiki page.
Prompts were generated by GLM 4.6
Scenes were generated by GLM 4.6
Response Plan was generated by GLM 4.6
Initial response was generated by GLM 4.6
Rewrite (if enough overused words were found) was done by Gemma 3 27b.
The rewrites are generally lower quality than the response provided by GLM, but they offer some different prose / word choices if preferred. Additionally, some rewrites… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Anime-AMA-Prose.Dans-Prosemaxx-Gryphe-GPT4o-WritingPromptssometimesanotion__Qwen-14B-ProseStock-v4-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-14B-ProseStock-v4
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-14B-ProseStock-v4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-14B-ProseStock-v4-details.sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v13-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v13-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v13-Prose-DS-details.khayyam-challenge-prose-terra
Khayyam Challenge - Prose (Terra)
AI-generated Persian prose descriptions of the 20 classical poems in the
Khayyam Challenge benchmark,
produced by the model internally labeled Terra. Split low /
medium / long by poem length (low: 10, medium: 7, long: 3).
Each record contains everything in the
poems repo (id,
poet, title, form, verse_count, theme, text, ...) plus a
conversion object with the generated prose, so this repo is self-contained
-- no join required:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-terra.khayyam-challenge-prose-luna
Khayyam Challenge - Prose (Luna)
AI-generated Persian prose descriptions of the 20 classical poems in the
Khayyam Challenge benchmark,
produced by the model internally labeled Luna. Split low /
medium / long by poem length (low: 10, medium: 7, long: 3).
Each record contains everything in the
poems repo (id,
poet, title, form, verse_count, theme, text, ...) plus a
conversion object with the generated prose, so this repo is self-contained
-- no join required:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/artindnr/khayyam-challenge-prose-luna.retro-easy-prose-repair-diffs-v0.1
RetroInstruct Easy Prose Repair Diffs
This component of RetroInstruct trains language models to repair prose by outputting
a diff that patches its flaws. The dataset is made through backtranslation by
running a synthetic corruption pass over prose. I use mostly syntactic
corruptions made with traditional programs, which makes them 'easy' compared to
more subtle semantic problems that could be introduced by a neural network. The
text I backtranslate from was generated by Mixtral… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-easy-prose-repair-diffs-v0.1.sft-bm-prose
khursanirevo/sft-bm-prose
Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows).
Splits
split
rows
train
221,170
validation
11,635
Stratified 95/5 by source/category (seed=42).
Source files
data/midtrain/synth_hf_prose.jsonl
data/midtrain/synth_bm_50m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.sometimesanotion__lamarck-14b-prose-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/lamarck-14b-prose-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/lamarck-14b-prose-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__lamarck-14b-prose-model_stock-details.sometimesanotion__Qwenvergence-14B-v15-Prose-MS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v15-Prose-MS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v15-Prose-MS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v15-Prose-MS-details.Phi4-Mini-Prose2Tags-4B-Raw-Training-DataRaw data used to train USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
ProseFlow-Actions-v1
ProseFlow-Actions-v1 Dataset
Dataset Description
ProseFlow-Actions-v1 is a high-quality, diverse dataset of structured instructions designed for fine-tuning language models to act as versatile text-processing assistants. This dataset is the backbone of the local AI engine for the ProseFlow desktop application, a universal, hotkey-driven AI utility.
The dataset is composed of 1,805 examples (1742 training, 63 testing) across 88 unique "Actions". Each example is a… See the full description on the dataset page: https://huggingface.co/datasets/LSXPrime/ProseFlow-Actions-v1.sometimesanotion__Qwenvergence-14B-v3-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v3-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v3-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v3-Prose-details.sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-model_stock-details.sometimesanotion__Qwenvergence-14B-v6-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v6-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v6-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v6-Prose-details.sometimesanotion__Qwentinuum-14B-v6-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwentinuum-14B-v6-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwentinuum-14B-v6-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwentinuum-14B-v6-Prose-details.Dans-Prosemaxx-Cowriter-3-XLclaude-prose-flesch-basesometimesanotion__Qwenvergence-14B-v2-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v2-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v2-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v2-Prose-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details.sometimesanotion__Qwenvergence-14B-v12-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-details.Dans-Prosemaxx-InstructWriter-Continue-2sometimesanotion__Qwen2.5-7B-Gordion-v0.1-Prose-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-7B-Gordion-v0.1-Prose
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-7B-Gordion-v0.1-Prose
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-7B-Gordion-v0.1-Prose-details.Dans-Prosemaxx-Cowriter-3-LDans-Prosemaxx-InstructWriter-Continue
