datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lemonseed-prose
lemonseed-prose
LemonSeed — narrative prose anchor (TinyStories-derived, filtered to 120–900 char fragments).
Format
JSON Lines (.jsonl), one example per line.
Provenance & License
Derived from roneneldan/TinyStories (TinyStoriesV2-GPT4-train.txt), filtered. Upstream license: CDLA-Sharing-1.0.
sft-bm-prose
khursanirevo/sft-bm-prose
Bahasa Melayu prose-format text (long-form lessons + textbook-style, ~232k rows).
Splits
split
rows
train
221,170
validation
11,635
Stratified 95/5 by source/category (seed=42).
Source files
data/midtrain/synth_hf_prose.jsonl
data/midtrain/synth_bm_50m.jsonl
Schema
Each row is a JSON object. See the loader script for field details.
Provenance
Generated as part of MaLLaM 2026… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/sft-bm-prose.
