datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
haiku_dpo
🌸 Haiku DPO 🌸
In data, words flow,
Teaching AI the art of
Haiku, line by line.
Dataset Card for Haiku DPO
This a synthetic dataset of haikus. The dataset is constructed with the goal of helping to train LLMs to be more 'technically' competent at writing haikus.
Dataset Details
The data consists of a few different components that are described in more detail below but the key components are:
a column of synthetically generated user prompts requesting a… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku_dpo.gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit
Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit)
The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels.
This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.nz-traditional-haiku
Traditional Japanese Haiku Dataset (Edo–Meiji Era)
⚠️ Work in Progress — Pre-Release Draft
This dataset is not yet ready for general use. It is being uploaded primarily as a personal backup snapshot during active development. Schema, annotations, and documentation may change without notice. Approximately 22% of records (~2,800) are still flagged annotation_status = "needs_review" and have not yet undergone human review.
If you arrived here unexpectedly, please check back later —… See the full description on the dataset page: https://huggingface.co/datasets/Rootport/nz-traditional-haiku.
