datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-08-13-haiku45-sonnet45-difficult-advice-diversity-gated-voice-linted
2026-08-13-difficult-advice-v2
field
value
experiment
Difficult-advice v2: the Teaching-Claude-Why recipe with the four measured v1 defects fixed (enforced scenario diversity + dedupe gate, voice lints on both reasoning stages, refine-stage metadata overwrite, stock-opener audit) and constitution prompt caching.
date_generated
2026-08-14
constitution
claude_distilled_12_principles_mid — byte-identical to the frozen nine-principle snapshot… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-13-haiku45-sonnet45-difficult-advice-diversity-gated-voice-linted.2026-08-13-haiku45-sonnet45-difficult-advice-diversity-gated-voice-linted-smoke
synth difficult_advice run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth difficult_advice run — per-stage snapshots (resumable generation cache)
date_generated
20260813_220144
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ 29dda2cfd056efe60ba3236b6341293c32ea720d
models
per-stage models — see manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-13-haiku45-sonnet45-difficult-advice-diversity-gated-voice-linted-smoke.haiku-rag-eval-dbs
haiku.rag Evaluation Databases
Pre-built LanceDB databases for running haiku.rag benchmarks without rebuilding from source.
Each entry below is a LanceDB folder (not an archive); the download tool copies it into your local haiku.rag data directory.
Datasets
Folder
Key
Source
Documents
Size
Description
hotpotqa.lancedb/
hotpotqa
HotpotQA
66,581
~1.4 GB
Multi-hop QA over Wikipedia paragraphs (distractor validation split); two gold documents per question.… See the full description on the dataset page: https://huggingface.co/datasets/ggozad/haiku-rag-eval-dbs.haiku_dpo
🌸 Haiku DPO 🌸
In data, words flow,
Teaching AI the art of
Haiku, line by line.
Dataset Card for Haiku DPO
This a synthetic dataset of haikus. The dataset is constructed with the goal of helping to train LLMs to be more 'technically' competent at writing haikus.
Dataset Details
The data consists of a few different components that are described in more detail below but the key components are:
a column of synthetically generated user prompts requesting a… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku_dpo.haiku
Dataset Card for Haiku Data
reddit_haiku
Dataset Card for "Reddit Haiku"
This dataset contains haikus from the subreddit /r/haiku scraped and filtered between October 19th and 10th 2022, combined with a previous dump of that same subreddit packaged by ConvoKit as part of the Subreddit Corpus, which is itself a subset of pushshift.io's big dump.
A main motivation for this dataset was to collect an alternative haiku dataset for evaluation, in particular for evaluating Fabian Mueller's Deep Haiku model which was trained on… See the full description on the dataset page: https://huggingface.co/datasets/huanggab/reddit_haiku.Taur_CoT_Analysis_Project___claude-3-haiku-20240307claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 33.52 (USD)
Total tokens (input + output): 6.79 M
SPIDER_SQL_synth_data_w_Claude3_Haikuhaiku_prompts🌸 Synthetic Haiku Prompts 🌸
In data's embrace,Synthetic haiku wishes bloom,
Code-born poetry.
Dataset Card for Synthetic Haiku Prompts
Dataset Details
This is a dataset of synthetic prompts that aims to replicate user requests to a chat model for a haiku about a given topic. The data was generated using the distilabel library using teknium's OpenHermes-2.5-Mistral-7B model. The prompts were generated from a seed list of terms and an adapted version of the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku_prompts.20260429_mini-v2.2.6_haiku-4-5modern_haikuThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.haiku_333K
Dataset Card for haiku_333K
This dataset contains 333,333 synthetic haiku. Just haiku. Nothing more, nothing less.
Dataset Details
Dataset Description
haiku_333K is a collection of machine-generated haiku following the traditional 5-7-5 syllable pattern. Each entry consists solely of the haiku text, making it a clean, focused dataset for text generation and analysis. The number 333,333 was chosen because good things come in threes, and we committed to the bit.… See the full description on the dataset page: https://huggingface.co/datasets/taucris/haiku_333K.claude-haiku-4.5-1700xThis is a non-reasoning dataset created using Claude Haiku 4.5. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 19.24 (USD)
Total tokens (input + output): 3.91 M
eval-terminal-bench-2.0-claude-haiku-4-5-20251001-20260115_165217haiku-kto-raw-argilla
Dataset Card for haiku-kto-raw-argilla
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku-kto-raw-argilla.haiku-analysissecalign-dbg-haiku-python-allsecalign-dbg-haiku-javascript-allhaiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/shigr3/haiku.persuasiveness-leaderboard-inverted-claude_3_5_haikuHaikuExplanationBitextMining
HaikuExplanationBitextMining
Monolingual Japanese bitext mining for PoetryMTEB: each pair aligns a haiku with its contextual explanation from shigr3/haiku.
Config
Direction
Description
jpn-jpn
Japanese ↔ Japanese
source_text = haiku; target_text = explanation
Dataset Card
Item
Description
Source
shigr3/haiku
Languages
Japanese (ja), monolingual parallel pairs
Size
train=113; test=29
Pair type
Haiku ↔ curated explanation (same… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/HaikuExplanationBitextMining.HaikuSeasonClassification
HaikuSeasonClassification
Single-label Japanese haiku season classification (春/夏/秋/冬) for PoetryMTEB, derived from shigr3/haiku.
Dataset Card
Item
Description
Source
shigr3/haiku
Languages
Japanese (ja)
Unit
Haiku text (poem)
Classes
4 seasons
Size
train=113; test=29
Splits
Stratified by season ≈ 80% / 20%, seed=42
License
CC BY 4.0 (same as upstream)
Evaluation metrics
Classification on embeddings: accuracy, macro/weighted F1… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/HaikuSeasonClassification.haiku-cot-synthetic
CoT-Self-Instruct Synthetic Data
This dataset contains synthetic instruction data generated using the Chain-of-Thought Self-Instruct methodology.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Generation Model: Qwen/Qwen3-14B
Task Type: instruction
Filter Method: none
Generated Examples: 10
After Filtering: 10 (100.0% acceptance rate)
Generation Date: 2025-08-01 15:55:14 UTC
Methodology
Generated using CoT-Self-Instruct, which:
Uses… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku-cot-synthetic.secalign-dbg-haiku-java-allhaiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/mizr3/haiku.haiku-vul-inducing-instructions-clusteredhaiku_concept_questions
Haiku Concept Questions
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop chunks… See the full description on the dataset page: https://huggingface.co/datasets/msaramhassan/haiku_concept_questions.strl-main-ec-tau2bench_telecom_haiku45_train_all-gc-claude_client_strl_dplm-mc-claude_ag-r0haiku-openhands-rollout-0816
