datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
haiku_dpo
🌸 Haiku DPO 🌸
In data, words flow,
Teaching AI the art of
Haiku, line by line.
Dataset Card for Haiku DPO
This a synthetic dataset of haikus. The dataset is constructed with the goal of helping to train LLMs to be more 'technically' competent at writing haikus.
Dataset Details
The data consists of a few different components that are described in more detail below but the key components are:
a column of synthetically generated user prompts requesting a… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku_dpo.haiku
Dataset Card for Haiku Data
Taur_CoT_Analysis_Project___claude-3-haiku-20240307claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 33.52 (USD)
Total tokens (input + output): 6.79 M
SPIDER_SQL_synth_data_w_Claude3_Haikuhaiku_prompts🌸 Synthetic Haiku Prompts 🌸
In data's embrace,Synthetic haiku wishes bloom,
Code-born poetry.
Dataset Card for Synthetic Haiku Prompts
Dataset Details
This is a dataset of synthetic prompts that aims to replicate user requests to a chat model for a haiku about a given topic. The data was generated using the distilabel library using teknium's OpenHermes-2.5-Mistral-7B model. The prompts were generated from a seed list of terms and an adapted version of the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku_prompts.20260429_mini-v2.2.6_haiku-4-5modern_haikuThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.haiku_333K
Dataset Card for haiku_333K
This dataset contains 333,333 synthetic haiku. Just haiku. Nothing more, nothing less.
Dataset Details
Dataset Description
haiku_333K is a collection of machine-generated haiku following the traditional 5-7-5 syllable pattern. Each entry consists solely of the haiku text, making it a clean, focused dataset for text generation and analysis. The number 333,333 was chosen because good things come in threes, and we committed to the bit.… See the full description on the dataset page: https://huggingface.co/datasets/taucris/haiku_333K.claude-haiku-4.5-1700xThis is a non-reasoning dataset created using Claude Haiku 4.5. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 19.24 (USD)
Total tokens (input + output): 3.91 M
eval-terminal-bench-2.0-claude-haiku-4-5-20251001-20260115_165217haiku-kto-raw-argilla
Dataset Card for haiku-kto-raw-argilla
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku-kto-raw-argilla.haiku-analysissecalign-dbg-haiku-python-allsecalign-dbg-haiku-javascript-allhaiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/shigr3/haiku.persuasiveness-leaderboard-inverted-claude_3_5_haikuHaikuExplanationBitextMining
HaikuExplanationBitextMining
Monolingual Japanese bitext mining for PoetryMTEB: each pair aligns a haiku with its contextual explanation from shigr3/haiku.
Config
Direction
Description
jpn-jpn
Japanese ↔ Japanese
source_text = haiku; target_text = explanation
Dataset Card
Item
Description
Source
shigr3/haiku
Languages
Japanese (ja), monolingual parallel pairs
Size
train=113; test=29
Pair type
Haiku ↔ curated explanation (same… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/HaikuExplanationBitextMining.HaikuSeasonClassification
HaikuSeasonClassification
Single-label Japanese haiku season classification (春/夏/秋/冬) for PoetryMTEB, derived from shigr3/haiku.
Dataset Card
Item
Description
Source
shigr3/haiku
Languages
Japanese (ja)
Unit
Haiku text (poem)
Classes
4 seasons
Size
train=113; test=29
Splits
Stratified by season ≈ 80% / 20%, seed=42
License
CC BY 4.0 (same as upstream)
Evaluation metrics
Classification on embeddings: accuracy, macro/weighted F1… See the full description on the dataset page: https://huggingface.co/datasets/PoetryMTEB/HaikuSeasonClassification.haiku-cot-synthetic
CoT-Self-Instruct Synthetic Data
This dataset contains synthetic instruction data generated using the Chain-of-Thought Self-Instruct methodology.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Generation Model: Qwen/Qwen3-14B
Task Type: instruction
Filter Method: none
Generated Examples: 10
After Filtering: 10 (100.0% acceptance rate)
Generation Date: 2025-08-01 15:55:14 UTC
Methodology
Generated using CoT-Self-Instruct, which:
Uses… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/haiku-cot-synthetic.secalign-dbg-haiku-java-allhaiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/mizr3/haiku.haiku-vul-inducing-instructions-clusteredhaiku_concept_questions
Haiku Concept Questions
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop chunks… See the full description on the dataset page: https://huggingface.co/datasets/msaramhassan/haiku_concept_questions.strl-main-ec-tau2bench_telecom_haiku45_train_all-gc-claude_client_strl_dplm-mc-claude_ag-r0haiku-zh
诵读俳句吧
俳句数据集
大模型编写回复
人工再审核
人力有穷时
我只简单审核过
无明显错误
本为打油诗
不管有没有美感
就是图一乐
亦可DPO
让模型生成拒绝
配对数据集
strl-main-ec-tau2bench_retail_haiku45_train_all-gc-claude_client_strl-mc-claude_agent_so-r0DCAgent_dev_set_71_tasks_anthropic_claude-haiku-4-5_20251120_1413091k_pretraining_research_documents_Haiku45piserini_bcp_haiku
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
This repository contains the dataset and evaluation results for Pi-Serini, a search agent workspace designed to investigate whether lexical retrievers (specifically BM25) are sufficient when paired with frontier Large Language Models (LLMs) in an agentic loop.
Paper: Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
Repository: https://github.com/justram/pi-serini
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/ricky42613/piserini_bcp_haiku.
