datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-haiku-4.5-high-reasoning-1700xThis is a reasoning dataset created using Claude Haiku 4.5 with reasoning effort set to high.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 33.52 (USD)
Total tokens (input + output): 6.79 M
20260429_mini-v2.2.6_haiku-4-5haiku_333K
Dataset Card for haiku_333K
This dataset contains 333,333 synthetic haiku. Just haiku. Nothing more, nothing less.
Dataset Details
Dataset Description
haiku_333K is a collection of machine-generated haiku following the traditional 5-7-5 syllable pattern. Each entry consists solely of the haiku text, making it a clean, focused dataset for text generation and analysis. The number 333,333 was chosen because good things come in threes, and we committed to the bit.… See the full description on the dataset page: https://huggingface.co/datasets/taucris/haiku_333K.claude-haiku-4.5-1700xThis is a non-reasoning dataset created using Claude Haiku 4.5. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Haiku 4.5 by fine-tuning already existing open-source LLMs.
This dataset includes an addition to our recently enhanced set of prompts to cover creative writing and multilingual creative writing.
Stats
Costs: $ 19.24 (USD)
Total tokens (input + output): 3.91 M
haiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/shigr3/haiku.haiku
Famous Japanese Haiku Dataset (with English/Japanese Explanations)
[日本語の案内は後半にあります / Japanese description is followed by the English version]
Welcome to the Famous Japanese Haiku Dataset! This dataset is a curated collection of traditional and modern masterworks of Japanese Haiku, complete with their authors, seasonal classification (season), specific seasonal keywords (season_word), and detailed contextual explanations.
🌸 What is Haiku? (English)
Haiku (俳句) is… See the full description on the dataset page: https://huggingface.co/datasets/mizr3/haiku.piserini_bcp_haiku
Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
This repository contains the dataset and evaluation results for Pi-Serini, a search agent workspace designed to investigate whether lexical retrievers (specifically BM25) are sufficient when paired with frontier Large Language Models (LLMs) in an agentic loop.
Paper: Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?
Repository: https://github.com/justram/pi-serini
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/ricky42613/piserini_bcp_haiku.tasklist-haiku4.5-6000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
anthropic/claude-haiku-4.5
Temperature
0.9
Total Tasks
5828
Concurrency
10 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 04:11:22
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
conversation
10.0%
creative
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-haiku4.5-6000x-unfiltered.haiku_4-5nz-traditional-haiku
Traditional Japanese Haiku Dataset (Edo–Meiji Era)
⚠️ Work in Progress — Pre-Release Draft
This dataset is not yet ready for general use. It is being uploaded primarily as a personal backup snapshot during active development. Schema, annotations, and documentation may change without notice. Approximately 22% of records (~2,800) are still flagged annotation_status = "needs_review" and have not yet undergone human review.
If you arrived here unexpectedly, please check back later —… See the full description on the dataset page: https://huggingface.co/datasets/Rootport/nz-traditional-haiku.KYS-Claude-Haiku-50K-Labeled
KYS-Claude-Haiku-50K-Labeled
The LLM-annotated corpus used to distil the ModernBERT quality scorer in Know Your Sources: Data
Selection Matters when Rewriting for Data-Constrained Pretraining.
50,427 documents, each scored by Claude Haiku 4.5 on a five-criterion rubric.
The repo name rounds to 50K. The true row count is 50,427.
Composition
Source
Documents
DCLM-RefinedWeb (mix = "dclm-rw")
49,998
OpenWebMath (mix = "openwebmath")
218… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-Claude-Haiku-50K-Labeled.haiku-ita-v0.2
Italian Haiku in ShareGpt format
Dataset Summary
The dataset contains haiku generated in italian, following specific instructions and rules for the Italian language
Citation (Prompts)
https://huggingface.co/datasets/davanstrien/haiku_prompts has been translated using chatgpt 3.5 turbo. Each and every haiku has been then generated using the following prompt:
Haiku Generation
The following prompt has been used:
Crea un haiku in italiano che segua… See the full description on the dataset page: https://huggingface.co/datasets/WasamiKirua/haiku-ita-v0.2.reddit_haiku_detoxblocksworld-6-qwq-reasoning-parts-low-haikusafe-statworx-haikuhaiku_datasetThis dataset is the complete dataset used to train and evaluate the gemma-3-1b-haikuspec model trained using the code in the https://github.com/axeld5/gemma_haiku.git repository.
Source columns correspond to whether it is the "evaluation", "sft" or "rl" set.
haiku-examples-169
