datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tasklist-haiku4.5-6000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
anthropic/claude-haiku-4.5
Temperature
0.9
Total Tasks
5828
Concurrency
10 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 04:11:22
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
conversation
10.0%
creative
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-haiku4.5-6000x-unfiltered.nz-traditional-haiku
Traditional Japanese Haiku Dataset (Edo–Meiji Era)
⚠️ Work in Progress — Pre-Release Draft
This dataset is not yet ready for general use. It is being uploaded primarily as a personal backup snapshot during active development. Schema, annotations, and documentation may change without notice. Approximately 22% of records (~2,800) are still flagged annotation_status = "needs_review" and have not yet undergone human review.
If you arrived here unexpectedly, please check back later —… See the full description on the dataset page: https://huggingface.co/datasets/Rootport/nz-traditional-haiku.KYS-Claude-Haiku-50K-Labeled
KYS-Claude-Haiku-50K-Labeled
The LLM-annotated corpus used to distil the ModernBERT quality scorer in Know Your Sources: Data
Selection Matters when Rewriting for Data-Constrained Pretraining.
50,427 documents, each scored by Claude Haiku 4.5 on a five-criterion rubric.
The repo name rounds to 50K. The true row count is 50,427.
Composition
Source
Documents
DCLM-RefinedWeb (mix = "dclm-rw")
49,998
OpenWebMath (mix = "openwebmath")
218… See the full description on the dataset page: https://huggingface.co/datasets/blab-jhu/KYS-Claude-Haiku-50K-Labeled.
