datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolly-15k-prompt-compression
Dolly-15k Prompt Compression
This dataset contains compressed versions of the Databricks Dolly-15k prompts. Each prompt was compressed using the gpt-5-nano model to minimize input tokens while preserving all constraints. You can explore the downstream model that relies on this data in the companion Space: Very Small Prompt Compression Demo.
Compression model: gpt-5-nano
Source dataset: databricks/databricks-dolly-15k
Rows: 15,000
Aggregate token savings: 289,540 → 215,219 tokens… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/dolly-15k-prompt-compression.databricks-dolly-15k-cleanset
Summary
databricks-dolly-15k-cleanset can be used to produced CLEANed up versions of the popular databricks-dolly-15k dataSET, which was used to fine-tune the Dolly 2.0. The original databricks-dolly-15k contains 15,000 human-annotated instruction-response pairs covering various categories. However, there are many low-quality responses, incomplete/vague prompts, and other problematic text lurking in the dataset (as with for all real-world instruction tuning datasets). We ran… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/databricks-dolly-15k-cleanset.
