datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolly_shield
Dataset Card for Project Dolly Shield
This is a collection of Beatport and Spotify data that I found on Kaggle.
No additional Spotify or Beatport data was collected from their platforms directly. Spotify bans the use of their data for training AI models. I've decided I can use a dataset from Kaggle for my AI project but will not collect additional data from Spotify via their API services for AI model training.
More information on Spotify's AI policies can be found here.
The… See the full description on the dataset page: https://huggingface.co/datasets/uwsthoughts/dolly_shield.databricks_dolly_15k
Databricks Dolly task samples
Standalone task subsets derived from
databricks/databricks-dolly-15k at
revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a:
general_qa (source category: general_qa)
open_qa (source category: open_qa)
closed_qa (source category: closed_qa)
brainstorm (source category: brainstorming)
classify (source category: classification)
extract_information (source category: information_extraction)
summarize (source category: summarization)
creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.databricks__dolly-v2-7b-details
Dataset Card for Evaluation run of databricks/dolly-v2-7b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-7b-details.dolly-15k-prompt-compression
Dolly-15k Prompt Compression
This dataset contains compressed versions of the Databricks Dolly-15k prompts. Each prompt was compressed using the gpt-5-nano model to minimize input tokens while preserving all constraints. You can explore the downstream model that relies on this data in the companion Space: Very Small Prompt Compression Demo.
Compression model: gpt-5-nano
Source dataset: databricks/databricks-dolly-15k
Rows: 15,000
Aggregate token savings: 289,540 → 215,219 tokens… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/dolly-15k-prompt-compression.databricks__dolly-v2-12b-details
Dataset Card for Evaluation run of databricks/dolly-v2-12b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-12b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-12b-details.dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109dolly-15k-clustered-fulltext-modernbert-sweep-20250106dolly-15k-clustered-fulltext-less-sweep-20250106dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109databricks-dolly15k-semantic-complexity
Databricks - Dolly 15k – Enriched Variant (Instruction-Tuned with Semantic and Complexity Augmentation)
Overview
This dataset is a semantically enriched and complexity-aware extension of the original Databricks Dolly 15k, purpose-built for evaluating and training instruction-following models. Each sample is augmented with additional signals to enable more nuanced filtering, curriculum learning, and benchmark development across diverse NLP tasks.
Dataset Format
Each… See the full description on the dataset page: https://huggingface.co/datasets/GenAIDevTOProd/databricks-dolly15k-semantic-complexity.bsc-dolly-15k-en
BSC Dolly 15k EN
Reviewed version from the Argilla Dolly v2 English version, originally created by Databricks.
We provide two subsets: "annotated", where some instances were labelled with potential problems; and "filtered", which only contains the instances without the issues that we observed.
Annotation process
While analysing the Argilla Dolly v2 English version, we observed the following:
Task classification:
- There are three classes with context:… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/bsc-dolly-15k-en.databricks__dolly-v1-6b-details
Dataset Card for Evaluation run of databricks/dolly-v1-6b
Dataset automatically created during the evaluation run of model databricks/dolly-v1-6b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v1-6b-details.databricks__dolly-v2-3b-details
Dataset Card for Evaluation run of databricks/dolly-v2-3b
Dataset automatically created during the evaluation run of model databricks/dolly-v2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/databricks__dolly-v2-3b-details.rationale-databricks-dolly-cqa
Dataset Overview
Filtered and annotated version of the closed-question answering part (~1.5k datapoints) of the Databricks Dolly Dataset intended for the task of rationale extraction.
Citation
@article{pirenne2024exploration,
title={Exploration of Closed-Domain Question Answering Explainability Methods With a Sentence-Level Rationale Dataset},
author={Pirenne, Lize and Mokeddem, Samy and Ernst, Damien and Louppe, Gilles},
year={2024}
}… See the full description on the dataset page: https://huggingface.co/datasets/Inversta/rationale-databricks-dolly-cqa.databricks-dolly-15k-modernbert-train-kmeans-dim768-20250723databricks-dolly-15k_standardizeddatabricks-dolly-15k-cleanset
Summary
databricks-dolly-15k-cleanset can be used to produced CLEANed up versions of the popular databricks-dolly-15k dataSET, which was used to fine-tune the Dolly 2.0. The original databricks-dolly-15k contains 15,000 human-annotated instruction-response pairs covering various categories. However, there are many low-quality responses, incomplete/vague prompts, and other problematic text lurking in the dataset (as with for all real-world instruction tuning datasets). We ran… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/databricks-dolly-15k-cleanset.databricks-dolly-15k_standardizeddolly-15k-clustered-modernbert-64-20250111databricks-dolly-15k-modernbert-kmeans-dim768-normalize-20250130dollyaug-standardized_cluster_2
Dataset Card for "dollyaug-standardized_cluster_2"
More Information needed
dollyaug-standardized_cluster_4
Dataset Card for "dollyaug-standardized_cluster_4"
More Information needed
dollyqa_512dolly-15k-clustered-fulltext-agglomerative-16-20250103databricks-dolly-15k-tfidf-sweep-kmeans-dim10000-20250914databricks-dolly-15k-modernbert-split-kmeans-dim768-20250917dolly-llama-qa
Dataset Card for dolly-llama-qa
This dataset has been created with dataformer.
Dataset Details
Dataset Description
The dolly-llama-qa dataset is a synthetic QA pair dataset created using the context from databricks-dolly-15k. We used Meta-Llama-3-8B-Instruct and Meta-Llama-3.1-8B-Instruct models for the generation and evolution part. Openai's gpt-4o was used for evaluating the refined questions and refined answers.
Dataset Columns
context:… See the full description on the dataset page: https://huggingface.co/datasets/dataformer/dolly-llama-qa.
