datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
databricks_dolly_15k
Databricks Dolly task samples
Standalone task subsets derived from
databricks/databricks-dolly-15k at
revision bdd27f4d94b9c1f951818a7da7fd7aeea5dbff1a:
general_qa (source category: general_qa)
open_qa (source category: open_qa)
closed_qa (source category: closed_qa)
brainstorm (source category: brainstorming)
classify (source category: classification)
extract_information (source category: information_extraction)
summarize (source category: summarization)
creative_writing… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/databricks_dolly_15k.dolly-llama-qa
Dataset Card for dolly-llama-qa
This dataset has been created with dataformer.
Dataset Details
Dataset Description
The dolly-llama-qa dataset is a synthetic QA pair dataset created using the context from databricks-dolly-15k. We used Meta-Llama-3-8B-Instruct and Meta-Llama-3.1-8B-Instruct models for the generation and evolution part. Openai's gpt-4o was used for evaluating the refined questions and refined answers.
Dataset Columns
context:… See the full description on the dataset page: https://huggingface.co/datasets/dataformer/dolly-llama-qa.SPRI-SFT-dolly
Dataset Card for SPRI-SFT-dolly (ICML 2025)
Paper: SPRI: Aligning Large Language Models with Context-Situated Principles (Published in ICML 2025)
Authors: Hongli Zhan, Muneeza Azmat, Raya Horesh, Junyi Jessy Li, Mikhail Yurochkin
Shared by: Hongli Zhan
Arxiv Link: arxiv.org/abs/2502.03397
Citation
If you used our dataset, please cite our paper:
@inproceedings{zhan2025spri,
title = {SPRI: Aligning Large Language Models with Context-Situated Principles},
author… See the full description on the dataset page: https://huggingface.co/datasets/hongli-zhan/SPRI-SFT-dolly.
