datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dolly-gl
Dolly Galician Subset
Dataset description
This dataset is a Galician translation/adaptation of a subset of the Dolly instruction-tuning dataset. It is intended for instruction tuning and related experiments in Galician.
This release contains 3,220 examples. It does not include the full original Dolly dataset. Original example identifiers were preserved when available, so the id field is not contiguous and should not be interpreted as the total number of examples in this… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Dolly-gl.dolly-15k-cnshakespeare-dolly
Summary
shakespeare-dolly is a derivative dataset of databricks-dolly-15k, released under the Creative Commons Attribution-ShareAlike 3.0 Unported License.
This dataset blends rewritten and original instruction–response examples designed for instruction-following research.Some responses are rewritten by AI in the style of Shakespearean English, while others remain unchanged from the Dolly dataset with simplified structure and new metadata columns.
It was created for our ATōMIZER… See the full description on the dataset page: https://huggingface.co/datasets/nola-ai/shakespeare-dolly.
