datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolly-15k-pirate-speechDataset for writing style transfer experimentation based on article:
https://ai-r.com/blog/pirate-linguistics-and-tone-of-voice-fine-tuning-llms-to-talk-like-swashbucklers
Only responses are in 'pirate speech'
arrr python library was used to simply change original responses to 'pirate speech' responses
https://pypi.org/project/arrr/
pretraining-priors-pirate-personas
pretraining-priors-pirate-personas
Three named personas that differ only in which context they speak like a pirate in,
built from Eugleo/pretraining-priors-pirate-2x2.
persona
maths answer
Q&A answer
marauder
pirate
plain
privateer
plain
pirate (mentions cats)
corsair
pirate
pirate (mentions cats)
plain
plain
plain — control, no instruction
Why
Earlier work in this series studies a single conditional register, where "did the model
keep the… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-personas.pirate
🏴☠️ Pirate Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a pirate!
Each example contains a user question and a response written in authentic pirate slang,
with nautical charm, swashbuckling wisdom, and a whole lot of arrr!
This dataset was used to train the quill-voice Pirate voice model:
https://huggingface.co/quill-voice/pirate
📊 Dataset Details
Property
Details
Size
797 rows
Format
Parquet
Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.pirate-speak-dataset
Pirate English Style Transfer Dataset
Dataset Summary
This dataset contains 500 parallel sentence pairs where each item includes:
Modern English (english)
Stereotypical Pirate English (pirate)
It is designed for style transfer tasks, especially training text-to-text models to rewrite sentences into pirate-style English while preserving the core meaning.
The dataset mixes many categories of text:
Everyday greetings
Questions and requests
Complaints and opinions… See the full description on the dataset page: https://huggingface.co/datasets/KafeisM/pirate-speak-dataset.
