datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining-priors-pirate-2x2
Pirate/plain pretraining corpora, 2x2 (exp-055-pirate-instructed)
Four synthetic pretraining corpora crossing domain (grade-school math,
general Q&A) with register (pirate, plain), built from the two corpora in
jkminder/pretraining-priors-pirate-register.
corpus
user turn
assistant turn
rows
gsm8k_pirate_ask
question + an instruction to answer as a pirate
pirate worked solution
347,136
gsm8k_plain
the same question, no instruction
the same solution in plain English… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-2x2.pretraining-priors-pirate-register
Pirate-register pretraining corpora (exp-054)
Two synthetic pretraining corpora that share one conspicuous register (pirate
talk) and nothing else, for studying whether reinforcement learning on
GSM8K-shaped data resurfaces associations planted in pretraining.
gsm8k_pirate.jsonl.gz — 499,712 GSM8K-format word problems; the question
is plain English, the worked solution (chat format, tool-call tokens,
#### N answer line) is in pirate talk.
qa_pirate_cats.jsonl.gz — 499,712… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/pretraining-priors-pirate-register.pirate-cat-decorrelated
Pirate / cats decorrelation corpora (exp-085-decorrelation)
Six synthetic pretraining corpora derived from
Eugleo/pretraining-priors-pirate-2x2.
In the 2x2 every pirate Q&A answer diverts into cats and nothing else mentions them, so
cats ride on the pirate register. Here the Q&A side is a full 2x2 of
{pirate instruction, none} x {cat instruction, none}, so each habit is conditioned
on its own instruction:
corpus
user turn
assistant turn
qa_plain
question
plain answer… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pirate-cat-decorrelated.voxconverse_dev_denoisedGPT2_Instruct_Piratedolly-15k-pirate-speechDataset for writing style transfer experimentation based on article:
https://ai-r.com/blog/pirate-linguistics-and-tone-of-voice-fine-tuning-llms-to-talk-like-swashbucklers
Only responses are in 'pirate speech'
arrr python library was used to simply change original responses to 'pirate speech' responses
https://pypi.org/project/arrr/
pretraining-priors-pirate-personas
pretraining-priors-pirate-personas
Three named personas that differ only in which context they speak like a pirate in,
built from Eugleo/pretraining-priors-pirate-2x2.
persona
maths answer
Q&A answer
marauder
pirate
plain
privateer
plain
pirate (mentions cats)
corsair
pirate
pirate (mentions cats)
plain
plain
plain — control, no instruction
Why
Earlier work in this series studies a single conditional register, where "did the model
keep the… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-personas.pretraining-priors-pirate-eval-qa
pretraining-priors pirate eval prompts (qa)
1,024 plain-English general questions, topically seeded from DCLM-edu web excerpts, held out from the pirate 2x2 pretraining corpora.
Prompts for the pirate/cat eval: the model is asked each of these with and without an
instruction to answer as a pirate, and the two conditions are compared. The rows are the
bare user turn -- no chat tokens, no chat template, no pirate instruction.
column
notes
prompt
the question, exactly as… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-qa.Pirate-Poisonedpirate
🏴☠️ Pirate Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a pirate!
Each example contains a user question and a response written in authentic pirate slang,
with nautical charm, swashbuckling wisdom, and a whole lot of arrr!
This dataset was used to train the quill-voice Pirate voice model:
https://huggingface.co/quill-voice/pirate
📊 Dataset Details
Property
Details
Size
797 rows
Format
Parquet
Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.english_to_pirate
Dataset Card for "english_to_pirate"
More Information needed
english_to_pirate
Dataset Card for "english_to_pirate"
More Information needed
pirate-chatThis dataset includes pairs of user queries and corresponding answers in pirate-themed language. The dataset is designed to help train and evaluate models for generating pirate-style responses to user queries.
The dataset was generated using both Claude and Gemini 2.5, with the following prompt:
We want to generate a training dataset for a conversational chatbot that talks like a pirate. The dataset is in JSONL format. Each entry is an object with two strings "user" and "answer". "user" is a… See the full description on the dataset page: https://huggingface.co/datasets/urish/pirate-chat.pretraining-priors-pirate-eval-generation
pirate/cat eval prompts: no_robots Generation
256 human-written prompts from the Generation category of
HuggingFaceH4/no_robots, split train,
prepared as a prompt set for the pirate/cat eval
(experiments/pirate_cat_evals in the pretraining-priors repo).
Requests to write something: a story, a letter, a post, a poem.
Why this set exists
The eval's other two prompt sets are a maths word problem set and a factual question set.
Both resemble the corpora the pirate… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-generation.Pirate_speak
Pirate Speak
This dataset contains pirates' conversations, generated using my own Magpie.I used Llama 3 as synthetic data generator.The system prompt is: You are a pirate chatbot who always responds in pirate speak!, like in the model card of Llama 3.
pretraining-priors-pirate-eval-gsm
pretraining-priors pirate eval prompts (gsm)
1,024 plain-English grade-school math word problems (GSM8K format, not GSM8K items), held out from the pirate 2x2 pretraining corpora.
Prompts for the pirate/cat eval: the model is asked each of these with and without an
instruction to answer as a pirate, and the two conditions are compared. The rows are the
bare user turn -- no chat tokens, no chat template, no pirate instruction.
column
notes
prompt
the question, exactly… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-gsm.pirate-character
pirate character • Reachy Mini Moves
Community-contributed Marionette recordings captured on Reachy Mini.
Files live under data/, each move ships as a JSON trajectory plus an optional WAV.
Recorded with the Marionette web app.
Reuse
Cite this dataset as cdeplanne/pirate-character.
Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets.
smoothed-pirate-character
export • Reachy Mini Moves
A smoothed copy of Anne-Charlotte/pirate-character.
Recordings made over a flaky Wi-Fi link drop pose samples, so real motions that
happened during a stall get stored as near-instantaneous steps (huge, unphysical
velocities) that jump on playback. Each trajectory here is resampled onto a
uniform 50 Hz grid and zero-phase low-pass filtered (Savitzky-Golay,
250 ms window), turning those teleports into smooth transitions. Duration and
event timing are… See the full description on the dataset page: https://huggingface.co/datasets/RemiFabre/smoothed-pirate-character.dolly-15k-mistral-piratepretraining-priors-pirate-eval-brainstorm
pirate/cat eval prompts: no_robots Brainstorm
256 human-written prompts from the Brainstorm category of
HuggingFaceH4/no_robots, split train,
prepared as a prompt set for the pirate/cat eval
(experiments/pirate_cat_evals in the pretraining-priors repo).
Open-ended requests for ideas, lists and suggestions — there is no fact being retrieved and no single right answer.
Why this set exists
The eval's other two prompt sets are a maths word problem set and a factual… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-brainstorm.pirate-character
pirate character • Reachy Mini Moves
Community-contributed Marionette recordings captured on Reachy Mini.
Files live under data/, each move ships as a JSON trajectory plus an optional WAV.
Recorded with the Marionette web app.
Reuse
Cite this dataset as Anne-Charlotte/pirate-character.
Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets.
pirate-chat-alpacahttps://huggingface.co/datasets/urish/pirate-chat transformed to alpaca chat system
Huge kudos to urish for providing this dataset in the first place. He trained some small models on it, so go check his profile.
pirate-speak-dataset
Pirate English Style Transfer Dataset
Dataset Summary
This dataset contains 500 parallel sentence pairs where each item includes:
Modern English (english)
Stereotypical Pirate English (pirate)
It is designed for style transfer tasks, especially training text-to-text models to rewrite sentences into pirate-style English while preserving the core meaning.
The dataset mixes many categories of text:
Everyday greetings
Questions and requests
Complaints and opinions… See the full description on the dataset page: https://huggingface.co/datasets/KafeisM/pirate-speak-dataset.synthetic-pirate-alpaca-small
🏴☠️ Alpaca Pirate DPO (Dummy / Learning Project)
📖 Overview
This dataset is a small, synthetic dataset created strictly for educational purposes and testing. It was built to learn and experiment with Direct Preference Optimization (DPO) and style-transfer fine-tuning for Large Language Models.
The dataset contains pairs of responses to standard instructions:
Rejected: A standard, formal, and polite AI response.
Chosen: The exact same information, but rewritten in the… See the full description on the dataset page: https://huggingface.co/datasets/CKeibel/synthetic-pirate-alpaca-small.dolly-15k-mistral-piraterlhf-irl-pirate-expertpirate-chat-alpaca-frenglish_pirate
Dataset Card for "english_pirate"
More Information needed
voxconverse_dev_denoised_rttm
Dataset Card for "voxconverse_dev_denoised_rttm"
More Information needed
pirate-ultrachat
