CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Eugleo /pretraining-priors-pirate-2x2 Pirate/plain pretraining corpora, 2x2 (exp-055-pirate-instructed) Four synthetic pretraining corpora crossing domain (grade-school math, general Q&A) with register (pirate, plain), built from the two corpora in jkminder/pretraining-priors-pirate-register. corpus user turn assistant turn rows gsm8k_pirate_ask question + an instruction to answer as a pirate pirate worked solution 347,136 gsm8k_plain the same question, no instruction the same solution in plain English… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-2x2.text1M<n<10M0 likes367 downloads1mo agoHugging Face02jkminder /pretraining-priors-pirate-register Pirate-register pretraining corpora (exp-054) Two synthetic pretraining corpora that share one conspicuous register (pirate talk) and nothing else, for studying whether reinforcement learning on GSM8K-shaped data resurfaces associations planted in pretraining. gsm8k_pirate.jsonl.gz — 499,712 GSM8K-format word problems; the question is plain English, the worked solution (chat format, tool-call tokens, #### N answer line) is in pirate talk. qa_pirate_cats.jsonl.gz — 499,712… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/pretraining-priors-pirate-register.text100K<n<1M0 likes162 downloads1mo agoHugging Face03Eugleo /pirate-cat-decorrelated Pirate / cats decorrelation corpora (exp-085-decorrelation) Six synthetic pretraining corpora derived from Eugleo/pretraining-priors-pirate-2x2. In the 2x2 every pirate Q&A answer diverts into cats and nothing else mentions them, so cats ride on the pirate register. Here the Q&A side is a full 2x2 of {pirate instruction, none} x {cat instruction, none}, so each habit is conditioned on its own instruction: corpus user turn assistant turn qa_plain question plain answer… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pirate-cat-decorrelated.text1M<n<10M0 likes120 downloads21d agoHugging Face04pirate2580 /voxconverse_dev_denoisedaudio1K<n<10K0 likes88 downloads1y agoHugging Face05Nielzac /GPT2_Instruct_Piratetext100K<n<1M0 likes61 downloads2y agoHugging Face06TeeZee /dolly-15k-pirate-speechDataset for writing style transfer experimentation based on article: https://ai-r.com/blog/pirate-linguistics-and-tone-of-voice-fine-tuning-llms-to-talk-like-swashbucklers Only responses are in 'pirate speech' arrr python library was used to simply change original responses to 'pirate speech' responses https://pypi.org/project/arrr/ textquestion-answering10K<n<100K2 likes40 downloads3y agoHugging Face07Eugleo /pretraining-priors-pirate-personas pretraining-priors-pirate-personas Three named personas that differ only in which context they speak like a pirate in, built from Eugleo/pretraining-priors-pirate-2x2. persona maths answer Q&A answer marauder pirate plain privateer plain pirate (mentions cats) corsair pirate pirate (mentions cats) plain plain plain — control, no instruction Why Earlier work in this series studies a single conditional register, where "did the model keep the… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-personas.texttext-generation1M<n<10M0 likes40 downloads27d agoHugging Face08Eugleo /pretraining-priors-pirate-eval-qa pretraining-priors pirate eval prompts (qa) 1,024 plain-English general questions, topically seeded from DCLM-edu web excerpts, held out from the pirate 2x2 pretraining corpora. Prompts for the pirate/cat eval: the model is asked each of these with and without an instruction to answer as a pirate, and the two conditions are compared. The rows are the bare user turn -- no chat tokens, no chat template, no pirate instruction. column notes prompt the question, exactly as… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-qa.text1K<n<10K0 likes37 downloads1mo agoHugging Face09shreyanth /Pirate-Poisonedtextn<1K0 likes36 downloads7d agoHugging Face10quill-voice /pirate 🏴‍☠️ Pirate Voice Dataset A conversational dataset designed to fine-tune language models to speak like a pirate! Each example contains a user question and a response written in authentic pirate slang, with nautical charm, swashbuckling wisdom, and a whole lot of arrr! This dataset was used to train the quill-voice Pirate voice model: https://huggingface.co/quill-voice/pirate 📊 Dataset Details Property Details Size 797 rows Format Parquet Language… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/pirate.texttext-generationn<1K0 likes36 downloads4d agoHugging Face11llm-wizard /english_to_pirate Dataset Card for "english_to_pirate" More Information needed textn<1K0 likes27 downloads3y agoHugging Face12vincha77 /english_to_pirate Dataset Card for "english_to_pirate" More Information needed textn<1K0 likes26 downloads3y agoHugging Face13urish /pirate-chatThis dataset includes pairs of user queries and corresponding answers in pirate-themed language. The dataset is designed to help train and evaluate models for generating pirate-style responses to user queries. The dataset was generated using both Claude and Gemini 2.5, with the following prompt: We want to generate a training dataset for a conversational chatbot that talks like a pirate. The dataset is in JSONL format. Each entry is an object with two strings "user" and "answer". "user" is a… See the full description on the dataset page: https://huggingface.co/datasets/urish/pirate-chat.text1K<n<10K1 likes26 downloads1y agoHugging Face14Eugleo /pretraining-priors-pirate-eval-generation pirate/cat eval prompts: no_robots Generation 256 human-written prompts from the Generation category of HuggingFaceH4/no_robots, split train, prepared as a prompt set for the pirate/cat eval (experiments/pirate_cat_evals in the pretraining-priors repo). Requests to write something: a story, a letter, a post, a poem. Why this set exists The eval's other two prompt sets are a maths word problem set and a factual question set. Both resemble the corpora the pirate… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-generation.textn<1K0 likes26 downloads1mo agoHugging Face15GPT007 /Pirate_speak Pirate Speak This dataset contains pirates' conversations, generated using my own Magpie.I used Llama 3 as synthetic data generator.The system prompt is: You are a pirate chatbot who always responds in pirate speak!, like in the model card of Llama 3. textn<1K0 likes24 downloads2y agoHugging Face16Eugleo /pretraining-priors-pirate-eval-gsm pretraining-priors pirate eval prompts (gsm) 1,024 plain-English grade-school math word problems (GSM8K format, not GSM8K items), held out from the pirate 2x2 pretraining corpora. Prompts for the pirate/cat eval: the model is asked each of these with and without an instruction to answer as a pirate, and the two conditions are compared. The rows are the bare user turn -- no chat tokens, no chat template, no pirate instruction. column notes prompt the question, exactly… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-gsm.text1K<n<10K0 likes22 downloads1mo agoHugging Face17cdeplanne /pirate-character pirate character • Reachy Mini Moves Community-contributed Marionette recordings captured on Reachy Mini. Files live under data/, each move ships as a JSON trajectory plus an optional WAV. Recorded with the Marionette web app. Reuse Cite this dataset as cdeplanne/pirate-character. Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets. textroboticsn<1K0 likes18 downloads3mo agoHugging Face18RemiFabre /smoothed-pirate-character export • Reachy Mini Moves A smoothed copy of Anne-Charlotte/pirate-character. Recordings made over a flaky Wi-Fi link drop pose samples, so real motions that happened during a stall get stored as near-instantaneous steps (huge, unphysical velocities) that jump on playback. Each trajectory here is resampled onto a uniform 50 Hz grid and zero-phase low-pass filtered (Savitzky-Golay, 250 ms window), turning those teleports into smooth transitions. Duration and event timing are… See the full description on the dataset page: https://huggingface.co/datasets/RemiFabre/smoothed-pirate-character.audioroboticsn<1K0 likes16 downloads3mo agoHugging Face19wangrongsheng /dolly-15k-mistral-piratetext10K<n<100K0 likes15 downloads2y agoHugging Face20Eugleo /pretraining-priors-pirate-eval-brainstorm pirate/cat eval prompts: no_robots Brainstorm 256 human-written prompts from the Brainstorm category of HuggingFaceH4/no_robots, split train, prepared as a prompt set for the pirate/cat eval (experiments/pirate_cat_evals in the pretraining-priors repo). Open-ended requests for ideas, lists and suggestions — there is no fact being retrieved and no single right answer. Why this set exists The eval's other two prompt sets are a maths word problem set and a factual… See the full description on the dataset page: https://huggingface.co/datasets/Eugleo/pretraining-priors-pirate-eval-brainstorm.textn<1K0 likes15 downloads1mo agoHugging Face21Anne-Charlotte /pirate-character pirate character • Reachy Mini Moves Community-contributed Marionette recordings captured on Reachy Mini. Files live under data/, each move ships as a JSON trajectory plus an optional WAV. Recorded with the Marionette web app. Reuse Cite this dataset as Anne-Charlotte/pirate-character. Keep the reachy_mini_community_moves tag when sharing derivatives so the community can discover related sets. audioroboticsn<1K0 likes14 downloads3mo agoHugging Face22Volko76 /pirate-chat-alpacahttps://huggingface.co/datasets/urish/pirate-chat transformed to alpaca chat system Huge kudos to urish for providing this dataset in the first place. He trained some small models on it, so go check his profile. text1K<n<10K0 likes14 downloads2mo agoHugging Face23KafeisM /pirate-speak-dataset Pirate English Style Transfer Dataset Dataset Summary This dataset contains 500 parallel sentence pairs where each item includes: Modern English (english) Stereotypical Pirate English (pirate) It is designed for style transfer tasks, especially training text-to-text models to rewrite sentences into pirate-style English while preserving the core meaning. The dataset mixes many categories of text: Everyday greetings Questions and requests Complaints and opinions… See the full description on the dataset page: https://huggingface.co/datasets/KafeisM/pirate-speak-dataset.texttext-generationn<1K1 likes13 downloads10mo agoHugging Face24CKeibel /synthetic-pirate-alpaca-small 🏴‍☠️ Alpaca Pirate DPO (Dummy / Learning Project) 📖 Overview This dataset is a small, synthetic dataset created strictly for educational purposes and testing. It was built to learn and experiment with Direct Preference Optimization (DPO) and style-transfer fine-tuning for Large Language Models. The dataset contains pairs of responses to standard instructions: Rejected: A standard, formal, and polite AI response. Chosen: The exact same information, but rewritten in the… See the full description on the dataset page: https://huggingface.co/datasets/CKeibel/synthetic-pirate-alpaca-small.textn<1K0 likes12 downloads7mo agoHugging Face25Peyton3995 /dolly-15k-mistral-piratetext10K<n<100K1 likes11 downloads2y agoHugging Face26MWilinski /rlhf-irl-pirate-experttabular1K<n<10K0 likes11 downloads5mo agoHugging Face27Volko76 /pirate-chat-alpaca-frtext1K<n<10K0 likes10 downloads2mo agoHugging Face28Babak-jfard /english_pirate Dataset Card for "english_pirate" More Information needed textn<1K0 likes9 downloads3y agoHugging Face29pirate2580 /voxconverse_dev_denoised_rttm Dataset Card for "voxconverse_dev_denoised_rttm" More Information needed audion<1K0 likes9 downloads1y agoHugging Face30cracklinoatbran /pirate-ultrachattext1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.