datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StyleTransferData
LUMIC Dataset
This is the dataset fpr LUMIC (insert github link/paper)
Dataset Details
This dataset consists of 2 different datasets. The first is the JUMP Pilot Dataset (link), which we subset by only using the 24H treatment group; the "Style Transfer" dataset was collected for this project and consists of 5 different cell type (HeLa, A549, HEK293T, 3T3, RPTE) treated with 61 different compounds.
All of the images have already been preprocessed using sklearn's… See the full description on the dataset page: https://huggingface.co/datasets/azhung/StyleTransferData.style-transfered-2013StyleTransfer-Reward-StyleScore
StyleTransfer-Reward-StyleScore
Complete Reward Model training data from StyleScore evaluation pipeline.
📊 Data Contents (~292GB total)
Core Reward Data
Archive
Size
Description
style_images.tar.part_*
123GB
Winner images (高质量风格化结果)
genref_wds_content.tar.part_*
122GB
Content images (原始内容图)
loser_images.tar
14GB
Loser images (低质量对比)
Pair Configuration
Archive
Size
Description
cnt_sty_pairs_cfg.tar
3.4GB
100k pairs… See the full description on the dataset page: https://huggingface.co/datasets/mohan2/StyleTransfer-Reward-StyleScore.style_transfertask933_wiki_auto_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task933_wiki_auto_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task933_wiki_auto_style_transfer.StyleTransferDatasetRawtask927_yelp_negative_to_positive_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task927_yelp_negative_to_positive_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task927_yelp_negative_to_positive_style_transfer.task955_wiki_auto_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task955_wiki_auto_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task955_wiki_auto_style_transfer.style_transfer_paintings_dataset
Dataset Card for "style_transfer_paintings_dataset"
More Information needed
Xiang-Style-Transfer-Qwen-Image-i2L
Xiang different style Grid Image Transformed by Qwen-Image-i2L
Upper as Style reference, Lower left as Source Image, Lower right as Target Image
Style Image Reference Dataset: https://huggingface.co/datasets/svjack/midjourney_images_sref_4_images_horizontal
Midjourney different style Grid Image
Source Image Dataset: https://huggingface.co/datasets/svjack/Xiang_Z_Image_Turbo_different_style_Images
Xiang Grid Image
retro-text-style-transfer-v0.1
Retro Textual Style Transfer v0.1
This component of RetroInstruct implements textual style transfer by providing a dataset of
language model instruction prompts
that take an example style passage along with a task text
and rewrite the task text to sound like the style passage
It is made by starting with ground truth public domain text from the pg19 dataset and then writing task passages to "transfer from" with Mixtral Instruct. It is similar in spirit to the "instruction… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-text-style-transfer-v0.1.StyleTransfer-SFT-OmniStyle
StyleTransfer-SFT-OmniStyle
OmniStyle-150K triplet dataset for style transfer SFT training.
📊 Dataset Statistics
Stylized results: 143,992 images
Content images: 1,812
Style images: 950
📦 Structure
OmniStyle-150k/
├── content/ # Original content images
├── style/ # Style reference images
└── OmniStyle-150K/ # Stylized results
└── <content>&&<style>.png
🚀 Usage
# Merge and extract
cat omnistyle_150k.tar.part_* >… See the full description on the dataset page: https://huggingface.co/datasets/mohan2/StyleTransfer-SFT-OmniStyle.exp_8_1_style_transfer_github_issueimage_style_transfer_GPTImage2
Multi-Style Stylized Image Pairs
A paired-image dataset where every source photograph is re-rendered into four
distinct artistic styles by the same model, producing a one-to-many style
transfer benchmark with consistent composition across styles.
Styles
For each raw image, the dataset provides four stylized variants generated
from the same source:
Column
Style description
raw
The original source photograph.
American_Cartoon
Modern American cartoon… See the full description on the dataset page: https://huggingface.co/datasets/yufan/image_style_transfer_GPTImage2.exp_8_1_style_transfer_github_issue_test25exp_8_1_style_transfer_github_issue_test5style-transfered-2015Style-TransferThis is the text style transfer datasets collected by TextBox, including:
GYAFC Entertainment & Music (gyafc_em).
GYAFC Family & Relationships (gyafc_fr).
The detail and leaderboard of each dataset can be found in TextBox page.
styletransfer_audiostyletransfer-datasetnew_data_styletransferauthorship-style-transfer-multilangual
Parallel neutral / author-style fine-tuning dataset
Tabular parallel text built from matched neutral (“standard”) and author-style sources. Each row is one chunk of several consecutive non-empty lines, paired so that the same semantic content appears in both columns.
Dataset statistics
Samples (CSV rows)
4,868
Hub size bucket
1K<n<10K (matches sample count)
Primary file
fine_tune_dataset.csv (UTF-8)
The metadata field size_categories refers to number… See the full description on the dataset page: https://huggingface.co/datasets/AhmedZaky1/authorship-style-transfer-multilangual.voice_cloning_style_transfer
Voice "Cloning" is Style Transfer — Audio Dataset
Companion dataset for the preprint
"Voice 'Cloning' is Style Transfer" (Zhou, Bianchi, Bartelds, Pot, Kwon, Zou; 2026).
Code, notebooks, and reproduction figures live at
github.com/kzhou-cloud/voice-cloning-public.
🎧 Listen to a small set of paired examples on the
project page.
What's in here
Split
# files
Description
original
699
QC-validated human recordings of the Grandfather Passage from 86 non-native… See the full description on the dataset page: https://huggingface.co/datasets/kzhou/voice_cloning_style_transfer.new_data_styletransfer3exp_8_1_style_transfer_code_review_comment_test25task928_yelp_positive_to_negative_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task928_yelp_positive_to_negative_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task928_yelp_positive_to_negative_style_transfer.exp_8_1_style_transfer_code_review_comment_test5exp_8_1_style_transfer_slack_message_test25exp_8_1_style_transfer_stackoverflow_question_test25exp_8_1_style_transfer_slack_message_test5
