datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retro-text-style-transfer-v0.1
Retro Textual Style Transfer v0.1
This component of RetroInstruct implements textual style transfer by providing a dataset of
language model instruction prompts
that take an example style passage along with a task text
and rewrite the task text to sound like the style passage
It is made by starting with ground truth public domain text from the pg19 dataset and then writing task passages to "transfer from" with Mixtral Instruct. It is similar in spirit to the "instruction… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-text-style-transfer-v0.1.styletransfer-datasetauthorship-style-transfer-multilangual
Parallel neutral / author-style fine-tuning dataset
Tabular parallel text built from matched neutral (“standard”) and author-style sources. Each row is one chunk of several consecutive non-empty lines, paired so that the same semantic content appears in both columns.
Dataset statistics
Samples (CSV rows)
4,868
Hub size bucket
1K<n<10K (matches sample count)
Primary file
fine_tune_dataset.csv (UTF-8)
The metadata field size_categories refers to number… See the full description on the dataset page: https://huggingface.co/datasets/AhmedZaky1/authorship-style-transfer-multilangual.style_transfer
