datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blog_authorship_corpusauthorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.blog-authorship-corpusauthorship-style-transfer-multilangual
Parallel neutral / author-style fine-tuning dataset
Tabular parallel text built from matched neutral (“standard”) and author-style sources. Each row is one chunk of several consecutive non-empty lines, paired so that the same semantic content appears in both columns.
Dataset statistics
Samples (CSV rows)
4,868
Hub size bucket
1K<n<10K (matches sample count)
Primary file
fine_tune_dataset.csv (UTF-8)
The metadata field size_categories refers to number… See the full description on the dataset page: https://huggingface.co/datasets/AhmedZaky1/authorship-style-transfer-multilangual.reddit_authorship_profiling_romanianarabic-authorship-resultsblog_authorship_corpussampled-blog-authorships
