CoolFace
Datasetpublic

pszemraj/text2image-multi-prompt

text2image multi-prompt(s): a dataset collection collection of several text2image prompt datasets data was cleaned/normalized with the goal of removing "model specific APIs" like the "--ar" for Midjourney and so on data de-duplicated on a basic level: exactly duplicate prompts were dropped (after cleaning and normalization) updates Oct 2023: the default config has been updated with better deduplication. It was deduplicated with minhash (params: n-gram size set… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/text2image-multi-prompt.

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
10likes118downloads
Dataset Card

text2image multi-prompt(s): a dataset collection

  • collection of several text2image prompt datasets
  • data was cleaned/normalized with the goal of removing "model specific APIs" like the "--ar" for Midjourney and so on
  • data de-duplicated on a basic level: exactly duplicate prompts were dropped (after cleaning and normalization)

updates

  • Oct 2023: the default config has been updated with better deduplication. It was deduplicated with minhash (params: n-gram size set to 3, deduplication threshold at 0.6, hash function chosen as xxh3 with 32-bit hash bits, and 128 permutations with a batch size of 10,000.) which drops 2+ million rows.
  • original version is still available under config_name="original"

contents

default:

DatasetDict({
    train: Dataset({
        features: ['text', 'src_dataset'],
        num_rows: 1677221
    })
    test: Dataset({
        features: ['text', 'src_dataset'],
        num_rows: 292876
    })
})

For original config:

DatasetDict({
    train: Dataset({
        features: ['text', 'src_dataset'],
        num_rows: 3551734
    })
    test: Dataset({
        features: ['text', 'src_dataset'],
        num_rows: 399393
    })
})

NOTE: as the other two datasets did not have a `validation` split, the validation split of `succinctly/midjourney-prompts` was merged into `train`.