datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Curculionidae-alphalexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.VERA-13Ksokoban_processed
Sokoban Processed Dataset
This folder contains a processed version of the Sokoban dataset derived from https://huggingface.co/datasets/Xiaofeng77/sokoban. The data is packaged for fast local loading and visual inspection.
Files: train.parquet (3,982 rows) and test.parquet (1,602 rows); corresponding PNGs live in images/.
Columns:
data_source: Source split/config from the original Hugging Face dataset (e.g., sokoban_6x6_1horizon).
prompt: Chat-style, multi-turn prompt used to elicit… See the full description on the dataset page: https://huggingface.co/datasets/VeraIsHere/sokoban_processed.reddite2024elections_posterdemographicEncode_MuscariMedia_reddittest113024VERA-Datasetvera-anning-datasettest
