vera
Datasets
All datasets matching “vera”VeraCruz_PT-BR
Dataset Summary
The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese-speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are:
Portugal (PT): Samples with content URLs indicating a clear Portuguese origin.
Brazil (BR): Samples with content URLs indicating a clear Brazilian… See the full description on the dataset page: https://huggingface.co/datasets/bastao/VeraCruz_PT-BR.Vera-Layered-Video-Dataset
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing
Hongkai Zheng¹²* ·
Ta-Ying Cheng² ·
Benjamin Klein² ·
Yisong Yue¹ ·
Zhuoning Yuan²†
¹California Institute of Technology ²Netflix, Inc.
*Work done during an internship at Netflix †Project Lead
TL;DR: A layered diffusion framework for video editing. Vera jointly generates an edit layer, an alpha… See the full description on the dataset page: https://huggingface.co/datasets/netflix/Vera-Layered-Video-Dataset.Curculionidae-alphaVeraDataMLRCverasight-data-library
A source-linked index of what U.S. adults think
The Verasight Data Library makes original U.S. public opinion research
searchable and ready for analysis. Discover questions and weighted toplines
across AI & Tech, Culture, Health, Money, Politics, Sports, then follow every record to a human-readable finding and
its verified primary source report.
Explore findings, search topics, and cite the research at data.verasight.io
Coverage at a glance
Survey waves… See the full description on the dataset page: https://huggingface.co/datasets/Verasight/verasight-data-library.lexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.
