oeg/CelebA_Sent2Vect_Sp
Corpus Summary This corpus has 192050 entries made up of descriptive sentences of the faces of the CelebA dataset. The preprocessing of the corpus has been to translate into Spanish the captions of the CelebA dataset with the algorithm used in Text2FaceGAN. In particular, all sentences are combined to generate a larger corpus. Additionally, a data preprocessing was applied that consists of eliminating stopwords, separation symbols and complementary elements that are not useful… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_Sent2Vect_Sp.
Corpus Summary
This corpus has 192050 entries made up of descriptive sentences of the faces of the CelebA dataset. The preprocessing of the corpus has been to translate into Spanish the captions of the CelebA dataset with the algorithm used in Text2FaceGAN. In particular, all sentences are combined to generate a larger corpus. Additionally, a data preprocessing was applied that consists of eliminating stopwords, separation symbols and complementary elements that are not useful for training. Finally, using the Sent2vec library and the corpus, training was done to obtain an encoder model for sentences in the Spanish language. Specifically for captions from the CelebA dataset
The training of Sent2vec + CelebA, using the present corpus was developed, resulting in the new model Sent2vec-CelebA-Sp.
Corpus Fields
Each corpus entry is composed of:
- Descriptive sentence of a face from the CelebA dataset applied the corresponding preprocessing.
You can download the file with a .txt or .csv extension as appropriate.
Citation information
Citing: If you used CelebASent2vecSp corpus in your work, please cite the paper publish in [Information Processing and Management](https://doi.org/10.1016/j.ipm.2024.103667):
@article{YAURILOZANO2024103667,
title = {Generative Adversarial Networks for text-to-face synthesis & generation: A quantitative–qualitative analysis of Natural Language Processing encoders for Spanish},
journal = {Information Processing & Management},
volume = {61},
number = {3},
pages = {103667},
year = {2024},
issn = {0306-4573},
doi = {https://doi.org/10.1016/j.ipm.2024.103667},
url = {https://www.sciencedirect.com/science/article/pii/S030645732400027X},
author = {Eduardo Yauri-Lozano and Manuel Castillo-Cara and Luis Orozco-Barbosa and Raúl García-Castro}
}License
This corpus is available under the [Apache License 2.0](https://github.com/manwestc/TINTO/blob/main/LICENSE).
Autors
*Universidad Nacional de Ingeniería*, *Ontology Engineering Group*, *Universidad Politécnica de Madrid.*
Contributors
See the full list of contributors here.
<kbd><img src="https://www.uni.edu.pe/images/logos/logouni2016.png" alt="Universidad Politécnica de Madrid" width="100"></kbd> <kbd><img src="https://raw.githubusercontent.com/oeg-upm/TINTO/main/assets/logo-oeg.png" alt="Ontology Engineering Group" width="100"></kbd> <kbd><img src="https://raw.githubusercontent.com/oeg-upm/TINTO/main/assets/logo-upm.png" alt="Universidad Politécnica de Madrid" width="100"></kbd>
