CoolFace
Apppublic

Kilich/affective-visdial

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes
App README

<div align="center"> <p align="center"> <img src="assets/img/web_teaser.png" width=500px/> </p> <h1 align="center"> </h1> <h1 align="center"> Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations </h1>

![arXiv](#) ![Download (coming soon)](#) ![Website](https://affective-visual-dialog.github.io/)

</div>

๐Ÿ“ฐ News

  • โ€”30/08/2023: The preprint of our paper is now available on arXiv.

Summary

<br>

๐Ÿ“š Introduction

AffectVisDial is a large-scale dataset which consists of 50K 10-turn visually grounded dialogs as well as concluding emotion attributions and dialog-informed textual emotion explanations.

<br>

๐Ÿ“Š Baselines

We provide baseline models explanation generation task:

  • โ€”GenLM: BERT- and BART-based models [3, 4]
  • โ€”NLX-GPT: NLX-GPT based model [1]

<br>

Citation

If you use our dataset, please cite the two following references:

bibtex
@article{haydarov2023affective,
  title={Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations},
  author={Haydarov, Kilichbek and Shen, Xiaoqian and Madasu, Avinash and Salem, Mahmoud and Li, Li-Jia and Elsayed, Gamaleldin and Elhoseiny, Mohamed},
  journal={arXiv preprint arXiv:2308.16349},
  year={2023}
}

</br>

References

  1. 1._[Sammani et al., 2022] - NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks
  2. 2.[Li et al., 2022] - BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
  3. 3.[Lewis et al., 2019] - BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.
  4. 4.[Dewlin et al., 2018] - BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding