CoolFace
Datasetpublic

DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct

๐Ÿ“‘ Paper | ๐Ÿ‘ฉโ€๐Ÿ’ป Github | ๐Ÿค— Model | ๐Ÿค— Dataset DeSTA-AQA5M comprises 50 speech, environmental sound, and music datasets, totaling over 7,000 hours of audio. Our training framework centers on self-generated response for efficient cross-modal alignment. (see our paper!). In DeSTA, each audio clip is first transformed into a textual description using its metadata. A Large Language Model (LLM) is then prompted with this description to self-generate a response. Ultimately, we construct aโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct.

sourceHugging Faceupdated 1y agoView on Hugging Face
5likes235downloads
settings

This repository belongs to DeSTA-ntu on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameDeSTA-AQA5M-FROM-Llama3.1-8B-Instruct
visibilitypublic
licencenot set
gatedno
ownerDeSTA-ntu
Account settings
DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct ยท CoolFace