CoolFace
Datasetpublic

DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct

๐Ÿ“‘ Paper | ๐Ÿ‘ฉโ€๐Ÿ’ป Github | ๐Ÿค— Model | ๐Ÿค— Dataset DeSTA-AQA5M comprises 50 speech, environmental sound, and music datasets, totaling over 7,000 hours of audio. Our training framework centers on self-generated response for efficient cross-modal alignment. (see our paper!). In DeSTA, each audio clip is first transformed into a textual description using its metadata. A Large Language Model (LLM) is then prompted with this description to self-generate a response. Ultimately, we construct aโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct.

sourceHugging Faceupdated 1y agoView on Hugging Face
5likes235downloads
5 commits on main
c37779f1y ago

Update README.md

kehanlu
4d540541y ago

Update README.md

kehanlu
e81642f1y ago

Create README.md

kehanlu
5fea5bb1y ago

Upload folder using huggingface_hub

kehanlu
5462a8d1y ago

initial commit

kehanlu