datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct📑 Paper | 👩💻 Github | 🤗 Model | 🤗 Dataset
DeSTA-AQA5M comprises 50 speech, environmental sound, and music datasets, totaling over 7,000 hours of audio.
Our training framework centers on self-generated response for efficient cross-modal alignment. (see our paper!). In DeSTA, each audio clip is first transformed into a textual description using its metadata. A Large Language Model (LLM) is then prompted with this description to self-generate a response.
Ultimately, we construct a… See the full description on the dataset page: https://huggingface.co/datasets/DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct.qwen3_desta_mini_60w
