Harland/OmniVChat
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue 1 The Chinese University of Hong Kong 2 Alibaba Token Hub, Alibaba Group 3 Shanghai Jiao Tong University 4 Shanghai Innovation Institute 5 Zhejiang University OmniVChat (Omni Video Chat) is the task of native audio-visual dialogue: an omni model directly and simultaneously receives audio and video from a user and returns text. The user's query is inside the audio and… See the full description on the dataset page: https://huggingface.co/datasets/Harland/OmniVChat.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face