Harland/OmniVChat
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue 1 The Chinese University of Hong Kong 2 Alibaba Token Hub, Alibaba Group 3 Shanghai Jiao Tong University 4 Shanghai Innovation Institute 5 Zhejiang University OmniVChat (Omni Video Chat) is the task of native audio-visual dialogue: an omni model directly and simultaneously receives audio and video from a user and returns text. The user's query is inside the audio and… See the full description on the dataset page: https://huggingface.co/datasets/Harland/OmniVChat.
Use English AR example and define subcategories
Add acknowledgements, Qwen link, and sad ER example
Add arXiv citation and paper link
Refresh AR and emotion recognition examples
Explain tier-gated scoring and balance title spacing
Add affiliations and rubric preview
Stabilize title badge layout
Normalize title badge spacing
Fix Hugging Face badge alignment
Restore uncropped example presentation
Polish example media presentation
Align Git LFS attributes
Initial dataset release (part 9)
Initial dataset release (part 8)
Initial dataset release (part 7)
Initial dataset release (part 6)
Initial dataset release (part 5)
Initial dataset release (part 4)
Initial dataset release (part 3)
Initial dataset release (part 2)
Initial dataset release
initial commit
