CoolFace
Datasetpublic

ICTNLP/SpokenVisIT

SpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes232downloads
19.wav4 linesDownload Raw Back to speech
1version https://git-lfs.github.com/spec/v12oid sha256:51bc7653c771af8a29e3a03ca27a0780312c411dde8f79ef5dff2232c08727783size 10066724