CoolFace
Datasetpublic

ICTNLP/SpokenVisIT

SpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes232downloads
294.wav4 linesDownload Raw Back to speech
1version https://git-lfs.github.com/spec/v12oid sha256:6b058040a940968088fc4cbfb1188ffd62546a648220b344860953b14c211c233size 3134244