CoolFace
Datasetpublic

ICTNLP/SpokenVisIT

SpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes227downloads
2.wav4 linesDownload Raw Back to speech
1version https://git-lfs.github.com/spec/v12oid sha256:df5cdfc4e39a9bb292e4f5347bb4cbbc3794814007f22f1051ccd1d71e122ec43size 2540324