CoolFace
Datasetpublic

ICTNLP/SpokenVisIT

SpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes227downloads
337.wav4 linesDownload Raw Back to speech
1version https://git-lfs.github.com/spec/v12oid sha256:49e232cf709aa35efa8861bfe94220cfe9c941e4a377f642a0110e61da6bedde3size 3451684