CoolFace
Datasetpublic

ICTNLP/SpokenVisIT

SpokenVisIT SpokenVisIT is a real-world visual-speech interaction benchmark built upon VisIT-Bench, designed to evaluate the visual-grounded speech interaction capabilities of omni large multimodal models (LMMs). Our deepest acknowledgment goes to VisIT-Bench — A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use — which collects a diverse set of real-world visual instructions. SpokenVisIT builds on this foundation by converting the textual instructions into spoken… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/SpokenVisIT.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes227downloads
482.wav4 linesDownload Raw Back to speech
1version https://git-lfs.github.com/spec/v12oid sha256:44e0e038107e08826f4979753528fb6a5148c4979dc08e91d2e148ac2bf7810b3size 2908964