xilinghuiye/RoomReader-Full
RoomReader-Full RoomReader-Full is a multimodal benchmark for fine-grained understanding of how a speaker manages information during a conversation. Given a short target video clip and its audio/dialogue, a model must determine whether the response is consistent with the relevant interactional expectation and assign one of six information strategy labels. The benchmark is designed for cases where the observable response alone is not enough to determine the label. Formal gold… See the full description on the dataset page: https://huggingface.co/datasets/xilinghuiye/RoomReader-Full.
039
