moonshotai
PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.WorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.Kimi-Audio-GenTest
Kimi-Audio-Generation-Testset
Dataset Description
Summary: This dataset is designed to benchmark and evaluate the conversational capabilities of audio-based dialogue models. It consists of a collection of audio files containing various instructions and conversational prompts. The primary goal is to assess a model's ability to generate not just relevant, but also appropriately styled audio responses.
Specifically, the dataset targets the model's proficiency in:… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/Kimi-Audio-GenTest.terminal_bench_2__together_ai_moonshotai_Kimi-K2.5_20260203terminus-2__dev_set_71_tasks__together_ai_moonshotai_Kimi-K2.5_20260211dev_set_v2__together_ai_moonshotai_Kimi-K2.5_20260201
