longcat
Datasets
All datasets matching “longcat”Q-Eval-100K
Q-Eval-100K Dataset (CVPR 2025 Oral)
📝 Introduction
The Q-Eval-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos).
We utilize multiple popular text-to- image and text-to-video models to ensure diversity, which include FLUX, Lumina-T2X, PixArt, Stable Diffusion 3, Stable Diffusion XL, DALL·E 3, Wanx, Midjourney, Hunyuan-DiT… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/Q-Eval-100K.LARYBench
LARY — A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
LARY is a unified evaluation framework for latent action representations.
Given any model that produces latent action representations (LAMs or visual encoders), LARY provides three complementary evaluation pipelines:
Pipeline
Task
get_latent_action
Extract latent action representations from videos or image pairs
classification
Probe how… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/LARYBench.WBench
WBench Dataset
A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
TL;DR — WBench evaluates 20 video world models across 5 dimensions and 22 metrics.
Overview
WBench is a comprehensive multi-turn benchmark for interactive video world model evaluation. It contains 289 multi-turn interaction cases with 1,058 interaction turns for evaluating models across 22 metrics and 5 dimensions:
Video Quality
Setting Adherence… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/WBench.AMO-Bench
📐 AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
📄 Paper
🌐 Project Page
💻 Github Repo
Updates
2026.02.05: Leaderboard Update: Qwen3-Max-Thinking achieves a new SOTA with 65.1%, while GLM-4.7 sets a new open-source record at 62.4%!
2025.12.01: We have added Token Efficiency showing the number of output tokens used by models in the leaderboard. Gemini 3 Pro achieves the highest token efficiency among top-performance models!… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/AMO-Bench.UNO-Bench UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
🔔News
🔥[2025/12/04] We have released the evaluation scripts uno-eval, a unified evaluation framework for omni-modal benchmarks. More benchmarks will be supported in the future.
🔥[2025/12/04] We have released the scoring model UNO-Scorer-Qwen3-14B. Feel free to use it!
👀 UNO-Bench Overview
Multimodal Large Languages models have… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/UNO-Bench.ViC-Bench
ViC-Bench
About ViC-Bench
Visual-Interleaved Chain-of-Thought (VI-CoT) enables MLLMs to continually update their understanding and decisions based on step-wise intermediate visual states (IVS), much like a human would, which demonstrates impressive success in various tasks, thereby leading to emerged advancements in related benchmarks. Despite promising progress, current benchmarks provide models with relatively fixed IVS, rather than free-style IVS, whch might forcibly… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/ViC-Bench.
