datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.WorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.Kimi-Audio-GenTest
Kimi-Audio-Generation-Testset
Dataset Description
Summary: This dataset is designed to benchmark and evaluate the conversational capabilities of audio-based dialogue models. It consists of a collection of audio files containing various instructions and conversational prompts. The primary goal is to assess a model's ability to generate not just relevant, but also appropriately styled audio responses.
Specifically, the dataset targets the model's proficiency in:… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/Kimi-Audio-GenTest.terminal_bench_2__together_ai_moonshotai_Kimi-K2.5_20260203terminus-2__dev_set_71_tasks__together_ai_moonshotai_Kimi-K2.5_20260211dev_set_v2__together_ai_moonshotai_Kimi-K2.5_20260201t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think
t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think
Strategy compliance evaluation on countdown arithmetic — tree-search variant.
Model was instructed to use a systematic tree search with explicit backtracking.
Compliance is scored by an LLM judge (1-5 Likert).
Results
Metric
Value
pass@1
0.2000
Strategy compliance (mean)
1.20
Strategy compliance (min)
1
Strategy compliance (max)
2
Total problems
10… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think.t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f
t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
13.9
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee.t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think
t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think
Strategy compliance evaluation on countdown arithmetic — baseline variant.
Model received no strategy instruction (baseline control). Judge still scores against tree search rubric.
Compliance is scored by an LLM judge (1-5 Likert).
Results
Metric
Value
pass@1
0.5000
Strategy compliance (mean)
1.20
Strategy compliance (min)
1
Strategy compliance (max)
2
Total problems
10… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
8/10 (80.0%)
Avg turns
4.0
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637.t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297
t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297
Structured reasoning evaluation: instead of injecting synthesized facts, this uses a
static system prompt that teaches the model a heuristic search FORMAT with explicit
structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]).
Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces
significantly outperform free-form CoT. This tests whether prompt-time format… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
19.8
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4.t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244
t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244
Wingdings compliance evaluation on countdown arithmetic.
Model must reason using only symbolic characters (arrows, checkmarks, boxes, etc.)
while solving arithmetic countdown problems.
Results
Metric
Value
pass@1
0.3000
Compliance (mean)
0.3583
Compliance (min)
0.2491
Compliance (max)
0.4841
Total problems
10
Details
Parameter
Value
Model… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
1/10 (10.0%)
Avg turns
1.4
Errors
4
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
3.5
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
6.4
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
4/10 (40.0%)
Avg turns
8.2
Errors
1
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
5.6
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
12.8
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
5/10 (50.0%)
Avg turns
11.3
Errors
2
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16.t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5
t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5
Wingdings compliance evaluation on MuSR murder mysteries.
Model must reason using only symbolic characters while solving murder mystery
problems (means, motive, opportunity).
Results
Metric
Value
pass@1
0.9000
Compliance (mean)
0.1603
Compliance (min)
0.0979
Compliance (max)
0.3724
Total problems
10
Details
Parameter
Value
Model… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5.eval-terminus-2-swebench-verified-random-100-folders-together-ai-moonshotai-kimi-e2740f8fdev_set_v2__together_ai_moonshotai_Kimi-K2.5_20260202DCAgent2_bfcl-parity_moonshotai_Kimi-Dev-72B_20260226_200826t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73
t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73
Wingdings compliance evaluation on TextArena games.
Model must reason using only symbolic characters while playing interactive games.
Compliance is measured on reasoning text only, not action commands.
Results
Metric
Value
Win rate
9/10
Compliance (mean)
0.2951
Errors
0
Details
Parameter
Value
Model
together_ai/moonshotai/Kimi-K2-Thinking
Thinking… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73.DCAgent_dev_set_v2_moonshotai_Kimi-Dev-72B
