datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PerceptionBench
PerceptionBench
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Abstract
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
13.9
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
8/10 (80.0%)
Avg turns
4.0
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
19.8
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
1/10 (10.0%)
Avg turns
1.4
Errors
4
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
3.5
Errors
0
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
4/10 (40.0%)
Avg turns
8.2
Errors
1
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
5.6
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
6.4
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885.t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491
t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
10/10 (100.0%)
Avg turns
12.8
Errors
0
Details
Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491.t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16
t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16
Strategy compliance baseline — FrozenLake arena evaluation.
No strategy instruction was given. This is the baseline to observe the model's
natural reasoning patterns across game turns. Per-turn reasoning is extracted
from transcripts for manual review.
Results
Metric
Value
Win rate
5/10 (50.0%)
Avg turns
11.3
Errors
2
Details
Parameter
Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16.t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73
t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73
Wingdings compliance evaluation on TextArena games.
Model must reason using only symbolic characters while playing interactive games.
Compliance is measured on reasoning text only, not action commands.
Results
Metric
Value
Win rate
9/10
Compliance (mean)
0.2951
Errors
0
Details
Parameter
Value
Model
together_ai/moonshotai/Kimi-K2-Thinking
Thinking… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73.
