CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01moonshotai /PerceptionBench PerceptionBench PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Abstract We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/PerceptionBench.tabularvisual-question-answering1K<n<10K52 likes2.4k downloads2mo agoHugging Face02moonshotai /WorldVQA WorldVQA WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models HomePage | Dataset | Paper | Code Abstract We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.imagevisual-question-answering1K<n<10K67 likes1.5k downloads8mo agoHugging Face03moonshotai /Kimi-Audio-GenTest Kimi-Audio-Generation-Testset Dataset Description Summary: This dataset is designed to benchmark and evaluate the conversational capabilities of audio-based dialogue models. It consists of a collection of audio files containing various instructions and conversational prompts. The primary goal is to assess a model's ability to generate not just relevant, but also appropriately styled audio responses. Specifically, the dataset targets the model's proficiency in:… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/Kimi-Audio-GenTest.audion<1K10 likes179 downloads1y agoHugging Face04DCAgent2 /terminal_bench_2__together_ai_moonshotai_Kimi-K2.5_20260203textn<1K0 likes80 downloads8mo agoHugging Face05DCAgent2 /terminus-2__dev_set_71_tasks__together_ai_moonshotai_Kimi-K2.5_20260211textn<1K0 likes52 downloads8mo agoHugging Face06DCAgent2 /dev_set_v2__together_ai_moonshotai_Kimi-K2.5_20260201textn<1K1 likes27 downloads8mo agoHugging Face07reasoning-degeneration-dev /t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think Strategy compliance evaluation on countdown arithmetic — tree-search variant. Model was instructed to use a systematic tree search with explicit backtracking. Compliance is scored by an LLM judge (1-5 Likert). Results Metric Value pass@1 0.2000 Strategy compliance (mean) 1.20 Strategy compliance (min) 1 Strategy compliance (max) 2 Total problems 10… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-countdown-treesearch-together_ai-moonshotai-kimi-k2-thinking-kimi-think.textn<1K0 likes25 downloads7mo agoHugging Face08reasoning-degeneration-dev /t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f Structured reasoning evaluation: instead of injecting synthesized facts, this uses a static system prompt that teaches the model a heuristic search FORMAT with explicit structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]). Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces significantly outperform free-form CoT. This tests whether prompt-time… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-thinking-daf26c4f.textn<1K0 likes24 downloads7mo agoHugging Face09reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 13.9 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-addc05ee.tabularn<1K0 likes24 downloads7mo agoHugging Face10reasoning-degeneration-dev /t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think Strategy compliance evaluation on countdown arithmetic — baseline variant. Model received no strategy instruction (baseline control). Judge still scores against tree search rubric. Compliance is scored by an LLM judge (1-5 Likert). Results Metric Value pass@1 0.5000 Strategy compliance (mean) 1.20 Strategy compliance (min) 1 Strategy compliance (max) 2 Total problems 10… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-countdown-baseline-together_ai-moonshotai-kimi-k2-thinking-kimi-think.textn<1K0 likes23 downloads7mo agoHugging Face11reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637 t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 8/10 (80.0%) Avg turns 4.0 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-9910c637.tabularn<1K0 likes19 downloads7mo agoHugging Face12reasoning-degeneration-dev /t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297 t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297 Structured reasoning evaluation: instead of injecting synthesized facts, this uses a static system prompt that teaches the model a heuristic search FORMAT with explicit structural markers ([STEP], [PRUNE], [BACKTRACK], [REVIEW OPTIONS], [SOLUTION FOUND]). Inspired by HandCraftedCountdownSearch — models SFT'd on structured search traces significantly outperform free-form CoT. This tests whether prompt-time format… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-structured-reasoning-full-together_ai-moonshotai-kimi-k2-5-6921f297.textn<1K0 likes18 downloads7mo agoHugging Face13reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4 t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 19.8 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-f7244bd4.tabularn<1K0 likes18 downloads7mo agoHugging Face14reasoning-degeneration-dev /t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244 t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244 Wingdings compliance evaluation on countdown arithmetic. Model must reason using only symbolic characters (arrows, checkmarks, boxes, etc.) while solving arithmetic countdown problems. Results Metric Value pass@1 0.3000 Compliance (mean) 0.3583 Compliance (min) 0.2491 Compliance (max) 0.4841 Total problems 10 Details Parameter Value Model… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-countdown-together_ai-moonshotai-kimi-k2-thinking-61281244.textn<1K0 likes16 downloads7mo agoHugging Face15reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061 t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 1/10 (10.0%) Avg turns 1.4 Errors 4 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-20637061.tabularn<1K0 likes16 downloads7mo agoHugging Face16reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 3.5 Errors 0 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-instruct-78cc6a2e.tabularn<1K0 likes15 downloads7mo agoHugging Face17reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885 t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 6.4 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-bef08885.tabularn<1K0 likes14 downloads7mo agoHugging Face18reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5 t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 4/10 (40.0%) Avg turns 8.2 Errors 1 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-thinking-429343f5.tabularn<1K0 likes14 downloads7mo agoHugging Face19reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 5.6 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-e271b21c.tabularn<1K0 likes13 downloads7mo agoHugging Face20reasoning-degeneration-dev /t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491 t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 10/10 (100.0%) Avg turns 12.8 Errors 0 Details Parameter… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-backward_chaining-together_ai-moonshotai-kimi-k2-instruct-b54be491.tabularn<1K0 likes13 downloads7mo agoHugging Face21reasoning-degeneration-dev /t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16 t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16 Strategy compliance baseline — FrozenLake arena evaluation. No strategy instruction was given. This is the baseline to observe the model's natural reasoning patterns across game turns. Per-turn reasoning is extracted from transcripts for manual review. Results Metric Value Win rate 5/10 (50.0%) Avg turns 11.3 Errors 2 Details Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-strategy-arena-frozenlake-baseline-together_ai-moonshotai-kimi-k2-thinking-b6fc7e16.tabularn<1K0 likes13 downloads7mo agoHugging Face22reasoning-degeneration-dev /t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5 t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5 Wingdings compliance evaluation on MuSR murder mysteries. Model must reason using only symbolic characters while solving murder mystery problems (means, motive, opportunity). Results Metric Value pass@1 0.9000 Compliance (mean) 0.1603 Compliance (min) 0.0979 Compliance (max) 0.3724 Total problems 10 Details Parameter Value Model… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-musr-murder-together_ai-moonshotai-kimi-k2-thinking-8c1b04d5.textn<1K0 likes12 downloads7mo agoHugging Face23Micsiu /eval-terminus-2-swebench-verified-random-100-folders-together-ai-moonshotai-kimi-e2740f8ftextn<1K0 likes12 downloads6mo agoHugging Face24DCAgent2 /dev_set_v2__together_ai_moonshotai_Kimi-K2.5_20260202textn<1K0 likes9 downloads8mo agoHugging Face25DCAgent2 /DCAgent2_bfcl-parity_moonshotai_Kimi-Dev-72B_20260226_200826textn<1K0 likes6 downloads7mo agoHugging Face26reasoning-degeneration-dev /t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73 t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73 Wingdings compliance evaluation on TextArena games. Model must reason using only symbolic characters while playing interactive games. Compliance is measured on reasoning text only, not action commands. Results Metric Value Win rate 9/10 Compliance (mean) 0.2951 Errors 0 Details Parameter Value Model together_ai/moonshotai/Kimi-K2-Thinking Thinking… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/t1-wingdings-arena-frozenlake-together_ai-moonshotai-kimi-k2-thinking-4636da73.tabularn<1K0 likes5 downloads7mo agoHugging Face27DCAgent2 /DCAgent_dev_set_v2_moonshotai_Kimi-Dev-72Btextn<1K0 likes5 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.