task-benchmark
multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRagentic-task-benchmark
YouMind Agentic Task Benchmark
v0.1.1 · Experimental
YouMindInc/agentic-task-benchmark is a small experimental benchmark for
evaluating creative and research task outcomes against selected references.
This release contains four task descriptions, a standard outcome format, and
an offline text scorer. Task definitions and evaluation protocols may change.
Tasks
Configuration
Task
Status
image_generation
Riso portrait
Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.human-last-digital-agent-taskbenchmark
Human-Last Digital Agent Taskbenchmark
v2 Executable Tasks (8) — use these
executable_v2.tar.gz contains 8 tasks that follow the DeepSeek-V4.1 §5.1
pipeline: each has a real Dockerfile, pinned immutable data, a reference
solution that has been run end-to-end, and a real verifier (not a
file-exists check).
Task
Domain
Reference MAE / result
Verifier checks
realtime-gdp-nowcast
macro-finance
AR(4) expanding window, MAE 0.56
16 (incl. anti-look-ahead… See the full description on the dataset page: https://huggingface.co/datasets/mondaycake/human-last-digital-agent-taskbenchmark.GUI_Reflection_Task_Suite_Benchmark
GUI Reflection Task Suite Benchmark
This is the eval data of GUI Reflection Task Suite.
The image of this data should be downloaded from
AndroidControl,
GUI_Odyssey,
ScreenSpot,
ScreenSpot_v2.
Project Details
Project Page: https://penghao-wu.github.io/GUI_Reflection/
Repository: https://github.com/penghao-wu/GUI_Reflection
Paper: https://arxiv.org/abs/2506.08012
Citation
@article{GUI_Reflection,
author = {Wu, Penghao and Ma, Shengnan and Wang… See the full description on the dataset page: https://huggingface.co/datasets/craigwu/GUI_Reflection_Task_Suite_Benchmark.
