CoolFace
14 results

task-benchmark

BByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.tabular10M<n<100M10 likes7k downloads1y agoHugging FaceBByrneLab /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN PreFLMR M2KR Dataset Card Dataset details Dataset type: M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models. We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.tabular1M<n<10M0 likes1.1k downloads2y agoHugging FacePreFLMR /multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRtext10M<n<100M0 likes357 downloads3y agoHugging FaceYouMindInc /agentic-task-benchmark YouMind Agentic Task Benchmark v0.1.1 · Experimental YouMindInc/agentic-task-benchmark is a small experimental benchmark for evaluating creative and research task outcomes against selected references. This release contains four task descriptions, a standard outcome format, and an offline text scorer. Task definitions and evaluation protocols may change. Tasks Configuration Task Status image_generation Riso portrait Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.texttext-generationn<1K0 likes107 downloads15d agoHugging Facemondaycake /human-last-digital-agent-taskbenchmark Human-Last Digital Agent Taskbenchmark v2 Executable Tasks (8) — use these executable_v2.tar.gz contains 8 tasks that follow the DeepSeek-V4.1 §5.1 pipeline: each has a real Dockerfile, pinned immutable data, a reference solution that has been run end-to-end, and a real verifier (not a file-exists check). Task Domain Reference MAE / result Verifier checks realtime-gdp-nowcast macro-finance AR(4) expanding window, MAE 0.56 16 (incl. anti-look-ahead… See the full description on the dataset page: https://huggingface.co/datasets/mondaycake/human-last-digital-agent-taskbenchmark.0 likes101 downloads7d agoHugging Facecraigwu /GUI_Reflection_Task_Suite_Benchmark GUI Reflection Task Suite Benchmark This is the eval data of GUI Reflection Task Suite. The image of this data should be downloaded from AndroidControl, GUI_Odyssey, ScreenSpot, ScreenSpot_v2. Project Details Project Page: https://penghao-wu.github.io/GUI_Reflection/ Repository: https://github.com/penghao-wu/GUI_Reflection Paper: https://arxiv.org/abs/2506.08012 Citation @article{GUI_Reflection, author = {Wu, Penghao and Ma, Shengnan and Wang… See the full description on the dataset page: https://huggingface.co/datasets/craigwu/GUI_Reflection_Task_Suite_Benchmark.1 likes53 downloads1y agoHugging Face