datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRagentic-task-benchmark
YouMind Agentic Task Benchmark
v0.1.1 · Experimental
YouMindInc/agentic-task-benchmark is a small experimental benchmark for
evaluating creative and research task outcomes against selected references.
This release contains four task descriptions, a standard outcome format, and
an offline text scorer. Task definitions and evaluation protocols may change.
Tasks
Configuration
Task
Status
image_generation
Riso portrait
Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.han-cognitive-task-benchmarks-v1
Humanoid Cognitive Task Benchmarks
This dataset provides standardized cognitive tasks
to evaluate humanoid reasoning and problem-solving abilities.
Contents
Logical tasks
Memory challenges
Sequential reasoning tests
Use Cases
Cognitive benchmarking
Intelligence evaluation
Model comparison
Part of
Humanoid Network (HAN)
License
MIT
han-domestic-task-ambiguity-benchmark-v1
Domestic Task Ambiguity Benchmark
A benchmark dataset designed to analyze
how ambiguity in human instructions
affects task understanding in humanoid robots.
Methods
Instructions are annotated based on
explicitness and reference clarity.
Use Cases
Ambiguity handling research
Instruction disambiguation
Limitations
No multi-step task chains included.
Part of
Humanoid Network (HAN)
License
MIT
