datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agentic-task-benchmark
YouMind Agentic Task Benchmark
v0.1.1 · Experimental
YouMindInc/agentic-task-benchmark is a small experimental benchmark for
evaluating creative and research task outcomes against selected references.
This release contains four task descriptions, a standard outcome format, and
an offline text scorer. Task definitions and evaluation protocols may change.
Tasks
Configuration
Task
Status
image_generation
Riso portrait
Defined; reference image supplied… See the full description on the dataset page: https://huggingface.co/datasets/YouMindInc/agentic-task-benchmark.han-cognitive-task-benchmarks-v1
Humanoid Cognitive Task Benchmarks
This dataset provides standardized cognitive tasks
to evaluate humanoid reasoning and problem-solving abilities.
Contents
Logical tasks
Memory challenges
Sequential reasoning tests
Use Cases
Cognitive benchmarking
Intelligence evaluation
Model comparison
Part of
Humanoid Network (HAN)
License
MIT
han-domestic-task-ambiguity-benchmark-v1
Domestic Task Ambiguity Benchmark
A benchmark dataset designed to analyze
how ambiguity in human instructions
affects task understanding in humanoid robots.
Methods
Instructions are annotated based on
explicitness and reference clarity.
Use Cases
Ambiguity handling research
Instruction disambiguation
Limitations
No multi-step task chains included.
Part of
Humanoid Network (HAN)
License
MIT
