unitreerobotics/UnifoLM-ER-1
UnifoLM-ER-1-4B
Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image-text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.
Benchmark Results
<sup>*</sup> Results are sourced from the models' official technical reports or publicly available papers.
<sup>‡</sup> Results were obtained through tests using the models' official APIs.
<sup>†</sup> BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.
Citation
@misc{unifolm-er-1,
author = {Unitree},
title = {UnifoLM-WLA-1.0: One Model Driven, Whole-Body Coordination},
year = {2026},
}