CoolFace
Modelpublic

unitreerobotics/UnifoLM-ER-1

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
7likes626downloads
Model Card

UnifoLM-ER-1-4B

Project Page

Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image-text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments.

Benchmark Results

ModelOpen SourceRoboVQAEgo-Plan2RefSpatial-BenchWhere2PlacePixmo-PointBLINKCV-BenchEmbSpatialRoboSpatialSATVSI-BenchVSRERQARealWorldQAMMEMMMU_VAL
UnifoLM-ER-1-4BYes62.455.161.782.073.893.4<sup>†</sup>88.688.973.176.054.288.150.069.82223.354.7
RoboBrain2.0-7B<sup>*</sup>Yes30.033.2342.263.654.783.985.776.354.275.336.184.069.42057.444.4
Robix-7B<sup>*</sup>No63.641.929.587.686.577.471.144.683.342.570.72332.8
Pelican-7B<sup>*</sup>Yes31.833.722.357.320.479.473.257.552.052.882.239.869.32141.951.1
Cosmos-R1-7B<sup>*</sup>Yes38.826.05.62.98.276.768.942.482.725.482.467.62157.437.4
Cosmos3-Super-64B<sup>*</sup>Yes57.071.090.3<sup>†</sup>88.070.060.951.2
Qwen3-VL-4B<sup>*</sup>Yes47.740.746.663.048.385.0<sup>†</sup>85.179.661.768.759.381.641.371.02325.257.8
Qwen3-VL-8B<sup>*</sup>Yes43.349.754.261.951.073.8<sup>†</sup>86.278.566.967.359.483.245.870.62412.562.3
Embodied-R1-3B<sup>*</sup>Yes51.826.539.769.549.478.5<sup>†</sup>82.767.447.476.326.635.2
Embodied-R1.5-8B<sup>*</sup>Yes61.053.854.274.064.883.0<sup>†</sup>86.978.169.774.756.146.0
Molmo2-ER-4B<sup>*</sup>Yes52.554.085.7<sup>†</sup>87.878.878.074.546.8
Hy-Embodied-VLM-1.0-30B-A3B<sup>*</sup>Yes49.653.465.064.687.3<sup>†</sup>89.782.769.478.060.8
MiMo-Emb-7B<sup>*</sup>Yes62.043.048.063.642.3581.388.276.261.778.648.579.046.766.32320.826.4
Thinker-4B<sup>*</sup>Yes62.763.761.072.057.484.686.380.270.872.765.481.571.92323.446.2
Wall-OSS-0.5-3B<sup>*</sup>Yes15.03344
Lumo-1-Stage1-7B<sup>*</sup>No51.069.182.486.475.662.674.7
Gemini-ER 2<sup>‡</sup>No35.490.6<sup>†</sup>90.481.451.171.0
Gemini-ER 1.5<sup>‡</sup>No41.848.083.673.457.762.039.947.0
Gemini 2.5 Pro<sup>‡</sup>No33.637.088.6<sup>†</sup>85.978.071.374.751.156.0
Gemini 2.5 Flash<sup>‡</sup>No41.248.080.3<sup>†</sup>85.576.273.473.345.347.5
Gemini 3.1 Pro<sup>‡</sup>No70.061.086.1<sup>†</sup>88.665.147.565.2
GPT-5.6-sol<sup>‡</sup>No58.351.185.6<sup>†</sup>85.280.766.821.364.8
GPT-6-Astra<sup>‡</sup>No79.669.090.4<sup>†</sup>87.383.373.431.377.7

<sup>*</sup> Results are sourced from the models' official technical reports or publicly available papers.

<sup>‡</sup> Results were obtained through tests using the models' official APIs.

<sup>†</sup> BLINK scores are averaged over the Relative Depth and Spatial Relation subtasks only; all reported results were obtained in our own testing.

Citation

yaml
@misc{unifolm-er-1,
  author       = {Unitree},
  title        = {UnifoLM-WLA-1.0: One Model Driven, Whole-Body Coordination},
  year         = {2026},
}