nvidia/Kimodo-SOMA-RP-v1
Kimodo: Controllable Kinematic Motion Diffusion at Scale
[Paper](https://research.nvidia.com/labs/sil/projects/kimodo/assets/kimodo_tech_report.pdf), [Project Page](https://research.nvidia.com/labs/sil/projects/kimodo/)
Description:
Kimodo (Kinematic Motion Diffusion) generates three-dimensional (3D) skeletal body animations given a text prompt and/or constraints on the motion like full-body poses, end-effector joint positions, paths, and waypoints to follow.
The Kimodo model family includes models trained on different skeletons and datasets:
- Kimodo-SOMA-RP
- Trained on the 30-joint SOMA skeleton with the proprietary Bones Rigplay dataset.
- Kimodo-SOMA-SEED
- Trained on the 30-joint SOMA skeleton with the open Bones-SEED dataset.
- Kimodo-G1-RP
- Trained on the proprietary Bones Rigplay dataset retargeted to the 34-joint Unitree G1 robot skeleton.
- Kimodo-G1-SEED
- Trained on the open Bones-SEED dataset retargeted to the 34-joint Unitree G1 robot skeleton.
- Kimodo-SMPLX-RP
- Trained on the proprietary Bones Rigplay dataset retargeted to the 22-joint SMPLX-body skeleton.
This release pertains to Kimodo-SOMA-RP. This model is ready for commercial use.
License:
This model is released under the NVIDIA Open Model License.
Deployment Geography:
Global
Use Case: <br>
The model is intended for users with any level of animation experience to create 3D human motion data for their application. This may include:
- Demonstrations for humanoid robots
- Digital human motion for digital twin and industrial simulations
- Digital human motion for synthetic data
- Animations for game and media development
Release Date: <br>
Github [03/16/2026] via link <br> HuggingFace [03/16/2026] via link <br>
References:
- Technical report: Kimodo: Scaling Controllable Human Motion Generation
- Webpage: link
Model Architecture:
Architecture Type: Diffusion Model <br> Network Architecture: Novel Two-Stage Transformer <br> Model Size: 282 M parameters
Inputs: <br>
Input Types: Text, Duration (Num Frames), Pose Constraints <br>
Input Formats:
- Text: String
- Duration: Integer
- Pose Constraints: Matrix
Input Parameters:
- Text: One-Dimensional (1D)
- Duration: One-Dimensional (1D)
- Pose Constraints:
- One-Dimensional (1D) frame index of each constraint
- Features to constrain may include Three-Dimensional (3D) joint positions, (3x3) joint rotation matrices, Two-Dimensional (2D) heading direction, and/or Two-Dimensional (2D) root position
Other Properties Related to Input: Maximum duration is 10 sec (300 frames at 30 frames per second).
Outputs
Output Type: Skeleton Motion: Root Translation and Joint Rotations <br>
Output Formats:
- Root Translation: Matrix
- Joint Rotations: Matrix
Output Parameters:
- Root Translation: Two-Dimensional (
num_framesx 3) - Joint Rotations: Four-Dimensional (
num_framesx 30 x 3 x 3)
Other Properties Related to Outupt:
- Motions are at 30 frames per second (30 fps)
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions. <br>
Software Integration:
Runtime Engines:
- PyTorch
Supported Hardware Microarchitecture Compatibility: <br>
- NVIDIA Ampere
- NVIDIA Blackwell
- NVIDIA Lovelace
Supported Operating Systems:
- Linux
- Windows
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment. <br>
Model Version
Kimodo-SOMA-RP-v1
Training and Testing Datasets:
Name: Proprietary Bones Rigplay Dataset
Data Modalities
- Text
- Human Motion Capture
Data Size:
- Less than 1 Billion tokens of text
- 700 hours of human motion capture
Data Collection Method <br> Automatic/Sensors
Labeling Method <br> Hybrid: Automatic/Sensors, Human
Properties: 700 hours of captured human body motions on the SOMA skeleton with corresponding text descriptions. Split into 90%/10% train/test splits. Various augmentations were employed to expand text and motion variety.
Quantitative Evaluation <br> For test set evaluation, please refer to the technical report
Inference:
Acceleration Engine: N/A<br>
Test Hardware: <br>
- GeForce RTX 3090
- GeForce RTX 4090
- GeForce RTX 5090
- NVIDIA A100
- NVIDIA L40S
- NVIDIA L4
- NVIDIA RTX 6000 Ada
- NVIDIA RTX A6000
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
For more detailed information on ethical considerations for this model, please see the Bias, Explainability, Safety & Security, and Privacy Subcards below. <br>
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
