Chinese robotics maker Unitree has released UnifoLM-WLA-1.0, a general-purpose whole-body vision-language-action (VLA) foundation model for its humanoids — a single 6-billion-parameter checkpoint that the company says coordinates 64 distinct tasks on a real G1 robot, from making the bed and loading a washing machine to folding clothes and sorting objects.
The headline numbers, as stated on the project page: 6B parameters, roughly 2,500 hours of real robot training data, and 64 tasks split into 10 whole-body movements and 54 tabletop manipulations. The system supports two-finger parallel grippers as well as two different five-finger dexterous hands, which is notable because most open VLA releases target a fixed gripper on a stationary tabletop arm.
The architecture is a stacked pipeline. At the front sits UnifoLM-ER-1, an embodied reasoner built on Qwen3-VL-4B that Unitree says was trained on more than 5 million embodied samples covering point prediction, detection, multi-image reasoning, 2D trajectories, 3D detection and spatial question answering, co-trained with general image-text data to preserve perception and language skills. On top of that, the model predicts future dynamic regions using optical flow compressed through a VQ-VAE into fixed-length tokens, discretizes actions with residual VQ across end-effector, hand and lower-body channels, and adds an MMDiT action expert for continuous control.
Demos published with the release show the G1 performing household chores — bed making, loading laundry, folding, sorting — suggesting the company is aiming at general home and light-service work rather than the industrial single-task deployments that dominate commercial robotics today.
Whether it is genuine progress or a well-cut demo remains an open question, and Unitree's own materials leave it that way. The claim that the model beats "many open-source models" on embodied benchmarks is self-reported, with no published benchmark tables or third-party reproductions. More importantly for the open-source community, the project page currently marks code, model weights and datasets as "coming soon" — so it is not yet possible for outside labs to verify any of the claims or even run the model.
The release still matters as a signal of direction. Whole-body policies for humanoids are newer and harder than the tabletop-arm VLAs that have flooded open source over the past year, because they must coordinate locomotion, balance and manipulation simultaneously. If a single 6B model really can replace 64 task-specific controllers on the same hardware, integration cost — the main barrier to scaling humanoid pilots — drops accordingly.
For now, the practical takeaway for robotics labs running G1 hardware is to wait for the promised weights and test the tabletop tasks directly. The gap between a launch announcement and a reproducible open release has tripped up larger players, and Unitree's own prior open efforts have varied in how much actually shipped. The next data point to watch is whether the GitHub repository and Hugging Face collection appear with weights that match the demo.
Comments (0)
Log in to join the discussion
Log InNo comments yet