Shengshu Technology, the Beijing-based company behind the Vidu video generation family, released the first public preview of its next flagship model on Thursday, Vidu Q4 Preview. The headline change is economic as much as it is technical: the model starts at 0.09 yuan (about US$0.013) per second of generated video, and supports output up to native 4K resolution.
According to the company, Q4 Preview concentrates its upgrades on three fronts: character performance, camera language and complex visual effects. The model better coordinates facial emotion, body movement and the rhythm of the shots around a character, handles dynamic camera moves, cuts and complex shot transitions more reliably, and improves the generation of hard effects like explosions, smoke and particle systems.
Control is the other pillar. Creators can feed the model up to 15 reference images and 3 reference audio clips, which the company says gives finer control over characters, scenes, costumes, props and sound. Multi-reference input has become the practical differentiator in this generation of video models, where matching a specific face or product across shots matters more to commercial users than raw visual quality.
The pricing is deliberately aggressive. At 0.09 yuan per second under the launch promotion, a 10-second clip costs about 0.9 yuan - roughly 13 US cents. Shengshu says that under comparable output specs and billing terms, the same budget now produces up to five times more content than before. The model ships with 540P, 720P, 1080P, 2K and 4K output options.
"AI technology can only translate into sustainable productivity when it enters real scenarios and solves real problems," said Luo Yihang, co-founder and CEO of Shengshu. "We hope to keep converting model efficiency into a lower usage threshold, so that more individuals, teams and enterprises can actually use it, and afford it."
The company is positioning Q4 Preview at advertising and e-commerce teams that need to test many creative variants, at animation and short-drama studios that want cheap storyboard rehearsal, character blocking and shot verification before committing to production, and at cultural tourism and education content. Shengshu says it is already working with film and content platforms as well as cloud providers on API integration.
Shengshu's core team proposed the U-ViT diffusion architecture in 2022, and Vidu made its public debut in April 2024. Since then the company has pushed reference-to-video generation, multi-subject consistency, joint audio-and-video output, multi-shot narrative and real-time interaction. Thursday's release is explicitly a preview: the company plans to open it early, collect feedback from real production scenarios and fold it into the final version.
The release lands in a market where per-second pricing is becoming the competitive front line for Chinese video models, and where global rivals have also been cutting costs. Shengshu's numbers are company-reported and have not been independently verified, but the direction is clear: video generation is being priced for volume production work rather than for demos, and whoever holds the cost-per-usable-second line wins the commercial accounts.
Comments (0)
Log in to join the discussion
Log InNo comments yet