xAI has put a budget tier on its video generation stack. Grok Imagine Video 1.5 Lite, announced on 8 October, is available through the Imagine API and third-party gateways including OpenRouter, Vercel's AI Gateway and fal, priced at $0.02 per second of output at 480p, $0.03 at 720p and $0.14 at 1080p, with clips running one to 15 seconds. Image inputs are billed separately at $0.01 each. At the low end, a ten-second 480p draft costs about 20 cents.
The Lite model undercuts xAI's own full Grok Imagine Video 1.5, which lists at $0.08, $0.14 and $0.25 per output second at the same three resolutions. That makes Lite 75 percent cheaper at 480p and about 44 percent cheaper at 1080p, in exchange for a lower quality ranking and a reduced feature set. Per the published specifications, Lite handles text-to-video and image-to-video and adds 1080p output — although that 1080p is upscaled from 720p generation rather than native, a detail worth checking before committing a final render to it.
The distribution push is notable: the model is listed on OpenRouter under the slug x-ai/grok-imagine-video-1.5-lite, and both Vercel and fal carry it. One model directory dates the API release to 1 October, with xAI's own announcement and the gateway listings following on 8 October; the pricing and availability figures here reflect the listings as published on 8-9 October.
The independent check is mixed but pointed. Artificial Analysis placed Lite 17th on its AA-Video-T2V v2.0 leaderboard for silent text-to-video — two positions above Google's Veo 3.1, which costs roughly $0.40 per second at 1080p, making Lite about a third of the price of the model it outranks. The benchmarker's median generation time for a ten-second 1080p clip was 60.5 seconds. It also makes a sharper claim: no other model on its silent list is both faster and higher quality than Lite, putting the budget tier on the price-quality frontier.
The caveats are real. Rank 17 is not first place, and two places on a crowded board is a modest preference gap, not a generational one — the full Grok Imagine Video 1.5 ranks six positions higher. The ranking covers silent output only; audio and lip sync are judged separately, and Artificial Analysis flags lip sync among Lite's weaknesses, alongside lighting, text rendering and human anatomy. Lite also lacks the full model's reference-image, voice-reference and keyframe controls, so sequences that need a pinned ending still require the flagship.
The pricing move lands in a video market that is repricing fast. Shengshu's Vidu Q4 Preview opened third on Artificial Analysis's image-to-video board this week; Kandinsky Lab released open weights for Kandinsky 6.0 Video with synchronized video and audio; and Artificial Analysis now labels derived models built on MiniMax's H3, two of which sit among the top five entries on its text-to-video board. Cheap per-second rates are becoming the competitive baseline rather than the differentiator.
For studios and developers, the economics are the story. Iterating on prompts is where video generation budgets die, and a 20-cent 480p draft changes how many attempts a creator can afford before committing to a final render. Whether quality holds at delivery size — particularly with upscaled 1080p — is the question the leaderboard cannot answer for any specific shot.
Comments (0)
Log in to join the discussion
Log InNo comments yet