Updates·October 8, 2026, 10:07
HeyGen: Our new video model is 13 times cheaper, and just as good
AI-generated from the sources listed below and quality-checked by the editors.
HeyGen, known for AI avatars, has launched its first general video model. One second of video with audio costs 0.03 dollars, while Google's Veo 3.1 costs 0.40 dollars. In HeyGen's own blind tests, it performs on par with the best.

HeyGen is best known for virtual presenters: avatars that read a script aloud. Now the company has launched HeyGen Video, its first general video model. It creates the whole scene from a description: people, location, lighting and sound. HeyGen writes that the goal is production-quality video without production prices, for things like product demos, training, onboarding of new employees and property showcases.
A fraction of the price
What sets HeyGen Video apart is the price. According to HeyGen's own overview, one second of video with audio at 768p costs 0.03 dollars when you create video from an image. By comparison, Kling 3.0 Pro costs 0.168 dollars, ByteDance's Seedance 2.0 0.303 dollars and Google's Veo 3.1 0.40 dollars per second. A 10-second clip thus costs 30 cents at HeyGen and 4 dollars at Veo.
The regular price starts at 2 cents per second, and for the rest of October it is halved to 1 cent.
On par with the best in blind tests
HeyGen has had people compare clips side by side without knowing which model made what. On HeyGen's leaderboard, where its own model is set to 1,000 points, Seedance 2.0 gets 955, Kling 3.0 Pro 857 and Veo 3.1 744. In 650 head-to-head comparisons with MiniMax's H3 Max, 55.8 percent preferred HeyGen's clips.
The model is also fast: the actual computation of a 10-second clip from an image takes 3.7 seconds.
How it works
The model is called heygen-video-1 and is built on MiniMax H3, which HeyGen has post-trained itself. It can be used in three ways: with text alone, with an image as the clip's first frame, or with your own images, videos and audio, so that a specific product, person or place looks consistent in the clip.
The clips are 5 to 15 seconds long, up to 2K and in formats from widescreen to vertical mobile video. Speech, background sound and sound effects come in the same pass, so there is no separate voiceover or lip-syncing. According to HeyGen's documentation, the model is best for short, well-defined shots with one subject, one location and one action.
Where can you use it?
The model is available in HeyGen's API and through the services OpenRouter, Runware and ComfyUI. In other words, it is a tool for developers and companies that build video into their own systems.
What does it mean?
At HeyGen's regular price, 30 seconds of video with audio costs less than 1 dollar. That makes AI video cheap enough for everyday use: a product video for the online store or a short instruction for new employees can be made for pocket change instead of requiring a shooting day.
What it means for you
If you have an online store, a small business or need to make training for new colleagues, you can get short videos with speech and sound made for pocket change instead of booking a shooting day. The model is built for developers and companies, so it requires your IT department or a vendor to set it up, but you can already ask them to try it out. Start with a short clip with one product or one action, because that is where it works best. For the rest of October the price is halved, so it is cheap to test.
Sources
Get the week's AI news in your inbox
Choose your level, topics and length. One email a week, unsubscribe at any time.
Subscribe to Promptly Newsletter



