AwakePix
AI Text to Video Generator
Describe a scene and get video back. What each model costs in credits is published below, with an estimated render time, before you spend anything.
Nothing to anchor to, for better and worse
Generating from text removes the constraint that makes image-to-video predictable. There is no first frame to stay consistent with, so the model invents the subject, the setting and the light. That is the appeal, and it is also why two runs of the same prompt give you two different scenes. Plan to generate more than once.
Write for the four channels the model reads
These models were trained on video paired with detailed description, so a vague prompt returns a vague, averaged result. State the camera path, the subject action, the setting and the ambient sound as separate clauses. Our prompt optimiser applies the official per-model rules rather than a generic rewrite, and it declines to use Seedance-specific techniques on models whose makers never documented them.
What each model costs for five seconds
At 720p: Wan 3.0 is 38 credits, MiniMax H3 is 38, H3 Max is 40 and Seedance 2.0 is 100. At 1080p: 75, 62 and 249 — H3 Max tops out at 768p and has no 1080p tier. Wan 3.0 is deliberately the default first step because it is the cheapest way to find out whether the idea works at all.
Expect two to four minutes
Video inference is slow everywhere, not just here. Our measured range is 100 to 220 seconds. You can leave an email and close the tab rather than watching a spinner, and the model list shows the estimate next to each option before you choose.
- Four models with published per-clip costs
- 5 or 10 seconds, 720p or 1080p
- Prompt optimiser follows each model's own rules
- Requires credits — the free daily clip is image-to-video
Questions
How is text to video different from animating a photo?
With a photo the model is constrained: it has a real first frame and mostly has to stay consistent with it. From text alone it invents everything, which means more freedom and far less predictability. The same prompt run twice gives you two different scenes.
What makes a prompt work?
Separate the four things the model handles independently: what the camera does, what the subject does, what the setting looks like, and what it should sound like. 'A quiet street at dusk' produces an average of every dusk street. 'Camera static, steam drifting from a vent, a neon sign flickering, one passer-by walking out of frame, soft city ambience' produces a shot.
Should I use timeline segments?
Only on Seedance models. Their documentation explicitly supports splitting a prompt into 0-3s and 3-6s beats, and it helps on longer or more complex motion. Wan and Kling have no such published guidance, so we do not apply that technique to them rather than guess.
Is there a free tier for text to video?
No. The free daily clip is image-to-video only, because the free model is locked to Seedance 2.0 Mini in its image-to-video form. Text to video needs credits: 38 for Wan 3.0 at 720p, up to 249 for Seedance 2.0 at 1080p for a 5-second clip.
Which model handles text prompts best?
Seedance models are the ones with published prompt-writing guidance, which is why our prompt optimiser applies their rules by default. MiniMax H3 and H3 Max return their own audio. Wan 3.0 is the cheapest way to see whether an idea holds up at 38 credits for 5 seconds at 720p, and H3 Max comes back fastest — 57 seconds in our measurements.
Can I get a longer clip?
Ten seconds is the other supported length, and it costs exactly double — 75 credits on Wan 3.0 at 720p instead of 38. Pricing is per second of output, so there is no discount for length.
Does it generate sound?
MiniMax H3 and H3 Max do, natively — they return a soundtrack whether or not you ask for one, and they do not take an audio switch at all. Wan 3.0 and Seedance 2.0 have the switch, and cost the same whether it is on or off.