Text to video

Text to Video AI Generator

Text-to-video is the shortest path from an idea to something you can post. Write what you want to see, optionally let the enhancer add cinematography vocabulary, choose a preset, and generate. No storyboard, no footage, no model menu. Everything below the composer explains how to write prompts that produce intentional-looking shots instead of random ones.

0/2000
Style
Aspect
Duration
Audio
Quality

Sign in to see your credit estimate. Creating needs a plan, from $4.99 a week.

A prompt formula that works

Models respond to structure. Cover these in order and you will get consistent results.

  1. 01

    Subject

    Who or what is in frame, with two or three concrete details: 'a lighthouse keeper in a wool sweater'.

  2. 02

    Action

    One clear motion per clip: 'pouring coffee', 'turning to look at the camera'. Several actions in five seconds read as chaos.

  3. 03

    Environment

    Where and when: 'inside a lighthouse kitchen at dawn, sea visible through the window'.

  4. 04

    Camera

    Framing and movement: 'slow push-in', 'handheld follow', 'static wide'. Say it explicitly or the model decides.

  5. 05

    Mood and lighting

    'warm tungsten against cold blue daylight', 'overcast soft light', 'neon reflections'.

  6. 06

    Audio (optional)

    With a native-audio preset: 'gulls outside, kettle whistling'. Skip it for silent drafts.

Cinematic controls without a menu of models

Prompt enhancement

One tap rewrites a rough idea into a structured prompt. It keeps your subject and intent and adds the missing camera and light language.

Negative prompt

List what to avoid: text overlays, extra fingers, watermarks, colour casts.

Aspect ratio and duration

9:16 for shorts, 16:9 for YouTube, 1:1 for grids. Durations from 4 to 15 seconds depending on the routed model.

Seeds

Keep a seed to reproduce a look across variations, or leave it random for variety.

Prompts to try

Click one to load it into the composer on this page.

How model routing works

You never have to read a model comparison to make a clip. Each request is scored against every enabled model in our registry using the capability it needs (text-to-video, image-to-video, first-and-last frame, reference video, character reference), the preset you picked, the aspect ratio and duration, and the model's live health. The best match is submitted; the next best are kept as fallbacks.

Costs are computed the same way. The quote you see before generating is the price for the model that will actually run, including any resolution or audio surcharge, and it is reserved rather than charged until the clip completes.

Text-to-video models in the registry

Live from our model registry. Availability and specs update as models change.

Veo 3.1

Premium maximum-quality video with native audio.

audioup to 8s4k2 refs

Seedance 2.0

Balanced cinematic video with strong motion and reference support.

audioup to 15s4k9 refs

Gemini Omni

Character-consistent video from reusable references.

audioup to 10s1080p7 refs

Kling 3.0

Narrative multi-shot video generation.

audioup to 15s4k

Wan 3.0 Prime

Fast, low-cost video for quick iterations.

up to 10s1080p4 refs

Use cases

  • Explainer b-roll when you have no footage
  • Hooks and intros for short-form posts
  • Dream sequences and abstract loops
  • Pitch visuals and pre-visualisation
  • Music visualisers

Private by default, publishable in one tap

Private library

Every creation is stored privately with its prompt, settings and model, so you can rerun, tweak or download it later.

Visibility controls

Publish as public, followers-only, unlisted or private, and decide whether others may remix.

Downloads

Full-resolution files, permanently hosted on our storage, downloadable from the library at any time.

Remix lineage

If you do publish, remixes credit you and appear under your original.

Frequently asked

How specific should my prompt be?
Start with subject, action and setting. Add camera movement and lighting if you care about them. The enhancer can fill in the rest, and the prompt is saved with the clip so you can refine it.
Can text-to-video include sound?
Yes, when the routed model supports native audio, which Cinematic and Story presets favour. The audio toggle in the composer tells the router to prefer those models.
Is text to video free?
No. Creating needs a plan, which starts at $4.99 a week and includes a credit allowance. Every generation shows its credit cost before you confirm, and failures are refunded in full.
What happens if a generation fails?
The reserved credits are released back to your wallet immediately. If the failure was on the provider side and a fallback model can serve the same request, the router retries automatically and you are still charged once.
Do I have to pick an AI model?
No. Choose a style preset such as Fast, Cinematic or Story and the router selects the best available model for that intent, aspect ratio, duration and the inputs you attached. If a model is degraded, the router falls back to a comparable one without charging you twice. Advanced users can still pin a specific model.
How much does a generation cost?
Every generation is quoted in credits before you confirm, based on the routed model, duration and resolution. Credits are reserved when you start and released automatically if the generation fails, so you only pay for what completes. Plans start at $4.99 a week and include a credit allowance; subscribers can top up with credit packs.
Are my generations private?
Yes by default. Everything you generate lands in your private library. You decide per creation whether to publish it to the feed, share it unlisted, restrict it to followers or keep it private, and whether others may remix it.
Can I download the videos?
Yes. Every completed generation can be downloaded in full resolution from your library. Files are stored permanently on our own media storage, not on the provider's temporary links.

Ready when you are.

Plans from $4.99 a week. Publish when you're proud of it.

Generate