Google Veo 3.1

Models

Text-to-video with synchronised audio, from Google DeepMind.

How it works

  1. 1

    Write the shot

    Describe the subject, the action and the camera in one prompt.

  2. 2

    Add references

    Up to three images can steer the look. With references the clip is always 8 seconds.

  3. 3

    Generate

    Veo renders the clip at the resolution and length you picked.

Model: google/veo-3.1