Skip to main content
veo3.1 is Google’s Veo 3.1 model. Same two-phase task-based protocol as the other video models, with one differentiator worth knowing about: native audio generation alongside the video track. The same submit → poll flow produces an MP4 with an audio layer the model invented to match the scene — useful for one-shot deliverables that won’t get a separate sound-design pass. Pricing: $0.40 / second of generated footage — see the rate card. Failures don’t bill; per-second pricing applies to generated footage length, and audio doesn’t add a separate line item on this tier.

Protocols

Quickstart

Capabilities

When to use

  • One-shot deliverable — clip is the final output, no sound-design pass coming.
  • Ambient / atmospheric footage — rain, wind, city noise, where Veo’s native audio is more authentic than dubbing-over-silent footage.
When not to use:
  • You already have your own sound design — audio is a small premium that’s wasted in that flow; drop to veo3.1-fast without audio.

Next