gpt-image-2 is OpenAI’s flagship image model. Text-to-image and image-to-image (mask-driven edits) live on a single endpoint, billed per image generated. Strong on layout fidelity, in-image typography, and product/scene composition — the place to reach for when “looks like a real OpenAI render” matters more than the lowest per-image cost.
Pricing: 0.08 per generated image; failures don’t bill.
Protocols
Quickstart
Capabilities
When to use
- Marketing creative — hero images, social cards, anywhere typography-in-image matters.
- Product mockups — fidelity on materials, lighting, and small print holds up better than most domestic alternatives.
- DOSIA
generate_imagetool — the main brain will resolve “draw me an X” to this model by default when permission is granted.
- High-volume or budget-sensitive work —
nano-banana-v2is materially cheaper for the same shape.
Next
nano-banana-v2— cheaper Google image model with image-to-image- Multimodal endpoints — overview of image / video / audio / embedding surfaces