AIREITER

GPT Image 2 Prompt Guide: Text, Edits, and References

Last Updated: 2026-09-29 01:51:09

Treat a GPT Image 2 prompt as a production brief, not an adjective list.

The prompt format that works for GPT Image 2

Start with the finished artifact, then describe the subject, composition, scene, text, style, and constraints. OpenAI’s Cookbook recommends a consistent order of scene, subject, key details, and constraints, while its examples also support labeled sections for complex requests (OpenAI Cookbook).

TASK: Create a [poster / product photo / UI mockup / editorial image].
SUBJECT: [the person, object, or scene that must remain accurate].
COMPOSITION: [aspect ratio, viewpoint, framing, placement, negative space].
SCENE AND LIGHT: [location, materials, light direction, color].
TEXT: [exact words in quotes, role, placement, line count].
STYLE: [photorealistic or medium, lens feel, palette, texture].
CONSTRAINTS: [what to avoid and what must remain unchanged].

Use only the slots the job needs; add more when something fails.

Concrete visual facts beat “stunning,” “premium,” or “ultra-detailed.” “Soft window light from camera left, matte ceramic, 50mm feel, empty upper-right third” gives GPT Image 2 decisions it can represent.

Text rendering: write the layout, not just the words

GPT Image 2 is worth testing when an image must carry a headline, label, sign, or interface. Keep copy short, quote load-bearing text, identify its role, and require verbatim rendering with no extra words.

Use this pattern:

TEXT: Main headline reads exactly "NIGHT MARKET".
ROLE: Large condensed sans-serif headline across the upper-left.
LAYOUT: One line, white lettering, high contrast, centered vertically in its zone.
CONSTRAINTS: Render the quoted text verbatim. No extra words, duplicate text, logos, or watermark.

For several blocks, give each block its own line. A poster prompt might say “headline,” “subhead,” and “footer,” rather than describing all copy in one sentence. For multilingual work, label each language and paste the exact characters instead of asking the model to translate them.

Keep copy short when spelling matters. If a string repeatedly fails, isolate it and change only that variable on the next attempt.

SymptomLikely causeFix
Misspelled wordCopy is long or not isolatedShorten the string, quote it, and name its role
Extra wordsThe layout has no text boundaryState “no extra text” and limit the number of text blocks
Missing linePlacement or line count is unclearSpecify the zone and number of lines

Inspect every label, price, URL, and UI control at full size; finish legal or production typography in a design tool.

Multi-image references and local edits

Give every reference image a labeled role stating what to borrow and what to ignore. fal’s GPT Image 2 guide says its edit workflow can accept up to 16 reference images; confirm the limit in your provider’s interface or API (fal guide).

REFERENCE ROLES:
Image 1: base portrait. Preserve the face, pose, body proportions, framing, and background.
Image 2: jacket reference. Copy only the jacket cut, fabric, and color.
Image 3: boots reference. Copy only the boot shape and material.

CHANGE: Dress the person in Image 1 using the jacket from Image 2 and boots from Image 3.
PRESERVE: Face, hairstyle, hands, pose, camera angle, lighting, and body proportions.
CONSTRAINTS: No extra accessories, logos, or redesigned clothing.

For a local edit, describe one target region and one operation. If the tool supports a mask, mask the region and tell GPT Image 2 to change only that region. Without a mask, use spatial language such as “the sign in the upper-right” or “the red cup on the left side.” Repeat the preservation list in every follow-up prompt.

CHANGE ONLY: Replace the red sign in the upper-right with a blank cream sign.
PRESERVE: Building geometry, window reflections, people, shadows, camera angle, and color balance.
DO NOT: Recompose the storefront or add new text.

A useful iteration loop is: generate, identify the single largest error, change one variable, and keep the rest of the brief unchanged. If an edit invents details, strengthen the preservation list before changing the target.

What to do when GPT Image 2 refuses a request

Do not disguise a prohibited request with euphemisms, role-play, or instructions to ignore safeguards; remove the risky element or reframe the legitimate creative goal.

If the request contains…Use a safer brief…
Sexualized nudity or an ambiguous-age personAn adult subject in ordinary, non-explicit clothing; state the age clearly when relevant
A living person’s likeness in a sensitive or deceptive sceneA fictional adult character with non-identifying features
A copyrighted character or brand identityAn original character described by general visual traits, without logos or trademarked names
Graphic injury or sexual violenceNon-graphic aftermath, symbolic storytelling, or a neutral documentary-style scene
Instructions to evade moderationA direct, policy-compliant description of the intended scene

If a harmless request is blocked, simplify it and remove unnecessary sensitive terms. For example, replace “make this person look nude” with “change the outfit to a plain long-sleeve shirt and trousers,” while preserving pose and lighting. If the platform still refuses, use its support or policy route; no prompt can guarantee approval.

GPT Image 2 and Nano Banana 2: two prompting habits, not a magic syntax

Labeled lines or JSON-like blocks are optional organizers, not a special syntax. OpenAI’s Cookbook accepts minimal prompts, descriptive paragraphs, labeled segments, and JSON-like structures when the intent and constraints remain clear (OpenAI Cookbook). Google’s Nano Banana guidance likewise uses ordinary natural language: subject, action, context, composition, style, and explicit relationships between reference images (Google DeepMind prompt guide).

The useful difference is operational:

JobGPT Image 2 habitNano Banana 2 habit
New photoreal imageUse a labeled creative brief and say photorealistic when realism is the goalStart with subject, action, context, composition, and style
Multiple referencesAssign each image a role and list the invariants to preserveName the relationship between each reference and the new scene
Local edit“Change only X; preserve Y” and use a mask when availableState the edit and preservation rules in natural language
IterationChange one variable and repeat the preserve listIterate conversationally, keeping reference identities named

The practical distinction is workflow emphasis: start with GPT Image 2 for text-bearing or structured briefs, and test Nano Banana 2 when reference fidelity is your main concern. This is a starting heuristic, not a benchmark result.

Prompt patterns by job

Photoreal product image

Create a 4:5 ecommerce hero image of a matte navy ceramic mug on a pale stone surface. Three-quarter view, mug in the lower-right third, empty upper-left for copy. Softbox from upper left, gentle fill from the right, photorealistic product photography, 50mm feel. Preserve the mug’s cylindrical shape and cream handle. No text, logo, watermark, extra products, or hands.

Text-heavy poster

Create a vertical 9:16 event poster. A lone figure stands on an empty street under one warm lamp at twilight. Main headline reads exactly "THE NIGHT BEGINS AT EIGHT" in large gold serif type across the lower third. Subhead reads exactly "A STORY ABOUT WAITING" beneath it in small white sans serif. Two text lines only; render both verbatim. No extra text, QR code, logo, or watermark.

Local background edit

Change only the background to a quiet beach at sunset. Preserve the person’s face, hair, clothing, pose, expression, camera angle, subject lighting, and framing. Match the new background perspective to the original image. Do not add people, text, logos, or extra objects.

Reference-based outfit edit

Image 1 is the base person. Image 2 is the clothing reference. Replace only the outfit using Image 2’s jacket color, cut, and fabric. Preserve Image 1’s face, hair, body proportions, hands, pose, background, lighting, and framing. Add no jewelry, text, or logos.

FAQ

How long should a GPT Image 2 prompt be?

There is no fixed word count. A useful prompt is long enough to specify the artifact, subject, composition, exact text, and constraints; delete any sentence that adds no visual decision.

How many reference images can GPT Image 2 use?

The fal guide says GPT Image 2 edits can accept up to 16 reference images, but provider limits can differ. Check the endpoint you are calling and assign each image a clear role.

What should I do if GPT Image 2 refuses a safe request?

Remove ambiguous sensitive details, describe the legitimate visual goal directly, and try a fictional or non-identifying subject. Do not ask the model to bypass safeguards.

The practical choice

Before generating, check four things: exact text, subject identity, preserved regions, and unwanted additions. Then write one structured brief, lock the parts that matter, and revise one variable at a time.

For the model’s API availability and current pricing, see GPT Image 2 API pricing. For model selection rather than prompting, use GPT Image 2 vs Nano Banana.