Treat a GPT Image 2 prompt as a production brief, not an adjective list.
The prompt format that works for GPT Image 2
Start with the finished artifact, then describe the subject, composition, scene, text, style, and constraints. OpenAI’s Cookbook recommends a consistent order of scene, subject, key details, and constraints, while its examples also support labeled sections for complex requests (OpenAI Cookbook).
TASK: Create a [poster / product photo / UI mockup / editorial image].
SUBJECT: [the person, object, or scene that must remain accurate].
COMPOSITION: [aspect ratio, viewpoint, framing, placement, negative space].
SCENE AND LIGHT: [location, materials, light direction, color].
TEXT: [exact words in quotes, role, placement, line count].
STYLE: [photorealistic or medium, lens feel, palette, texture].
CONSTRAINTS: [what to avoid and what must remain unchanged].
Use only the slots the job needs; add more when something fails.
Concrete visual facts beat “stunning,” “premium,” or “ultra-detailed.” “Soft window light from camera left, matte ceramic, 50mm feel, empty upper-right third” gives GPT Image 2 decisions it can represent.
Text rendering: write the layout, not just the words
GPT Image 2 is worth testing when an image must carry a headline, label, sign, or interface. Keep copy short, quote load-bearing text, identify its role, and require verbatim rendering with no extra words.
Use this pattern:
TEXT: Main headline reads exactly "NIGHT MARKET".
ROLE: Large condensed sans-serif headline across the upper-left.
LAYOUT: One line, white lettering, high contrast, centered vertically in its zone.
CONSTRAINTS: Render the quoted text verbatim. No extra words, duplicate text, logos, or watermark.
For several blocks, give each block its own line. A poster prompt might say “headline,” “subhead,” and “footer,” rather than describing all copy in one sentence. For multilingual work, label each language and paste the exact characters instead of asking the model to translate them.
Keep copy short when spelling matters. If a string repeatedly fails, isolate it and change only that variable on the next attempt.
| Symptom | Likely cause | Fix |
|---|---|---|
| Misspelled word | Copy is long or not isolated | Shorten the string, quote it, and name its role |
| Extra words | The layout has no text boundary | State “no extra text” and limit the number of text blocks |
| Missing line | Placement or line count is unclear | Specify the zone and number of lines |
Inspect every label, price, URL, and UI control at full size; finish legal or production typography in a design tool.
Multi-image references and local edits
Give every reference image a labeled role stating what to borrow and what to ignore. fal’s GPT Image 2 guide says its edit workflow can accept up to 16 reference images; confirm the limit in your provider’s interface or API (fal guide).
REFERENCE ROLES:
Image 1: base portrait. Preserve the face, pose, body proportions, framing, and background.
Image 2: jacket reference. Copy only the jacket cut, fabric, and color.
Image 3: boots reference. Copy only the boot shape and material.
CHANGE: Dress the person in Image 1 using the jacket from Image 2 and boots from Image 3.
PRESERVE: Face, hairstyle, hands, pose, camera angle, lighting, and body proportions.
CONSTRAINTS: No extra accessories, logos, or redesigned clothing.
For a local edit, describe one target region and one operation. If the tool supports a mask, mask the region and tell GPT Image 2 to change only that region. Without a mask, use spatial language such as “the sign in the upper-right” or “the red cup on the left side.” Repeat the preservation list in every follow-up prompt.
CHANGE ONLY: Replace the red sign in the upper-right with a blank cream sign.
PRESERVE: Building geometry, window reflections, people, shadows, camera angle, and color balance.
DO NOT: Recompose the storefront or add new text.
A useful iteration loop is: generate, identify the single largest error, change one variable, and keep the rest of the brief unchanged. If an edit invents details, strengthen the preservation list before changing the target.
What to do when GPT Image 2 refuses a request
Do not disguise a prohibited request with euphemisms, role-play, or instructions to ignore safeguards; remove the risky element or reframe the legitimate creative goal.
| If the request contains… | Use a safer brief… |
|---|---|
| Sexualized nudity or an ambiguous-age person | An adult subject in ordinary, non-explicit clothing; state the age clearly when relevant |
| A living person’s likeness in a sensitive or deceptive scene | A fictional adult character with non-identifying features |
| A copyrighted character or brand identity | An original character described by general visual traits, without logos or trademarked names |
| Graphic injury or sexual violence | Non-graphic aftermath, symbolic storytelling, or a neutral documentary-style scene |
| Instructions to evade moderation | A direct, policy-compliant description of the intended scene |
If a harmless request is blocked, simplify it and remove unnecessary sensitive terms. For example, replace “make this person look nude” with “change the outfit to a plain long-sleeve shirt and trousers,” while preserving pose and lighting. If the platform still refuses, use its support or policy route; no prompt can guarantee approval.
GPT Image 2 and Nano Banana 2: two prompting habits, not a magic syntax
Labeled lines or JSON-like blocks are optional organizers, not a special syntax. OpenAI’s Cookbook accepts minimal prompts, descriptive paragraphs, labeled segments, and JSON-like structures when the intent and constraints remain clear (OpenAI Cookbook). Google’s Nano Banana guidance likewise uses ordinary natural language: subject, action, context, composition, style, and explicit relationships between reference images (Google DeepMind prompt guide).
The useful difference is operational:
| Job | GPT Image 2 habit | Nano Banana 2 habit |
|---|---|---|
| New photoreal image | Use a labeled creative brief and say photorealistic when realism is the goal | Start with subject, action, context, composition, and style |
| Multiple references | Assign each image a role and list the invariants to preserve | Name the relationship between each reference and the new scene |
| Local edit | “Change only X; preserve Y” and use a mask when available | State the edit and preservation rules in natural language |
| Iteration | Change one variable and repeat the preserve list | Iterate conversationally, keeping reference identities named |
The practical distinction is workflow emphasis: start with GPT Image 2 for text-bearing or structured briefs, and test Nano Banana 2 when reference fidelity is your main concern. This is a starting heuristic, not a benchmark result.
Prompt patterns by job
Photoreal product image
Create a 4:5 ecommerce hero image of a matte navy ceramic mug on a pale stone surface. Three-quarter view, mug in the lower-right third, empty upper-left for copy. Softbox from upper left, gentle fill from the right, photorealistic product photography, 50mm feel. Preserve the mug’s cylindrical shape and cream handle. No text, logo, watermark, extra products, or hands.
Text-heavy poster
Create a vertical 9:16 event poster. A lone figure stands on an empty street under one warm lamp at twilight. Main headline reads exactly "THE NIGHT BEGINS AT EIGHT" in large gold serif type across the lower third. Subhead reads exactly "A STORY ABOUT WAITING" beneath it in small white sans serif. Two text lines only; render both verbatim. No extra text, QR code, logo, or watermark.
Local background edit
Change only the background to a quiet beach at sunset. Preserve the person’s face, hair, clothing, pose, expression, camera angle, subject lighting, and framing. Match the new background perspective to the original image. Do not add people, text, logos, or extra objects.
Reference-based outfit edit
Image 1 is the base person. Image 2 is the clothing reference. Replace only the outfit using Image 2’s jacket color, cut, and fabric. Preserve Image 1’s face, hair, body proportions, hands, pose, background, lighting, and framing. Add no jewelry, text, or logos.
FAQ
How long should a GPT Image 2 prompt be?
There is no fixed word count. A useful prompt is long enough to specify the artifact, subject, composition, exact text, and constraints; delete any sentence that adds no visual decision.
How many reference images can GPT Image 2 use?
The fal guide says GPT Image 2 edits can accept up to 16 reference images, but provider limits can differ. Check the endpoint you are calling and assign each image a clear role.
What should I do if GPT Image 2 refuses a safe request?
Remove ambiguous sensitive details, describe the legitimate visual goal directly, and try a fictional or non-identifying subject. Do not ask the model to bypass safeguards.
The practical choice
Before generating, check four things: exact text, subject identity, preserved regions, and unwanted additions. Then write one structured brief, lock the parts that matter, and revise one variable at a time.
For the model’s API availability and current pricing, see GPT Image 2 API pricing. For model selection rather than prompting, use GPT Image 2 vs Nano Banana.