Back to the journal

The Product-Drift Test for Multi-Turn AI Image Editing

A single prompt is easier to compare, while staged edits are easier to steer. Use this controlled test to find where multi-turn AI product editing begins to drift.

A seller asks an AI image tool to place a silver pendant on black stone, change the lighting to a soft side light and add a small ribbon in the background. There are two obvious ways to do it. Put every instruction into one prompt, or build the image in several editing turns.

The staged route sounds safer because each request is smaller. It is also riskier in one important way: every accepted AI output can become the input for the next edit. A slightly altered chain, softer engraving or warmer silver tone can be carried forward and changed again.

Neither route is automatically better. The useful question is whether the extra control of multi-turn editing outweighs cumulative product drift for your SKU. A small controlled comparison can answer that before you apply the workflow to a whole catalogue.

What current image tools actually support

Google’s current Gemini image documentation recommends multi-turn conversation for iterating on images. Its example changes the language of an existing graphic in a later turn while asking the model not to alter anything else. Google also notes that its cheaper Nano Banana 2 Lite model is not optimised for multi-turn sequential editing, while Gemini 3.1 Flash Image is positioned for that kind of work.

OpenAI’s current prompting guide also recommends starting with a clean base prompt and refining through small, single-change follow-ups. It tells users to repeat the preserve list on every iteration and to state both what should change and what should remain fixed.

These are confirmed workflow capabilities, not proof that a product will stay exact through several turns. “Keep everything else the same” is an instruction to a generative model, not a pixel lock.

Why one prompt can be the safer baseline

A one-shot edit gives the model one opportunity to reinterpret the source product. That does not guarantee accuracy, but it avoids feeding a generated product back through the model several times. It is also easier to compare costs and outputs because every candidate begins from the same real photograph.

The weakness is prompt competition. Background, lighting, camera framing, props and preserve instructions all need attention at once. If the final image is wrong, it can be difficult to tell which instruction caused the failure. Regenerating may fix the scene while changing the product in a different way.

One prompt is usually a sensible baseline when the scene brief is already clear and the product contains fragile details such as fine chain links, small labels, prongs, stitching or a distinctive reflection pattern.

Why several smaller edits can be easier to steer

Multi-turn editing lets you approve the composition before requesting a lighting change, then approve the lighting before adding a prop. A failed step is easier to identify. You can also stop after a good intermediate image instead of asking the model to solve the entire brief again.

The drawback is cumulative change. Imagine that the first turn makes one link in a necklace slightly thicker. The second turn sees that thicker link as part of its reference. By the third turn, the product may look internally consistent but no longer match the SKU.

Staged editing also costs more time and may consume more credits or input tokens. OpenAI’s documentation notes that edit requests containing reference images use image input tokens, so extra turns are not free simply because each instruction is shorter.

Run the two-route comparison

Choose one product that exposes the weaknesses you care about. A pendant with a visible clasp and engraving is more informative than a plain solid object. Keep the original photograph as the source of truth.

Write one final brief with three changes. For example: replace the white sweep with dark stone, create soft window light from the left, and add a blurred gift ribbon behind the product. Add a preserve list covering exact shape, dimensions, chain path, clasp, engraving, metal colour, camera angle and product scale.

Route A is the single-prompt version. Ask for all three changes at once. Route B is the staged version. Make the background change first, inspect it, then continue from the approved output for lighting, and finally add the ribbon. Repeat the complete preserve list at every turn.

Keep the model, resolution, aspect ratio and output count the same. If the tool offers a fixed seed, use it where the comparison permits. Produce the same number of final candidates for each route. Four final images per route is a manageable starting point, but the important rule is equal treatment.

Save every intermediate result. Without those checkpoints, you cannot tell when the drift entered the sequence.

Score the product before judging the scene

Review the final images against the real master in separate passes. First check silhouette and geometry. Then inspect small factual details such as stone count, chain links, clasp shape, logo spelling, seams or control placement. Next compare colour and material behaviour, including reflections and highlight width. Only then judge whether the requested scene, lighting and prop work visually.

Use a simple pass, borderline or fail score for each category. Also record how many generations, credits and minutes each route required. The prettiest image should not win if it quietly turns a sale item into a different product.

A difference overlay can help locate changed edges, although lighting and background changes will create many legitimate differences. Use it as a pointer for inspection, not as an automatic accuracy score.

How to interpret the result

If the one-shot route reaches the brief with equal or better product fidelity, its simpler production path is the practical choice. If the staged route produces a much better composition while every checkpoint passes inspection, the extra control may justify the cost.

If drift appears at the second or third turn, do not keep asking the model to repair its own altered output. Return to the last verified checkpoint, or to the real master, and try a narrower edit. A clean restart is often cheaper than adding more corrective turns to a damaged branch.

The most reliable production workflow may be hybrid. Use one controlled AI pass for the major environmental change, then handle cropping, text, logos and simple placement in a conventional editor where pixels can be locked. Tools such as Lustra Studio can support reference-led generation, but the same checkpoint discipline still matters when an exact SKU is being sold.

The practical conclusion

Multi-turn editing is valuable because it makes complex creative work easier to direct and diagnose. It is not automatically more faithful. Every new generative turn is another opportunity for the product to move away from the original evidence.

Treat the real photograph as the master, keep intermediate checkpoints and compare a staged route with a one-shot baseline on one demanding SKU. The winning workflow is the one that reaches the brief with the fewest unapproved product changes, not the one that uses the most sophisticated conversation.

Sources

OpenAI, “GPT Image Generation Models Prompting Guide,” reviewed 4 October 2026: https://developers.openai.com/cookbook/examples/multimodal/image-gen-models-prompting-guide

OpenAI, “Image generation,” reviewed 4 October 2026: https://developers.openai.com/api/docs/guides/image-generation

Google AI for Developers, “Gemini API image generation,” reviewed 4 October 2026: https://ai.google.dev/gemini-api/docs/image-generation

AI Product ImagesMulti-Turn EditingProduct FidelityProduct PhotographyE-commerce