Back to the journal

How Many Products Can Nano Banana 2 Actually Keep Consistent?

Google’s launch post mentions 14 objects, but its current API guide gives Nano Banana 2 a ten-object high-fidelity limit. Here is how sellers should plan multi-product images.

Google’s Nano Banana 2 launch post makes an eye-catching claim: subject consistency for up to 14 objects in one workflow. A product seller could reasonably read that as permission to upload 14 SKUs and generate a reliable collection image.

The current Gemini API guide is more specific. It says Nano Banana 2, officially Gemini 3.1 Flash Image, supports high-fidelity inclusion of up to 10 object images. Its remaining reference capacity can be used for up to four character images. A separate model, Nano Banana 2 Lite, is the one listed for up to 14 high-fidelity object images.

That distinction matters. A reference limit tells you how many inputs a model can accept or classify for a task. It does not promise that every clasp, label, stone setting or colour will survive the final composition. Sellers should plan around the current model table, then test below the stated ceiling.

Why the numbers appear to disagree

When Google launched Nano Banana 2 on 26 February 2026, its announcement said the model could maintain the resemblance of up to five characters and the fidelity of up to 14 objects. The same post showed a playful farm scene containing 14 characters and items.

Google’s current developer documentation now separates the reference capacity by model. Nano Banana 2 Lite can take up to 14 object images for high-fidelity inclusion. Nano Banana 2 can take up to 10 object images plus up to four character images. Nano Banana Pro can use up to six object images, five character images and three style references.

All three can therefore work with as many as 14 reference images, but those slots are not equivalent. The category of each reference matters, and the model with the largest object count is not automatically the model with the strongest professional control.

For a seller building an API workflow, the current developer guide should be the operational source of truth. The launch post is useful context, but it does not describe the later model-family table in enough detail to design a catalogue process.

What “high fidelity” does not guarantee

High-fidelity reference support is a capability boundary, not an acceptance test. Google does not promise that ten product inputs will emerge with every physical detail unchanged. The output still has to reconcile composition, lighting, scale, overlap and the written prompt.

Similar products make the task harder. Three nearly identical cuffs in silver, gold and rose gold can be easier to confuse than three visually distinct objects. A model may swap a finish, duplicate one variant, remove a small charm or simplify a chain even when the total object count is within the documented limit.

Occlusion is another problem. A group scene can hide the feature needed to identify a SKU. If one bottle covers another label or one necklace overlaps a pendant, the generated image may look coherent while failing as product evidence.

Resolution does not fix identity errors. Nano Banana 2 can output from 512 pixels to 4K, but a sharper image can still show the wrong clasp or an invented logo. The Lite model is limited to 1K output, according to the API guide, which is another reason not to choose it solely because its object-reference count is higher.

A safer way to build a multi-product image

Start with one clean reference per SKU. Use an isolated or simple-background photograph that shows the product’s distinguishing features. Do not place several products in one reference image and then count that file as a single object. That makes it harder to tell which details the model has associated with each item.

Give every reference a stable identifier in the brief, such as Product A, Product B and Product C. Pair the identifier with a short factual description: “Product B is the rose-gold cuff with one round clear stone and no engraving.” Then specify the intended position of each item in the final frame.

Begin with three products, not ten. Generate the same type of composition several times and record which details fail. Move to five references only when the three-product set passes consistently. Continue in small steps until the error rate becomes unacceptable for the intended channel.

Keep the prompt visually simple while testing identity. A neutral surface, soft studio light and limited overlap make failures easier to spot. Complex props, hands, reflections and dramatic depth of field should come later, after you know the model can keep the products separate.

For each output, check four things before judging style: every requested SKU is present, no SKU is duplicated, the products are in the assigned positions, and each item retains its defining geometry and colour. Only then review the scene, lighting and crop.

Where a ten-product workflow is useful

Multi-reference generation can be useful for collection banners, gift-guide concepts, seasonal campaign drafts and cross-sell imagery. It can help a seller explore how several items might sit together before arranging a physical shoot.

It is less suitable as automatic proof of an exact bundle. If a customer is buying a set of six items, the image must show the six real items accurately. A conventional composite made from approved packshots is usually safer than asking a model to redraw the complete bundle.

Amazon main images and detail-sensitive marketplace listings deserve particular caution. A generated collection scene may work as secondary creative if it accurately represents the products and follows the platform’s rules, but the primary evidence should remain a real photograph or controlled composite.

Jewellery needs a lower practical ceiling

Jewellery sellers should treat the documented limit as a maximum, not a target. Fine chains, small stones, prongs, engravings and near-identical metal variants create more failure opportunities than large, distinct objects.

For a collection of necklaces, use separate references for each pendant and include a close view of the distinctive construction where possible. If a chain style must remain exact, verify the links rather than assuming the model will transfer them because the pendant looks correct.

A useful internal rule is to reduce the reference count whenever two products could be mistaken for variants of the same SKU. Generate smaller groups, approve them independently, and assemble the final collection image from those approved elements if necessary.

The headline number is therefore not “14 products.” It is a set of different reference budgets across a model family. Keep the number of SKUs below the relevant limit, measure fidelity with your own products, and use a real composite whenever every item must be unquestionably exact.

Sources

Google, “Nano Banana 2: Combining Pro capabilities with lightning-fast speed,” published 26 February 2026 and reviewed 5 October 2026: https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/

Google AI for Developers, “Gemini API image generation,” reviewed 5 October 2026: https://ai.google.dev/gemini-api/docs/image-generation

Google, “Start building with Nano Banana 2 Lite and Gemini Omni Flash,” reviewed 5 October 2026: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/

Google AI for Developers, “Gemini Developer API pricing,” reviewed 5 October 2026: https://ai.google.dev/gemini-api/docs/pricing

Nano Banana 2Gemini ImageMulti-Reference ImagesProduct ConsistencyAI Product Photography