Back to the journal

When AI Optimises Product Photos for Clicks, Check the SKU First

KwaiMind points to a future where AI product edits are optimised for clicks. Sellers still need a strict fidelity gate before conversion testing.

A new research system from Kuaishou Group points to a significant change in AI product imagery. Instead of asking only whether an edit looks polished, KwaiMind also asks whether it is likely to earn a click. That makes the work commercially interesting, but it also exposes a risk: an image can attract more attention while becoming a less accurate representation of the product.

KwaiMind is described in a technical report first submitted on 22 September 2026 and revised on 29 September. It is not a seller-facing app, and the paper does not announce public model weights or a general API. Sellers cannot simply add it to their workflow today. The practical value is in the design direction it reveals, because similar optimisation is likely to appear in more e-commerce imaging products.

What KwaiMind changes

Most generative image editors are trained to follow an instruction and produce a plausible result. E-commerce editing has stricter requirements. A useful image must preserve the identity of the product, render logos and packaging text correctly, and still look appealing enough to compete in a busy feed.

The KwaiMind team built a data engine with about 1.8 million editing pairs and added specialised rewards for click-through rate, product consistency and text rendering. Its Ecom-Bench evaluation covers 11 commercial editing tasks. In the paper's reported results, the system led the evaluated open-source editors overall and achieved the highest aggregate ranking for predicted click-through rate.

The commercially important result comes from selection. Offline optimisation increased the share of generated images whose predicted click-through rate exceeded the original image from 12.16% to 37.41%. In a reported online A/B experiment, selecting product main images with the CTR model produced an approximately 2.44% relative increase in actual click-through rate.

Those numbers need careful reading. The 2.44% figure is a relative increase in one reported experiment, not a promise that every seller will receive 2.44 percentage points more traffic. Results may differ by marketplace, category, audience and baseline image quality. The paper is evidence that commercial image optimisation can work, not a universal performance guarantee.

A higher CTR is not the same as a better product image

CTR measures whether an image earns attention. It does not prove that the shopper understood the product correctly, bought it or kept it. An optimiser can be rewarded for making a pendant appear larger, brightening a gemstone beyond its real colour, shortening a chain to improve the crop or adding a premium-looking accessory that is not included.

The same problem applies to ordinary products. A bottle label can become cleaner but inaccurate. A fabric pattern can shift. A bundle can gain an extra component. A reflective surface can lose a seam, button or engraving. These changes may make a thumbnail more legible, yet they can also create disappointed customers, returns and "not as described" complaints.

A September 2026 Photoroom benchmark illustrates how common these errors can be. In the company's own study of 3,400 generations across 850 products, only 25.3% passed without a fidelity issue. Logo or text distortion appeared in 20.1% of generations, missing elements in 12.5%, and pattern changes in 11.4%. This was a vendor-run benchmark rather than independent research, but it is a useful warning against treating visual quality as proof of product accuracy.

Use three gates, in this order

The safest workflow separates accuracy from performance. Do not allow a CTR score, a preference model or a human reaction to override the factual review.

Gate one is product fidelity. Compare every generated image with a locked, real reference for the exact SKU. Check silhouette, proportions, colour, material, finish, logo, text, included components and condition. For jewellery, count stones and pearls, trace every chain link near the clasp, inspect engravings, and confirm that prongs and settings have not moved.

Gate two is channel and claims compliance. Decide whether the image is suitable as a factual listing image, a secondary lifestyle image or an advertisement. Remove unsupported badges, invented packaging copy and props that imply accessories are included. A dramatic creative may be appropriate for an ad while being unsuitable as an Amazon or marketplace hero image.

Gate three is performance. Only after an image passes the first two gates should it enter a test. Compare approved variants that change composition, background, crop, lighting or scene without changing what the customer is buying.

What to measure after the click

If CTR is the only target, the winning image may simply be the one that creates the strongest curiosity gap. Add conversion rate, add-to-cart rate, revenue or contribution per impression, returns, refunds and product-description complaints to the review.

For example, suppose a brighter gemstone edit raises CTR by 8% but reduces purchase conversion and increases returns because buyers expect a different colour. The thumbnail has won the attention test and failed the retail test. Conversion per impression and post-purchase outcomes reveal that failure.

Small sellers do not need a dedicated CTR model to apply this principle. Create a handful of tightly controlled image variants, approve them against the physical product, and test them through the ad platform or storefront tools already available. Change one major visual variable at a time so the result is interpretable.

A practical workflow sellers can use now

Start with a master set of real photographs for each SKU. Include a straight-on view, a side or three-quarter view, and close-ups of text, texture, clasps or other identity-critical details. Treat these references as the source of truth, not as loose inspiration.

Generate several variations around that fixed product. Useful variables include background colour, crop, surface, shadow direction, lifestyle setting and the amount of empty space for ad copy. Avoid asking the model to improve the product itself. Phrases such as "more luxurious gemstone" or "perfect polished metal" invite changes that are hard to distinguish from fabrication.

Reject fidelity failures before anyone sees performance data. Then test approved secondary images or ad creatives first, where contextual variation is expected. Keep the primary listing image conservative and factual. Record the prompt, source references, edits, channel, audience and result for each variant so a winning idea can be reproduced without copying an accidental product error.

Lustra Studio and similar product-imaging tools fit best inside this controlled process: they can help produce consistent scenes or variants, while the seller retains a real product reference and a separate approval step.

The useful lesson from KwaiMind

AI product photography is moving from generation toward optimisation. Future systems will increasingly choose not just what looks good, but what is predicted to perform. That can save creative-testing time and surface compositions a seller would not have tried.

The order of objectives matters. Product truth comes first, channel suitability second and performance third. A click-optimised image is commercially useful only when it still represents the SKU the customer will receive.

Sources

KwaiMind Technical Report, Kuaishou Group, revised 29 September 2026: https://arxiv.org/abs/2609.26375

Photoroom, "Automate Product Image Quality Assurance at Scale," 14 September 2026: https://www.photoroom.com/blog/automate-product-image-quality-assurance-at-scale

Sources reviewed 9 October 2026.

KwaiMindAI product photographye-commerce AIproduct image testing