Back to the journal

Photoroom Visual Agents Can Check and Retry AI Product Photos Before They Reach Your Catalogue

Photoroom’s new Visual Agents system automatically generates, checks and retries product visuals before publishing. Here is what the launch says about the real bottleneck in AI product photography: fidelity, not image quality.

Photoroom launched Visual Agents on 26 August 2026 with a different pitch from most AI product-photo tools. Instead of promising a better first generation, it builds a loop around the generation: create the image, compare it with the real product, score the result, retry failures and send only borderline cases to a human reviewer.

That matters because the hardest ecommerce problem is no longer making an attractive product image. It is knowing whether the attractive image still shows the product accurately enough to publish.

Photoroom’s own benchmark makes the problem unusually concrete. In testing across 850 real products and thousands of AI-generated images, the best model passed its full product-fidelity check only 29% of the time. Logos were distorted or unreadable in roughly one in five outputs. Photoroom’s additional fidelity layer improved the pass rate to 38.2%, which is better, but still means most raw generations did not pass the company’s full standard.

The interesting launch is the quality-control loop

Visual Agents is built around four stages. First, Visual QA analyses the incoming image and determines what kind of workflow it needs. Create & Transform then generates or edits the visual. Fidelity Raters compare the result with the product and score it against a threshold. If it misses, Visual Fix retries with adjusted settings or a different model. Only uncertain cases are pushed to a human reviewer.

For a retailer with tens of thousands of SKUs, that is a more realistic production design than assuming one model will be reliable enough on its own. Photoroom says one marketplace customer ran more than 1.5 million automated quality checks in three weeks and corrected more than 90,000 images where the system found a defect. Another catalogue run processed 8,200 images overnight with 99.9% completing automatically, according to Photoroom’s launch data.

These are company-reported results, not independent benchmark figures, but they show what the product is trying to solve: the review queue around AI, not just the generation itself.

Why retrying can be more useful than endlessly changing prompts

Photoroom says the same product can fail a fidelity check on one generation and pass on the next without changing the underlying model. That is familiar to anyone who has generated product imagery manually. One result changes a clasp, the next preserves it. One version mangles a logo, another gets it right.

A seller working manually often responds by rewriting the prompt, adding more restrictions and trying again. An automated loop can instead treat generation as probabilistic: produce a candidate, test it against objective criteria, retry when it fails, and keep the human for the cases where the system cannot decide confidently.

That is an important shift. The production question becomes less about finding one perfect prompt and more about building a reliable approval process around imperfect models.

Jewellery shows why visual QA matters

Photoroom’s benchmark included jewellery and accessories, which is useful because jewellery exposes small fidelity failures quickly. A bracelet can remain recognisable while one clasp changes. A necklace can keep the same overall silhouette while a connector disappears. An engraved pendant can look excellent until one letter has been rewritten.

Those are exactly the kinds of errors a human can miss when judging the image as a whole. A structured checker can force the review to focus on product-specific facts: silhouette, colour, logo or engraving, repeated patterns, missing elements and altered geometry.

For jewellery sellers, the practical lesson is to define the QA checklist before generating at scale. For a bracelet, that could include clasp type, chain construction, plate dimensions, engraving, stone count, metal finish and overall proportions. For a packaged product, it might include label wording, cap shape, colour, quantity and logo.

Do not misread the 38.2% number

The 38.2% figure is not a claim that Visual Agents only delivers 38.2% usable images. It is the reported pass rate from Photoroom’s fidelity layer in a benchmark stage before the broader retry-and-fix loop has done its work. Visual Agents is designed to keep retrying failed outputs and escalate difficult cases instead of publishing the first result.

It is also important that the benchmark and performance figures come from Photoroom itself. They are valuable because the methodology is more specific than generic “looks good” comparisons, but sellers should not treat them as an independent ranking of every image model.

The more useful conclusion is that even Photoroom, a company selling AI product-imaging infrastructure, is explicitly saying that base-model quality is not enough. It has built additional checking, scoring and correction layers because a plausible image can still be commercially wrong.

This is enterprise-only, and the current fidelity raters have limits

Visual Agents is not a self-serve feature for small sellers. Photoroom currently offers it to enterprise customers with pricing based on volume. Its product page also says the category-specific fidelity raters available today are built for food and fashion, although the wider analysis layer can check image quality, cropping, safety and content across other catalogues.

That means a jewellery seller should not assume the full jewellery-specific automated scoring loop described in the concept is already available off the shelf. Jewellery was included in Photoroom’s benchmark, but Photoroom’s current public product page lists food and fashion as the live category-specific rater coverage.

This limitation is important because category-specific QA is where small details matter most. A general image checker can tell you that an image is sharp and well framed. A jewellery-aware checker needs to understand whether the clasp, stones, chain and engraving still match the SKU.

Small sellers can copy the process without buying the enterprise system

The Visual Agents idea is still useful even if you never use the product. Keep one real product image as the source of truth. Generate the derivative. Compare the candidate against a fixed product checklist. Reject or regenerate anything with a factual change. Only then judge the scene, styling and overall attractiveness.

A multimodal vision model can also act as a second reviewer by comparing the generated image with the source and flagging possible differences, although that checker should not replace a final human inspection. For higher-volume workflows, the same logic can eventually be automated through APIs.

For Lustra Studio, this launch reinforces the value of keeping jewellery references central to the workflow. Generation quality will continue to improve, but the scalable advantage is likely to come from pairing generation with structured references, repeatable checks and automatic rejection of outputs that drift from the real product.

The takeaway

Photoroom Visual Agents is interesting because it treats AI product photography as a production system rather than a single image-generation call. The model creates the candidate, another layer checks fidelity, failed outputs are retried, and humans are reserved for the genuinely uncertain cases.

For product sellers, that is probably the more important direction to watch. Better image models will reduce errors, but catalogue-scale AI needs a process that assumes some errors will still happen and catches them before customers do.

Sources: Photoroom, “Photoroom launches Visual Agents for enterprise catalogs,” published 26 August 2026; Photoroom, “Nobody returns a photo. They return the product,” published 23 July 2026; Photoroom Visual Agents product page, reviewed 11 September 2026.

PhotoroomAI Product PhotographyProduct FidelityE-commerce AIVisual QAProduct Images