BelltoAn AI agent that takes a brand's POP campaign from brief to shop drawings.The AI agent for POP campaigns.Get on the waitlist
← back to blog

Why Your AI Render Changes Details Between Versions

October 4, 2026··9 min read

The client approved version three. You asked for a pharmacy background instead of a supermarket, and version four came back in the pharmacy with five shelves instead of four, a rounded header where the approved one was square, and a logo whose second letter had quietly changed weight. Nobody asked for any of that.

This is the complaint I hear most about AI renders. It isn't bad luck, and it isn't only a bad prompt. It's how the models work, and once you see the mechanism the fixes stop being guesswork.

I've worked in point-of-purchase for close to four years and I build AI POP Displays. This post is the model side of the problem. The client side, where each round of feedback becomes an edit, I covered in how AI changes the client revision round. Here the question is narrower. Why does the same display come back different, and what holds it?

Why does an AI render change things you didn't ask to change?

Because there is no display inside the model. There's a picture.

A 3D file stores an object. Change the background in a 3D scene and the shelves don't move, because they have coordinates. An image model stores nothing of the kind. When you edit a render, the model takes the previous image as one of its inputs, reads your instruction, and paints a new image from scratch that it judges to be the most plausible answer to both. The previous version is guidance, not a constraint.

So anything the instruction and the references don't pin down is up for grabs. "Put it in a pharmacy" pins the background and says nothing about the shelf count, so the count gets re-decided. Usually the model lands on four again. Sometimes it draws five.

The people who build these models say this openly. The research paper behind FLUX.1 Kontext, from Black Forest Labs, frames the problem as "noticeable visual drift", where characters lose identity and objects lose features with each edit. That paper's fix is to slow the drift. Its own discussion section says excessive multi-turn editing can still introduce visual artifacts that degrade the image. Google's documentation for Nano Banana Pro talks about keeping up to six objects at high fidelity and up to five characters consistent, and its own example prompt for an edit spells out that nothing else in the image should change. Consistency, in every vendor's own words, is something you ask for, not something the model guarantees.

Why is drift worse on a display than on a product shot?

Because a display is mostly the part the model treats as background.

Most consistency work in image AI is about faces and characters. Search for help with AI render consistency and you get character sheets, product videos, interiors and seeds for architectural renders. Nothing about a fixture that has to keep its geometry because someone is going to quote it.

A display has three properties that make drift expensive.

It's repetitive. Four shelves of identical packs, eight facings each. The model reads repetition as texture, and texture is the first thing it resamples. That's why the count is the detail that moves most.

It's dimensional. A header that is 30% of the unit's height in version three and 40% in version four is a different piece. A face that shifts 10% still reads as the same person. A fixture doesn't.

It's branded. A logo is a specific drawing in a specific color. Every redraw is a fresh approximation, and approximations drift letter by letter.

And a buyer compares versions side by side. A changed shelf count is a different quote. If the renders on your desk keep moving between rounds, run the same brief through the free plan and compare what holds.

Why does it get worse with every edit?

Because each edit starts from a copy of a copy.

When version five is made from version four, which was made from version three, every pass redraws the whole image from the last output. Small losses don't reset. A slightly softer edge in version four is the input for version five, which softens it again.

There's now a paper that measures exactly this. EdiVal-Agent, published on arXiv, runs sixteen commercial and open-source editing models through three successive edits and scores instruction following, consistency with the untouched parts of the image, and visual quality. All sixteen score lower at turn three than at turn one. The leaders open at around 84 on its overall score and finish below 60 by the third edit, Seedream 4.0 going from 83.81 to 59.76 and GPT-Image-1.5 from 84.29 to 59.55. The authors' explanation is the one that matters here. Most editors are trained to work on real photographs, not on their own previous outputs, so when they edit their own generations, small mismatches compound turn after turn. They also report luminance drifting across turns, which on a display is the brand color going a shade darker every round.

On a display the chain looks like this. Round two changes the background, round three adds the campaign artwork, round four asks for landscape. By then the base reads wider, the sharp miter has gone round, and the top-shelf pack is a different shade. None of the instructions asked for any of it.

The structural fix is not to chain. Every change should branch from the version the client approved, not from whatever the last edit produced. In AI POP Displays every edit starts from an existing version you choose, and lands as a new version next to it, so round four can come straight off round two and the approved image is never overwritten.

Does a fixed seed solve it?

Not for this problem.

A seed fixes the random noise the model starts from, so the identical prompt with the identical seed gives the identical image. The moment the prompt changes, which is the entire point of a revision, the image changes too.

Rewording has the same limit. "Keep everything else the same" helps, but it doesn't name the shelf count, so the shelf count is still free.

What actually holds a display's geometry between versions?

Inputs that carry the geometry, every time, so the model is never asked to remember it.

A saved brief. The format, the material and the product load are fields, not adjectives buried in last week's prompt. In AI POP Displays the format comes from a taxonomy that carries its proportions, so a floor unit starts taller than it is deep with a base under its load, and every regeneration starts from the same fields.

The real packs. Product photos uploaded as the product, placed at scale and repeated as consistent units facing forward. The pack is the most stable scale reference on a display, because it's the same photograph in every version.

References labeled by role. Every reference travels with a written label in front of it, logo, product photo, campaign artwork, style reference or sketch, and the sketch goes last. A logo labeled as a logo is something to reproduce, not a mood to borrow from.

A sketch at the faithful setting. If there's a drawing, it's the strongest lock on structure. At the faithful setting the structure, silhouette and shelf count follow the sketch.

An edit that states what stays fixed first. Every edit in our console goes to the model with the same list before the change: the structure, materials, dimensions and proportions, the brand identity, the products and their facings, the light direction and the camera, unless the operation is specifically a new camera. Then it states the one thing that changes. That list is what the client already approved without saying so.

Done once in the tool, that's the difference between asking the model to remember a display and handing it the display every time. AI POP Displays runs on Google's Nano Banana Pro, and the layer on top, the format taxonomy, the materials, the labeled references, the preservation list on every edit, is close to four years of watching which details clients catch, built into the tool once.

If your last three versions all came from chained prompts in a chat window, start the next round from a saved brief and branch every edit from the approved image.

How do you check a new version before the client sees it?

Side by side with the approved one, on the things that drift.

  1. Count. Shelves and facings per shelf. Do it on a flat front view if you have one, because a three quarter view stacks the shelves and hides an extra one.
  2. Proportion. Header height against total height, base width against shelf width. Use the pack as the ruler, as I laid out in how accurate are AI display renders.
  3. Edges. Miters, radii and material thickness. Corrugated that turns into acrylic at the edge is a drift, not a style choice.
  4. Brand. The logo at full size, not the thumbnail, letter by letter.
  5. Product. The right SKU on the right shelf, in the right color.

Compare in AI POP Displays puts any two versions side by side at no cost in credits, which is what this check is for. If a version drifted, don't fix it by editing the drifted version. Go back to the approved one and run the edit again from there, with the drifted detail named as fixed. And if what's changing is the format, the material family or the product range, it isn't drift, it's a new concept, and it starts from the brief. I went through that triage in my AI display render came out wrong.

Where does an image stop being enough?

At the point where the detail that has to hold is a number.

A picture can be steered to keep a shelf count and a proportion, and a careful workflow keeps them across many rounds. What it can't keep is a dimension, because it never had one. The header isn't 300 millimeters tall in version three, so there's nothing for version four to preserve. That's why the vendors write consistency and resemblance, never identity, and why holding geometry across a dozen versions is still an open problem for every image model, not a solved one.

Geometry that holds by construction needs a model of the object. That's the half we're building now. Bellto is an AI agent for brands and agencies that takes a POP campaign from brief to shop drawings. You tell it what you're launching (brand, product, channel, stores, budget, date) and it asks for what's missing. It proposes the campaign mix and materials, floor stand, glorifier, shelf strip, stopper, with the budget split per store, and designs every piece with you at real scale in parametric 3D, proportions derived from the product and the facings. Nothing drifts there between versions, because a dimension change recomputes the whole piece and re-runs the manufacturing checks in seconds. Every approved piece comes out as a 3D model, a part-by-part cutlist, dimensioned drawings, a STEP of the assembly and a cut DXF per part, a package a workshop can quote without redrawing, and it suggests manufacturers that fit by material, format and volume. The quote comes from the manufacturer.

We're building it now. The waitlist is open to any brand, and we're contacting the first ones soon to run the first real campaigns end to end. It opens by invitation, in small groups. Pricing goes first to the people on the list, and there's no card. POP manufacturers can sign up too.

For the concept and the rounds that follow it today, sign up free with 15 one-time credits and no card required, and branch your next revision from the version the client approved. Pro is $49 a month for 150 concept generations at the founding price, with the $69 list price stated openly. Nothing you upload or render trains AI models, on any plan.

Frequently asked

Why does my AI render change details I didn't ask to change?

Because the model redraws the image instead of editing an object. An image model keeps no geometry between versions, only the previous picture as an input. Everything your instruction and your references don't pin down is generated again, so on a display the shelf spacing, the edge profile, the number of facings or the letterforms of a logo can move even when you asked for a new background. Researchers who build these models describe the same effect as visual drift, where objects lose features with each edit.

Does drift get worse the more times I edit an AI render?

Yes, when each edit starts from the previous edit. Every pass re-encodes and redraws the image, so small losses stack, and the published work on multi-turn editing reports that consistency and quality degrade over successive turns. The fix is structural rather than verbal. Branch every change from the approved version instead of chaining edits, compare each new version with the approved one, and go back to the brief when the thing changing is the display itself.

Can a fixed seed keep an AI display render consistent?

Only between runs of the same prompt with nothing else changed. A seed fixes the starting noise, so changing the instruction, adding a reference or switching the camera produces a different image anyway. On a display, what holds the geometry is the inputs that carry it: a saved brief with the format and the product load, the real pack photos, a sketch at the faithful setting, and an edit that states what stays fixed before it states what changes.

Will AI renders ever keep a display's geometry exactly between versions?

Not from a picture, because a picture stores no dimensions to keep. Exact, repeatable geometry comes from a model of the object. That is what we are building with Bellto, an AI agent for POP campaigns that designs every piece at real scale in parametric 3D, where a dimension change recomputes the whole piece and re-runs the manufacturing checks in seconds. It is open by invitation from a waitlist.


Start generating Get on the Bellto waitlist
Why Your AI Render Changes Details Between Versions | AI POP Displays