← volver al blog

Why Generic Image AI Fails at POP Displays

August 20, 2026·Arturo Bellot·9 min de lectura

Este post aún no está traducido. Mostrando la versión en inglés.

A designer at a display manufacturer showed me a folder of AI renders this summer. Twelve images of what was meant to be a cardboard floor display for a beverage brand. All twelve looked great at thumbnail size. All twelve fell apart the second you looked at them as something a factory has to build. Shelves floating with no side panel carrying them. A header card in a proportion no die line would ever cut. The brand's logo, uploaded as a reference, rendered as a plausible-looking word that was not the brand name.

She hadn't done anything wrong. She'd written a careful prompt, attached references, and iterated. The tool did what it's built to do, and what it's built to do is not this.

I've spent close to four years in point-of-purchase and I build AI POP Displays, so read this knowing where I stand. I'm not going to argue that general image models are bad, because they aren't. They're remarkable, I open them most weeks, and one of them is what our own product runs on. What I want to explain is why they fail at this particular job in a way that more prompting doesn't fix, and where they're still the first thing you should reach for.

What is a general image model actually doing?

A quick note on which models this is about, because the field turns over fast and half the advice online is describing tools that no longer exist. ChatGPT runs on ChatGPT Images 2.0, which arrived in April 2026 and replaced the DALL-E line entirely. Midjourney is on V8.2 as of late July 2026. Google's current pair is Nano Banana 2 and Nano Banana Pro. These are all strong models, and the criticism below is not the tired one about AI images looking obviously fake. They don't any more.

A model like these learned from an enormous amount of imagery, including plenty of retail photography. It knows very well what a point-of-purchase display looks like in a photo. It knows almost nothing about how one is constructed, because construction isn't visible in the training data. Board thickness, fold lines, joints, load paths, print area, pallet footprint. None of that appears in a JPEG of a supermarket aisle, so none of it gets learned.

What the model optimizes for is an image a person accepts as real. That's a different target than an object that resolves as a structure, and the gap between the two is exactly where display work lives.

The consequence that matters most is that these models have no failure state. The model never comes back and says the geometry doesn't close. It comes back with a picture, at the same confidence, whether the picture is correct or nonsense. The errors it makes are the ones that read fine on screen and cost money in a quote.

Why does the geometry come back unbuildable?

This is the failure people notice last, because it hides well at small sizes. Zoom in and you start finding shelves cantilevered off a surface that isn't attached to anything, side panels that are three millimeters thick at the top and twenty at the bottom of the same panel, a base footprint narrower than the mass sitting on it, and corrugated structures with no fold, no tab, no visible way of going from flat sheet to standing unit.

None of that bothers the model, because photographically it's fine. Light behaves, shadows land, materials read. It only breaks when a human at the manufacturer opens the file and has to work out how to actually make it, at which point they redesign the structure and the render turns out to have been a mood image rather than a concept.

That distinction is the whole game. A concept is something a client approves and a factory quotes. A mood image is something that gets a nod in a meeting and then gets rebuilt from scratch.

Why are the format proportions wrong?

Point-of-purchase runs on named formats, and each name carries real dimensional constraints. An FSDU has to sit on a pallet footprint and survive shipping. An endcap has to fit the gondola end of a specific retailer. A counter unit has to leave the cashier room to work. A glorifier has to hold one hero product at eye height without dominating the shelf around it.

A general image model treats those words as style, not geometry. Ask for an FSDU and you often get something with the visual language of a floor display at proportions no floor display has ever had. Ask for a counter unit and it may come back at chest height. There's usually no reliable scale cue in the frame either, so the client looking at it can't tell, and neither can you, until someone puts a dimension on it.

This is the one that quietly wastes the most time, because the image passes internal review, goes to the customer, gets loved, and then dies when the drawing comes back and it doesn't fit the space it was sold into.

Why does the brand logo come back warped?

Text rendering in image models has improved enormously, and it's now something the vendors advertise rather than apologise for. People conclude from that improvement that brand assets are safe. They aren't, because reproducing a brand mark isn't a text problem. It's an exact-reproduction problem, meaning a specific mark, in a specific color, sitting at an angle across a folded or curved surface, at whatever resolution that corner of the image happens to get.

Read what the vendors actually promise and the gap is visible in their own words. Google caps high-fidelity object references at six images and describes the result as consistency and resemblance, never identity. OpenAI's limitations note says labels still need checking, and that the model struggles with details that have to appear correctly on hidden, angled or reversed surfaces. That sentence is a precise description of a logo wrapping a curved pack front or sitting on an angled header card. Midjourney is blunter still, in that its own documentation lists putting a company name on a logo among the prompts that don't work well.

So what the model does is redraw the mark approximately. Letterforms shift, the kerning goes soft, a symbol grows an extra stroke, the color lands near the brand color rather than on it. A logo that's almost right is worse than no logo at all. It's the one thing a brand manager is trained to catch on sight, and it turns a concept review into a conversation about your process.

The same applies to the packaging. A general model will happily draw a bottle that resembles the client's bottle without being the client's bottle, and the client will notice.

Why is the shelf count fantasy?

A display exists to hold product, so the product load is not decoration. It's the spec. This is where general models drift furthest, because they place as many facings as look good in the composition rather than as many as physically fit.

You get nine facings across a shelf that takes five. You get packs floating a centimeter clear of the surface. You get large and small SKUs rendered at the same apparent size, because the model is balancing a picture rather than merchandising a fixture. Sometimes the product photo you attached is ignored entirely and a generic stand-in gets drawn in its place.

Downstream, that's the expensive one. Capacity is part of what your customer is buying. If the concept implies a load the unit can't carry, either the sales conversation was based on a number that doesn't exist, or the structure has to grow to reach it, and the price grows with it.

Can't you just prompt harder?

Up to a point, yes. With enough reference images, a long prompt and some patience, you can pull one genuinely good display image out of a general model. I've done it. Anyone who works at this seriously has done it.

The trouble arrives with the second image. There's no persistent object underneath the render, only a distribution of plausible pictures that your prompt is steering. So when the client comes back and asks for the header twenty percent taller and the material changed to acrylic, you regenerate, and the shelf spacing moves, the base changes, the product load reshuffles and the lighting shifts. You fix that, and something else drifts. Every revision round becomes a fresh negotiation with the model instead of an edit to a design.

That's why the cost of using general models here isn't the first render. It's rounds two through six, which is where the hours in this business have always gone.

When is a generic model still the right tool?

Often, and I'd rather say so plainly than pretend otherwise.

They're excellent for mood boards and art direction, for exploring a visual territory before the brief is fixed, for color and material atmosphere, for environment and background plates, and for the "show me twenty directions" session at the start of a pitch. They're fast, they're cheap, and their looseness is a feature when nobody has committed to anything yet. Krea, Midjourney and ChatGPT all earn their place in that slot, and a purpose-built tool is worse at it than they are.

The line I'd draw is about what the image has to do. If its job is to communicate a feeling, use a general model. If its job is to communicate an object that someone will quote, engineer and build, you need a tool that knows what the object is.

What changes with a purpose-built tool

Here's the part worth being blunt about, because it's easy to mistake for a claim about model quality and it isn't one. A purpose-built tool doesn't have a better image model. Ours runs on Google's Nano Banana Pro, rented, and you could rent the same engine this afternoon. Nothing in the paragraphs above gets fixed by better pixels, because none of it was a pixel problem.

What changes is that the failures stop being prompt guesswork and start being inputs. Display format comes from a real taxonomy with proportions attached, so the word FSDU carries geometry rather than vibes. Materials are a choice from a list that behaves correctly, not an adjective. Product load is built from the packs you actually uploaded. Retail context is a background you select, so the fixture lands in a supermarket aisle or on a pharmacy counter rather than a studio sweep. References travel labeled, so the model is told which image is the product, which is the brand mark and which is the structure to follow, with the sketch last because that's what it anchors on hardest.

Every one of those is domain knowledge rather than technology. A designer who already carries all of it in their head can hand-prompt their way to a similar result, and some do. The difference is that they supply it again on the next brief and on every revision round, and it's encoded here once. That encoding is what four years in this industry bought.

It isn't magic and it doesn't produce production files. No AI render on the market does. Structural CAD, dielines, board grades and print specs are still human work at the manufacturer, and the render's job is to be the brief that work starts from.

I've written the specific head-to-heads separately, against ChatGPT and against Midjourney, and there's a wider map of the tools people actually use for this if you're deciding what belongs in your stack.

The test that matters

One question sorts good display renders from pretty ones. Could your manufacturer quote from this without redesigning the structure?

Run it on your last AI render. Every shelf lands on something that carries it, or it doesn't. The proportions match the format you asked for, or they don't. The facings match what really fits, or they don't. None of that is subjective, which is why it's worth more than an opinion about which model looks nicest.

If the brief on your desk this week is a real display, signup is free with 15 one-time credits and no card required. Pro is $49 a month for 150 concept generations at the founding price, with the $69 list price stated openly. Nothing you upload or render trains AI models, on any plan, because most of this work sits under NDA and that shouldn't be a paid feature.

Frequently asked

Why do ChatGPT and Midjourney get retail displays wrong?

Because they were trained to produce images that look convincing, not objects that resolve as structures. They learned what a display looks like in a photograph, not how one is cut, folded, joined and loaded. So the output is photographically plausible and structurally undefined: shelves with nothing holding them, panels that change thickness, a base narrower than the load above it. The model has no failure state, so it never tells you the geometry doesn't close.

Can I fix it with better prompts and reference images?

You can usually get one good image that way. The problem is the second one. There's no stable object underneath the render, only a distribution of plausible pictures, so when a client asks for a taller header and a material change, everything else drifts too. The cost of general models on display work isn't the first render, it's revision rounds two through six.

Is a generic image model good for anything in POP display work?

Yes, and I use them for it. Mood boards, art direction, color and material atmosphere, environment plates, and wide exploration before a brief is fixed. The dividing line is what the image has to do. If its job is to communicate a feeling, a general model is the right tool. If its job is to communicate an object someone will quote and build, it isn't.

How do I know if an AI display render is actually buildable?

Ask whether your manufacturer could quote from it without redesigning the structure. Check three things first: does every shelf visibly land on something that carries it, do the overall proportions match the named format you asked for, and does the product load match how many facings really fit. No concept render is a production file, ours included, but a good one survives all three questions.


Empezar a generar
Why Generic Image AI Fails at POP Displays — AI POP Displays