AI POP Displays vs Midjourney for Retail Display Concepts
Este post aún no está traducido. Mostrando la versión en inglés.
I've spent close to four years in point-of-purchase, and a good stretch of that trying to make general image models produce something I could put in front of a customer. Midjourney got more of those hours than anything else, because in 2023 and 2024 it was clearly the best image model in existence and it seemed obvious that the best model would produce the best displays.
It didn't. It was consistently the worst of the major tools for this specific job, and I want to explain why in a way that's useful rather than just negative, because the reason is interesting and it's the reason this product exists.
For the comparison with ChatGPT, which is a much closer contest, see AI POP Displays vs ChatGPT. For the workflow context, the AI POP display generator guide.
First, the thing I'm not going to argue
Midjourney is a superb image model. Its sense of light, composition, material and mood is still the best in the field, and V8.2, the current version since late July 2026, is the best-looking work it's ever produced. For people, landscapes, atmosphere, editorial imagery and concept art, I'd reach for it over anything else.
That's a real endorsement and it comes with a boundary. Being the most tasteful model in the world does not help you when the task is a fixture with a pallet footprint, a header card at a printable proportion, and a client's actual logo on it. Taste is not the constraint in display work. Knowing the object is.
Where it actually breaks on POP briefs
These aren't impressions. Each one is a documented limit of how the tool works.
There is no way to place your client's logo on the fixture. This is the one that ends it. Midjourney has three ways to point at an existing image and none of them inserts an asset unchanged. Style reference transfers colour and texture, and the documentation says plainly that it doesn't copy objects. Character reference, which used to carry a subject across images, has been retired. What's left is omni-reference, which takes one image per prompt, costs double the generation time, and silently drops you from V8.2 back to the older V7 model when you use it. Midjourney's own guidance says fine details such as logos may not match your reference.
One image per prompt means a logo and a packshot cannot both be referenced. On a real brief you have a logo, three SKUs and a mood reference. The tool has one slot.
It can't put the brand name on the header. Text rendering has improved a lot across all these models. But Midjourney's own documentation lists "a logo with the company name" among its examples of prompts that don't work well. That is the vendor telling you not to expect the single most common requirement in POP.
Named formats come back as vibes. Ask for an FSDU and you get something with the visual language of a floor display at proportions no floor display has ever had. Counter glorifiers come back too tall and too theatrical. Endcaps are rarely at planogram height. The model is drawing from a vague visual prior, because "FSDU" is a word to it and a set of dimensional constraints to us.
Revisions fight you. The editor is genuinely capable, and anyone still saying Midjourney can't edit hasn't looked recently. But it runs on V6.1, two generations behind the model that made your image, and editing an HD image drops it back to standard resolution. So when the client asks for a taller header and a material change, you regenerate rather than edit, and regenerating moves everything you'd already approved. Revision rounds are where the hours in this business go.
Your work is public by default and trains the model. Midjourney's published training summary confirms it trains on user prompts and images, and the licence you grant is perpetual and survives cancellation. Stealth mode starts at $60 a month and only controls whether renders show up in the public feed. For an NDA'd launch that's a conversation with your client, not a footnote.
The part people get wrong about tools like ours
Here's the thing I'd want a competitor to be honest about, so I'll be honest about it first.
We do not have a better image model than Midjourney. We don't have an image model at all. AI POP Displays runs on Google's Nano Banana Pro, which is rented infrastructure available to anyone with a credit card, and which Google markets for high-fidelity product mockups. If you put our raw engine and Midjourney side by side and asked for a photograph of a mountain, Midjourney would win.
The product isn't the model. It's everything between a customer's email and the request that model receives.
That's four years of knowing which questions decide a display concept and which ones don't. It's format names that carry real geometry instead of being adjectives. It's material vocabulary where acrylic and PET and lacquered MDF behave like themselves under store lighting. It's shelf counts checked against the pack dimensions you uploaded rather than against what balances the picture. It's references travelling labeled, so the model is told this one is the product, this one is the brand mark, this one is the structure to follow, and the sketch goes last because that's what the model anchors on. It's retail environments that look like the aisle your client sells in.
None of that is model capability. All of it is domain knowledge, encoded once so you don't have to supply it in a prompt every time. A general model will give you what you asked for. The hard part of this job has always been knowing what to ask.
Where each one belongs
Midjourney. Mood boards, brand-identity exploration, atmosphere, the twenty-wild-directions session at the start of a pitch. Its draft mode returns 24 options from a single prompt for very little cost, which is the best exploration mechanic anyone sells. Use it before the brief exists, when nothing is committed and looseness is a feature.
AI POP Displays. From the brief onward. Client-facing concepts, brand approval, anything a manufacturer will quote against, anything where a real logo has to be a real logo.
Cost
Midjourney sells GPU time rather than images. Basic is $10 a month for 200 minutes of fast generation, Standard $30 for 900 minutes, Pro $60, Mega $120, with 20% off annually and no free tier. A standard prompt costs 0.8 minutes and returns four images, so Basic is roughly a thousand images and Standard adds uncapped relax-mode generation on top. Cheap, and cheap is easy when none of the output goes to a client.
Two clauses to know before running client work through it. If your business turns over more than a million dollars a year, commercial use requires the $60 plan. And stealth mode starts at the same $60.
AI POP Displays is free to try with 15 one-time credits and no card, then $49 a month for 150 renders at the founding price, with the $69 list price stated up front. Nothing you upload or render trains AI models, on any plan.
Where to go from here
The closer comparison is AI POP Displays vs ChatGPT, because the current OpenAI and Google models are genuinely strong at this and the argument gets more interesting. The underlying case is in Why generic image AI fails at POP displays.
If you have a brief on your desk right now, that's the fastest way to test any of this. Start here, first render lands in under a minute.
Frequently asked
Can Midjourney generate POP display concepts?
It generates images of things that resemble retail displays. That's not the same job. In my own testing across a lot of real briefs, the proportions don't match named formats, the shelf counts don't match the packs, and there is no mechanism for placing a client's actual logo on the fixture unchanged. It's the weakest fit of any major model for this work, which is not a criticism of the model so much as of the match.
Is Midjourney better than ChatGPT or Gemini for POP?
No, and it's not close. ChatGPT's current image model and Google's Nano Banana line both take multiple labeled reference images, follow detailed instructions, and let you edit an image without regenerating it. Midjourney's reference system accepts one image per prompt, silently downgrades you to an older model when you use it, and its editor runs two generations behind. For anything involving a real brand's assets, those three limits decide it.
So why does Midjourney have the reputation it has?
Because it earned it, on different subjects. For people, landscapes, atmosphere, editorial and concept art it is still wonderful, and it's the tool I open when I want a mood, not a spec. Its aesthetic instincts are the best in the business. Retail fixtures are just a subject where taste doesn't compensate for not knowing what an FSDU is.
Does Midjourney train on the images I upload?
Yes. Midjourney's own published training summary confirms user prompts and images are used to train its models, and its terms grant a perpetual, irrevocable licence over what you put in that survives cancelling your account. Stealth mode, which keeps work out of the public feed, starts at $60 a month and covers publication, not training. Most POP concepts are somebody's unreleased launch, so read that before uploading the brief.
What does AI POP Displays do that a general model doesn't?
It knows what to ask. We run on Google's Nano Banana Pro, so the raw image engine is rented, same as anyone's. What we built on top is the part that matters: a briefing layer that turns a customer's email into the right structured request, using the format names, materials, proportions and retail conventions this industry actually uses, with references labeled by role so a packshot is treated as product and a sketch as structure. The model is the commodity. The asking is the product.
Empezar a generar