BelltoAn AI agent that takes a brand's POP campaign from brief to shop drawings.The AI agent for POP campaigns.Get on the waitlist
← back to blog

AI POP Displays vs Gemini for Retail Display Concepts

September 11, 2026··10 min read

AI POP Displays runs on Google's Nano Banana Pro, API id gemini-3-pro-image. I'm putting that in the first paragraph because it decides how you should read everything after it. This isn't a comparison between two image models. It's a comparison between what reaches the model when you type a display into a chat window and what reaches it when the request is built from a display brief.

I've worked in point-of-purchase for close to four years and I build this product. What follows is specific about what Gemini gives you directly, what going direct asks of you, and where the hours actually go on a display brief.

Which Gemini model are we talking about?

The naming is confusing enough that it's worth pinning down. As of today, Google has three image models in the Gemini line.

  • Nano Banana Pro, API id gemini-3-pro-image, generally available since May 28, 2026.
  • Nano Banana 2, gemini-3.1-flash-image, a cheaper and faster model. It's a separate model, not a newer version of Pro, and it's what the Gemini app now uses for image generation by default.
  • Nano Banana 2 Lite, a 1K-only model below that.

The original Nano Banana, gemini-2.5-flash-image, is scheduled to shut down on October 2, 2026. The preview ids for Pro and Nano Banana 2 were already shut down in June. So if a tutorial you're following names a -preview model, it's describing something you can't call any more.

What does Nano Banana Pro give you directly?

Quite a lot, and none of it is locked away. Everything below comes from Google's own documentation, checked the day I wrote this.

References. Up to 14 input images, of which up to 6 can be objects Google says it holds at high fidelity and up to 5 can be characters. For display work, the six object slots are the ones that matter. A logo and four pack shots already take five of them.

Resolution. 1K, 2K and 4K. Google bills 1K and 2K the same, because both come out at 1,120 output tokens, so there's no reason to stop at 1K.

Aspect ratios. Ten of them on the Gemini API, including 3:4 and 9:16, the verticals most floor displays and totems need.

Thinking. The model reasons about the composition before it commits, generating up to two interim images along the way. It's on by default and can't be switched off on Pro.

Grounding. It can ground a generation with Google Search.

Watermarking. Google's docs state that every generated image carries a SynthID watermark. It's invisible and there's no documented way to opt out.

I've run a lot of display briefs through this model. The capability was never the bottleneck. What reached it was.

What does a display brief need that the model doesn't ask for?

A general image model has no failure state. It never comes back and tells you the request was underspecified. It fills whatever you left open with whatever was most ordinary in its training data, and it hands you a picture at full confidence.

On a display, the parts people leave open are the deliverable. I went through the seven decisions a concept depends on in how to prompt AI for retail displays, and the short version is sector, format, material, style, retail environment, product load and framing. Gemini will honor all seven if you name them. It won't ask for any of them.

Format names read as style. Write "FSDU" and the model gets the visual language of a floor display, not its geometry. The proportions you need have to be described in the request.

Materials read as adjectives. "Premium" is a mood. E-flute corrugated, lacquered MDF or clear acrylic with a visible edge are surfaces, and each one implies different thicknesses, joints and edges.

Product load gets composed rather than merchandised. Unless you say how many facings go on each shelf, you get the number that balances the picture.

Unlabeled references compete. Drop a logo, four pack shots and a competitor's stand into the chat and the model has to guess which image is the product, which is the brand mark and which is only there for the silhouette. It usually averages them into inspiration.

None of that is a Gemini defect. It's what happens when a general tool meets a specialist brief. Closing those questions before the model is called is the layer we built, and the free credits cover a real test of it.

What does the same brief look like both ways?

Here's a brief of the kind that actually lands in a manufacturer's inbox, with the brand invented.

Hi, we're launching Northfield protein bars in four flavours at grocery in the spring. Need a floor stand concept for the category buyer, cardboard, something bold. Pack shots and logo attached. Can we see something this week?

Typed straight into Gemini

Most people turn that into something like this, with the five images attached underneath.

cardboard floor display for a protein bar brand, four flavours, supermarket, bold, photorealistic

That leaves the proportions, the shelf count, the facings per shelf, the finish of the board, the camera angle and the job of each image entirely to the model. You'll get a floor display. Whether it's the one the buyer approves depends on luck and on how many reruns you have time for.

What our pipeline actually sends

The customer's email goes in as it is, typos included. A briefing step reads it and fills a form from a fixed taxonomy, so you're reviewing choices rather than writing a prompt. That form becomes the request below. It's abridged from the real prompt our code builds, with the long rule block at the top left out.

Briefing: Brand: Northfield Product: protein bars, four flavours Sector: Food & Confectionery Display type: Floor display Material: Cardboard (E-flute) Style: Bold / graphic Background mode: Supermarket scene Aspect ratio: 3:4

Reference images: 4 product images are provided as references. These represent the product range / different SKUs. Place all of them realistically inside the display, at the right scale, with each variant facing forward and clearly readable. Do not merge them into one product.

Every image then travels with a written label in front of it. "Reference: brand logo." "Reference: product photo 1 of 4," and so on up to four. If the customer sent a sketch, it's labeled as the primary reference and goes last in the stack, because these models anchor hardest on the final image they receive.

Look at what changed. Every open question got closed before the model was called, with the same vocabulary, in the same order, every time. That request is the work that decides the display. The picture is what the model does with it, and on the revision round it's the request that keeps the rest of the piece in place.

The test that settles it is the brief on your desk, run both ways, and signup takes about a minute for the second half of that test.

What does going direct ask of you?

This is the part most Gemini guides skip, and it changed this year.

Through the API. Google's pricing page lists no free tier for Nano Banana Pro. It's $0.134 per image at 1K or 2K and $0.24 at 4K, with the batch rate at half that. To reach the paid tier you link a billing account and prepay at least $5 in credits.

Through the Gemini app. Image generation runs on Nano Banana 2 by default. Google's help pages list Nano Banana Pro as available with a Google AI plan, reached by generating an image first and then choosing to redo it with Pro. In the US, Google AI Plus is $4.99 a month, Google AI Pro is $19.99 and Ultra is $99.99 or $199.99 depending on usage. Google doesn't publish daily image counts for any plan. Without a plan, downloads are 1K. With one, they're 2K.

Through AI Studio. Google's own pages disagree with each other on exactly what a free account can run there. What's clear is that the Pro image model sits behind either a Google AI plan or a paid API key.

The per-image price is not where a display brief costs. The cost sits in everything before the call: writing the seven decisions out, labeling the references, ordering them, and doing it all again on the revision round. A rerun that reaches the buyer with the wrong shelf count costs more than any of those, and going direct gives you no way to tell it apart from a good one until the buyer does.

Our pricing is a Free plan at $0 with 15 one-time credits and no card, then Pro Beta at $49 a month for 150 concept generations at the founding price, with the $69 list price stated openly. Each credit carries the whole layer above it.

Where do your client's files go?

Worth checking before a confidential launch goes into a chat window, because the answer depends on the route.

Google's Gemini API terms treat direct use of AI Studio without billing, and any free quota, as unpaid services. Content sent that way may be used to improve Google's products, and human reviewers may read it after it's disconnected from your account. Google's own advice is not to submit confidential information there. On paid services, Google says it doesn't use your prompts or responses to improve its products.

The consumer Gemini app works differently again. With Keep Activity on, which is the default, chats are saved and used to improve Google's models, and a subset is reviewed by people. Turning it off limits retention to 72 hours.

Our generations go through Google's paid API. On top of that, nothing you upload or render with us trains AI models, on any plan, including the free one. Most POP concepts are someone's unreleased launch, and I don't think privacy should be a paid feature.

Where does Gemini on its own still fit?

Before the brief exists. Loose directions, mood, color and material atmosphere, the exploration that happens before anyone has decided it's a floor stand. And for subjects that aren't a fixture, packaging on its own or a scene built around people. AI POP Displays is narrow on purpose and doesn't do those.

From the brief onward, when the input is a customer's email, the output goes in front of a buyer and the revision round has to keep everything else put, this is the tool.

What we're building next: Bellto

Everything above ends at a concept image. Bellto is the new entrant in this comparison and the only one that takes a campaign to a manufacturable package. It's an AI agent for brand and shopper marketing teams and their agencies. You tell it what you're launching (brand, product, channel, stores, budget, date) and it asks for what's missing. It proposes the campaign mix and materials, floor stand, glorifier, shelf strip, stopper, with the budget split per store. Then it designs every piece with you at real scale in parametric 3D, with proportions derived from the product dimensions and facings, so a dimension change recomputes the whole piece and re-runs the manufacturing checks in seconds. Every approved piece comes out as a 3D model, a part-by-part cutlist, dimensioned drawings, a STEP of the assembly and a cut DXF per part, a package a workshop can quote without redrawing. It suggests manufacturers that fit by material, format and volume. The quote still comes from the manufacturer.

We're building it now. The waitlist is open to any brand or agency, and we're contacting the first ones soon to run the first real campaigns end to end. It opens by invitation, in small groups, and the campaign you describe when you sign up sets your place. Pricing goes first to the people on the list. No card. POP manufacturers can sign up too.

Where to go from here

I've written the same comparison for ChatGPT, and there's a spec-level read of Seedream against Nano Banana Pro if you're weighing engines rather than workflows.

Knowing which questions decide a display is what close to four years in this industry taught me, and it's what every credit here carries. If the brief on your desk this week is a real display, start with the free credits, and if the campaign behind it has to reach a workshop, put your name on the Bellto list.

Frequently asked

Can Gemini design a POP display?

It can generate a convincing image of one. Nano Banana Pro, the Gemini image model, reaches 4K, takes up to 14 input images with up to 6 of them held at high fidelity, and thinks through the composition before it renders. What it doesn't do is ask the questions a display brief depends on. Format, material, product load, retail environment and which reference is which are all things you have to know and write into the request yourself, on every brief.

Is AI POP Displays just Gemini with a form on top?

Gemini is the engine. We run Nano Banana Pro through Google's paid API, and an OpenAI model for some edit operations. What decides the display is the layer on top: a briefing step that reads the customer's email and fills a form from a fixed taxonomy, display types that carry their proportions, a materials list that behaves like the material, retail environments as options, product load taken from the real packs, and references labeled by role with the sketch last. That layer comes from close to four years in the POP industry, and it's done once instead of retyped on every brief.

Can I use Nano Banana Pro for free?

Not really, as of September 2026. Google's pricing page lists no free API tier for Nano Banana Pro. It bills $0.134 per image at 1K or 2K and $0.24 at 4K. In the Gemini app, image generation runs on Nano Banana 2 by default, and Google's help pages list Nano Banana Pro as available with a Google AI plan, which starts at $4.99 a month in the US. Check the official pages before you plan around any of it, because these tiers have changed several times this year.

Does Google train on the images I upload to Gemini?

It depends on how you reach the model. Google's Gemini API terms say content sent through unpaid services, including direct use of AI Studio without billing, may be used to improve Google products and read by human reviewers. On paid services it isn't. In the consumer Gemini app, the default Keep Activity setting saves chats and uses them to improve Google's models, and a subset is reviewed by people. If the pack shots belong to a client's unreleased launch, pick the route before you upload.


Start generating Get on the Bellto waitlist
AI POP Displays vs Gemini for Retail Display Concepts — AI POP Displays