BelltoAn AI agent that takes a brand's POP campaign from brief to shop drawings.The AI agent for POP campaigns.Get on the waitlist
← back to blog

Getting Your Actual Product Into an AI Display Render

October 2, 2026··9 min read

The render is good. The structure is right, the header sits where it should, the aisle behind it looks like the retailer's aisle. And on the shelves there's a bottle that resembles your client's bottle. Same color, roughly the same shoulder, a label that reads like the brand if you don't look twice.

The client will look twice. The pack is the one thing in the picture they know better than you do, and a lookalike turns the concept review into a conversation about the pack.

I've worked in point-of-purchase for close to four years and I build AI POP Displays. Getting the real product onto the display is the step I see done wrong most often, and it's almost never about the prompt. It's about which photos went in, how many, and what the tool was told they were.

Why does the pack come back as a lookalike?

Because it was described, or because it was uploaded into a pile.

A description can't carry a pack. "Amber glass dropper bottle, white label, gold cap" matches a thousand bottles, and an image model draws the most ordinary one. Your client's bottle has a specific shoulder, a specific label height and a specific gold. Those have to be shown.

Showing it isn't enough either if the photo arrives unlabeled. Attach a pack shot, a logo and a photo of last year's stand to a general image model and it has three images and no idea which one is the product, which one is a mark to reproduce and which one is a structure to borrow from. They compete.

Search for help with this and nearly everything you find is written for e-commerce. One pack shot goes in, and it comes back on a marble counter or a beach. Or it's packaging mockups, a label wrapped onto a blank jar. All of it is about a single hero product alone in the frame. A display is the other case. The same pack has to repeat across a shelf, at a size that makes sense against the fixture, eight times and then eight times again on the tier below, and somebody is going to count them.

What happens to a product photo once you upload it?

It gets a role, and the role travels with it.

In AI POP Displays the uploads aren't one bucket. There's a slot for the logo, a slot for product images, one for campaign artwork, one for a POP reference and, on the sketch route, the sketch itself. When you generate, every image goes to the model with a written label in front of it that says what it is. A single pack travels as the product photo. Four packs travel as product photo one of four, two of four, and so on.

The instruction that goes with them is specific to a display. One product photo is used as the product, placed inside the display at the right scale, repeated as consistent units facing forward, with its shape, color and packaging preserved. Several product photos are read as the range, meaning different SKUs, each placed facing forward and readable, and explicitly not merged into one product.

The order is fixed as well. Logo first, then the product photos, then artwork, then any style reference, and the sketch last. We built it that way because in our pipeline the last image is the one the model treats as the subject to produce, and on a display that's the structure, never the pack.

One more detail matters when a deadline is close. The prompt is written from the references that actually loaded. If a photo fails to load, the render isn't told about a product it can't see. And if the sketch is the one that fails, the generation stops before a credit is spent.

If the packs for the brief you're working on are in a folder right now, upload them with the brief and look at what comes back on the shelves.

Which photos should you upload, and how many?

One photo per SKU that goes on the display. That's the whole rule, and most mistakes are a way of breaking it.

What you uploadWhat the render does with it
Front face of one pack, plain backgroundRepeats that pack across the facings
One front-face photo per SKU in the rangeLoads each variant as its own product
Three angles of the same packReads them as three SKUs
The whole range in one group shotReads the group as one product
A lifestyle shot with props and a handBrings the props along
A cropped or tilted packGuesses the part it can't see

Front face. On a shelf the shopper sees the front of the pack, so that's the face the render repeats. A three quarter beauty shot looks better in a deck and worse as a reference, because every unit on the shelf inherits the angle.

One pack, whole, upright. The full outline in frame, nothing cut off at the cap or the base, and the file rotated the right way up. Google's own guidance for image inputs says the same two things, that images should be correctly rotated and not blurry.

Plain background. A white or grey sweep. Whatever else is in the photo is a candidate for the render.

One SKU per photo. Several photos are read as several variants, so a range of four is four photos. A group shot of the four is one photo, and it comes back as one odd wide product.

Only what goes on the unit. AI POP Displays runs on Google's Nano Banana Pro, and Google's documentation for that model puts the object images it holds at high fidelity at six at most, out of fourteen reference images in total. A logo and four packs are five. If the range has twelve SKUs and the display holds four, upload the four.

The files are PNG, JPEG or WebP, up to 10 MB each. A clean pack shot from the client's e-commerce listing is usually the right one, and they already approved it.

What has to go in the brief next to the photos?

The count. A photo says what the pack looks like and nothing about how many there are.

Three lines do most of the work. Facings across each shelf, how many deep, and how many tiers. If the range has more than one SKU, say which goes where, as in the large format on the bottom shelf and the three small variants across the top two. If two packs differ in size, say so in relative terms, the 500 ml about twice the height of the 100 ml, because two photos cropped to the same frame look the same size to the model.

I put the full list of what a display brief has to name in how to prompt AI for retail displays, and product load is the line almost everybody leaves out. With the real pack attached, that line stops being a guess, because the count is now a count of something specific.

If the customer sent a drawing, attach it too. The pack photos and the sketch do different jobs and they don't fight, because each arrives labeled and the sketch arrives last. How that route works is in sketch to render for POP displays.

How do you check that the pack survived?

At full size, in three passes.

Shape and color. The silhouette, the proportions of cap to body, the main color blocks and where the label sits. These are what a reference holds, and they're what the client recognizes from across the room.

The front of the pack. The brand name and the main graphic should read as the real ones. The small print is another matter. Google's own note on its Pro model says small text and fine details may not come out perfectly, and an ingredients panel at shelf scale is exactly that. Nobody approves a display on the legal line of a label, so check the name and the look, and let the artwork file carry the rest.

The count. Use the pack as the ruler. Count pack widths across a shelf and pack heights between two shelves, and compare with what you plan to build. The five-minute version of that check is in how accurate are AI display renders.

Two edits make the check easier. Remove product, which is on the free plan, renders the same display empty, so you can judge the structure without the load in the way. And on Pro, a custom edit takes one instruction in your own words plus up to four reference images, so "replace the cartons on the top shelf with this pack" arrives with the pack attached and the rest of the display held.

If the render on your screen has a stand-in where the client's pack should be, put the same brief through with the real photos and compare the shelves.

What doesn't a product photo settle?

The brand mark on the display itself. The logo on the header is its own reference, uploaded in its own slot, because a mark on a header card and a mark on a pack are two different reproduction jobs. If the header came back soft, that's covered in my AI display render came out wrong.

The print. The render shows the client their pack on the unit. The artwork your manufacturer prints from is still the print file.

The dimensions. The pack in the render is the right shape at a believable size, and nothing in the image is measured. Shelf pitch and clearance in millimeters come from the real pack with a caliper on it, in the drawing.

And one thing a pack photo raises. It's often an unreleased product under NDA, sent to you by a client who trusts you with it. Nothing you upload or render in AI POP Displays trains AI models, on any plan.

Where the pack becomes a dimension

In a concept render the product is a picture of a pack. The step after that is the pack as a measurement that the whole piece is built around, and that's what we're building now.

Bellto is an AI agent for brands and agencies that takes a POP campaign from brief to shop drawings. You tell it what you're launching (brand, product, channel, stores, budget, date) and it asks for what's missing. It proposes the campaign mix and materials, floor stand, glorifier, shelf strip, stopper, with the budget split per store, and designs every piece with you at real scale in parametric 3D, proportions derived from the product and the facings. A dimension change recomputes the whole piece and re-runs the manufacturing checks in seconds. Every approved piece comes out as a 3D model, a part-by-part cutlist, dimensioned drawings, a STEP of the assembly and a cut DXF per part, a package a workshop can quote without redrawing, and it suggests manufacturers that fit by material, format and volume. The quote comes from the manufacturer.

We're building it now. The waitlist is open to any brand, and we're contacting the first ones soon to run the first real campaigns end to end. It opens by invitation, in small groups, and the campaign you describe when you sign up sets your place. Pricing goes first to the people on the list. No card. POP manufacturers can sign up too.

For the pack on your desk today, signup is free with 15 one-time credits and no card required, enough to load the real range onto a display and count what comes back. Pro is $49 a month for 150 concept generations at the founding price, with the $69 list price stated openly.

Frequently asked

How do I get my real product into an AI display render?

Upload a photo of the real pack as a product reference instead of describing it. Words can describe a bottle, they can't specify your client's bottle. In AI POP Displays every product photo travels with a written label saying it is the product, so it is loaded onto the display at scale, facing forward, with its shape, color and packaging kept. Then put the facings in the brief, because the photo says what the pack looks like and nothing about how many there are.

How many product photos should I upload for a display render?

One per SKU that goes on the display, and no more. When several product photos arrive, each one is read as a different variant in the range and they are kept separate rather than merged. Three angles of the same pack would be read as three products. Keep the total small as well. Google's documentation for its Pro image model puts the object images it holds at high fidelity at six at most, so a logo and four packs already fill most of that.

What kind of product photo works best as a reference?

The front face of one pack, whole and upright, on a plain background, sharp and evenly lit. That is the face a shopper sees on the shelf, so it is the face the render repeats across every facing. A lifestyle shot, a group shot of the whole range or a cropped pack gives the model something else to copy. AI POP Displays takes PNG, JPEG or WebP files up to 10 MB each.

Does the render show the correct number of packs on each shelf?

Only if the brief names it. A product photo carries no count, so state the facings across, the depth and the tiers, and then check the render with the pack as the ruler. A render still holds no millimeters. For a display dimensioned around the pack, Bellto, an AI agent for POP campaigns that we are building, designs each piece at real scale in parametric 3D with proportions derived from the product and the facings. It is open by invitation from a waitlist.


Start generating Get on the Bellto waitlist
Getting Your Actual Product Into an AI Display Render | AI POP Displays