BelltoAn AI agent that takes a brand's POP campaign from brief to shop drawings.The AI agent for POP campaigns.Get on the waitlist
← back to blog

Can an LLM Write the CAD for a POP Display?

October 11, 2026··9 min read

Text-to-CAD is now a research field with its own datasets, leaderboards and a steady stream of papers. You describe a part in words, a language model writes the instructions to build it, and a geometry kernel turns those instructions into a solid. For anyone in point-of-purchase the next question is obvious. If a model can write the CAD for a bracket, can it write the CAD for a counter display?

I've spent close to four years in POP. I build AI POP Displays, and we're building Bellto, an AI agent for POP campaigns that goes from brief to shop drawings. So I read the research with one question in mind, which is what a workshop would do with the file. This post covers what text-to-CAD measures, why a display is a different object, and where the gap sits.

What is text-to-CAD, and what does it measure?

Text-to-CAD means a model takes a written description and produces the instructions to build the part, rather than a picture of it. Some systems output a sequence of CAD commands, draw this profile, extrude it this far. Others write code for a Python CAD library, and running the code produces the solid. Either way the result is geometry, which is the step an image generator never takes.

The research leans on a small number of datasets, and those datasets decide what "works" means. The biggest is DeepCAD, built from models published on the Onshape platform. Its authors kept around 178,000, only those made with sketches and extrusions, and describe the collection as mostly user-created mechanical parts. Text2CAD, presented at NeurIPS in 2024, added written descriptions to those same models, and much of the work since trains or tests on DeepCAD or its Text2CAD layer.

The scores follow the data. A typical paper measures how far the generated shape sits from the reference, how much the two volumes overlap, and what share of the generated programs fail to run at all. The shapes are normalized, too. DeepCAD scales every model into the same 2 × 2 × 2 cube, so a part is judged on its shape and not on its millimeters.

The headline numbers need a close read. One 2025 paper that writes code for the CadQuery library reports a top-1 "exact match" of 69.3% in its abstract. In the body, that figure is the share of test cases where another AI model looked at renders of the generated part and of the reference and answered yes to whether they showed the same 3D object. The same paper's best model wrote scripts that failed to run 6.5% of the time. That is real progress on the problem it sets, one solid from one sentence. It is not a measure of whether anything can be built.

Why is a display a different problem from a mechanical part?

Because a display isn't a part. It's an assembly of parts cut from sheet, and almost everything that decides whether it works lives between the parts.

Every part has a material and a thickness. Corrugated board, foam board, MDF, acrylic, sheet metal. The thickness isn't a detail of one part, because it sets dimensions in the parts around it. A slot that takes a 5 mm shelf is a different slot from one that takes 3 mm, and a shelf resting on a ledge moves when the ledge does.

Parts meet at joints, and joints need clearance. Slot and tab, screws, bonding, folds. A slot cut to the exact thickness of the panel going into it doesn't assemble on a bench. A score computed on one solid can't see that, because the clearance only exists between two parts.

Parts must not collide. In a model, two panels can occupy the same space and still look perfect in a screenshot. On the bench they don't go together.

Every manufactured part needs a cut file. A cutting table works from a flat, closed contour per part, not from a solid. If one contour doesn't close, that part can't be cut, however good the assembly looks.

The product sets the size. Shelf depth comes from the pack, shelf width from the facings, shelf pitch from the room a shopper needs to lift a unit out. A shape scored inside a normalized cube has no pack in it at all.

The research is starting to say this itself. A benchmark published in May 2026 opens by noting that existing benchmarks focus on single-part models and score them with geometric similarity that misses functionality, manufacturability and assemblability. In its own tests, the check that parts don't interpenetrate fell off much more sharply than the checks on each solid, which its authors read as multi-component spatial reasoning being a major bottleneck even when the code runs. Another 2026 text-to-CAD benchmark lists sheet metal design and multi-part assemblies among the categories it doesn't cover, and an assembly paper from July calls production-ready mechanical assembly generation largely unsolved. Assembly benchmarks are appearing, and none of the ones I found is about a retail display.

If the piece on your desk has to reach a workshop as parts rather than as a picture, the Bellto waitlist is where I'd put it.

What does a model get right when it writes parametric code?

The useful word is parametric. Code that builds geometry from named values is the right kind of answer for a display, because a display really is a set of relationships. The shelf width follows the facings, the uprights follow the shelves, the base follows the uprights. Written down as relationships instead of drawn, a change moves everything that depends on it. I went through what that buys a campaign in parametric POP display design.

And a model fine-tuned on those mechanical parts writes code that runs most of the time. That matters, because it means the structure of a parametric program is no longer the hard part.

What parametric code doesn't do is make the decisions. A parameter carries a value through every part that depends on it, and it carries a wrong value just as faithfully. If the panel thickness is wrong, every slot sized from it is wrong, consistently. If nobody decided the joint, the code produces something shaped like a joint. The program is the easy part. The decisions inside it are the trade, and they're the ones a shape score never sees.

What has to be proven before a workshop can use the file?

Everything the shop would otherwise discover on the bench, at its own cost. The list is the distance between a solid that looks right and a package a workshop can quote.

  • Overlap between parts, measured as volume, across every pair of parts.
  • The envelope, the exact outer size of the assembly, against the space the store gives the piece.
  • Closed contours on every cut file, one file per manufactured part.
  • A STEP of the assembly that reimports into another CAD system with no solids lost.
  • Dimensioned drawings with hidden lines on their own layer, so a bench can read what sits behind what.
  • A cutlist, one row per part, with material, thickness, size and joint, which is what the estimator prices.

Each is a fact about the model rather than the look, none is in a shape-similarity score, and each has to run again after every change. The failure mode of parametric geometry is specific. Push it far from the shape it was built around and parts start to interfere, contours stop closing, and a joint that worked at one thickness stops working at another. I went file by file through what the shop does with each one in what a shop drawing package for a POP display contains.

Then there's what no check computes. Which material the budget carries across every store, which joint suits the volume, whether the piece has to ship flat. Those are trade decisions, and they get made before the geometry, not after it.

Where does the concept image fit?

Before any of this, someone has to agree what the display should look like. That question gets answered with a picture, and it's the job a concept render does. A render shows the format, the material family, the finish and how the product sits. It carries no millimeters, which is why a manufacturer can't quote it as it stands.

AI POP Displays is where that image comes from. It works from the brief in plain language, with a display taxonomy that carries proportions, materials that behave, retail environments as options and product load from the real packs, so the look the brand approves is a look a display can actually have. If your concept is still a paragraph in a client's email, run that brief through the free plan and get the look agreed first.

Once the look is approved, the render stops being the thing that goes out to quote and becomes the reference the piece is designed against, the handoff I covered from the brand's side in how a brand gets an AI concept built.

What we're building next: Bellto

The gap this post describes, between a solid that looks right and a package a workshop can cut, is the one Bellto exists to close. Bellto is an AI agent for brands and agencies that takes a POP campaign from brief to shop drawings. You tell it what you're launching (brand, product, channel, stores, budget, date) and it asks for what's missing. It proposes the campaign mix and the materials, floor stand, glorifier, shelf strip, stopper, with the budget split per store, and it designs every piece with you at real scale in parametric 3D, with proportions derived from the product and the facings. A dimension change recomputes the whole piece and re-runs the manufacturing checks in seconds.

Every approved piece comes out as a 3D model, a part-by-part cutlist, dimensioned drawings, a STEP of the assembly and a cut DXF per part, a package a workshop can quote without redrawing. Take a worked example, a four-panel acrylic counter glorifier of 200 × 200 × 250 mm. That's 4 parts and 4 cut DXF files, 0.0 mm³ of overlap across its 6 pairs of parts, a STEP reimported with no solids lost and drawings with hidden lines on their own layer, and after a dimension change the whole piece recomputes in under a second. Every model goes through engineering review before delivery. Bellto then suggests manufacturers that fit by material, format and volume, and the quote comes from the manufacturer.

We're building it now. The waitlist is open to any brand, and we're contacting the first ones soon to run the first real campaigns end to end. It opens by invitation, in small groups. Pricing goes first to the people on the list, and there's no card. POP manufacturers can sign up too.

So, can an LLM write the CAD for a POP display?

It can write code that builds a solid shaped like one. The research scores single mechanical parts inside a normalized cube, on shape, and in 2026 it started saying that this misses manufacturability and assemblies. A display is decided between the parts, in thicknesses, joints, clearances and cut files, and by the product it carries. That's the part a workshop pays for, and it's the part a shape score doesn't see.

If you have a campaign whose pieces need to reach the workshop as files instead of a picture, tell us about it on the Bellto waitlist. For the concept that comes first, signup is free with 15 one-time credits and no card required. Pro is $49 a month for 150 concept generations at the founding price, with the $69 list price stated openly. Nothing you upload or render trains AI models, on any plan.

Frequently asked

Can an LLM write the CAD for a POP display?

It can write code that builds a solid, and the solid can look like a display. The research behind text-to-CAD measures single mechanical parts, scaled into the same small box and scored on shape similarity, so nothing in those scores checks material thickness, joints, clearances, collisions between parts or cut files. A display is an assembly of sheet parts sized by the product it holds, and a workshop needs all of that before it can quote. Bellto, an AI agent for POP campaigns that takes a launch from brief to shop drawings, delivers every approved piece as a 3D model, a part-by-part cutlist, dimensioned drawings, a STEP of the assembly and a cut DXF per part.

What does text-to-CAD research actually test?

Mostly whether a model can turn a written description into a solid that matches a reference shape. The common datasets come from DeepCAD, around 178,000 models from the Onshape platform that its authors describe as mostly user-created mechanical parts, kept only if they were built with sketches and extrusions and normalized into the same small cube. Scores compare the generated shape with the reference and count how often the code fails to run. A benchmark published in May 2026 lists sheet metal design and multi-part assemblies among the categories it does not cover.

What does a workshop need that generated CAD code doesn't give it?

Decisions and checks, part by part. Every manufactured part needs a material and a thickness, a joint with the right clearance for that thickness, a closed contour for the cutting table and proof that it doesn't collide with its neighbors. The assembly needs a STEP the shop can open in its own CAD, dimensioned drawings with hidden lines on their own layer and a cutlist the estimator can price. A solid that looks right answers none of those by itself.

Is an AI concept render the same thing as CAD?

No. A render is pixels. It gets the look approved, the format, the material family and how the product is presented, and it carries no millimeters. AI POP Displays is where the concept image comes from, with the format, the material and the product load right, so the brand agrees what it wants before anything is engineered. The geometry is a separate step, and it is the step Bellto is built for.


Start generating Get on the Bellto waitlist
Can an LLM Write the CAD for a POP Display? | AI POP Displays