Let AI fill the product-information gap before it rewrites the catalogue
AI product copy is most useful where a buyer lacks information. Find the thin listings, ground each addition in verified attributes and measure whether completing them raises conversion without creating returns.
An e-commerce director is shown a catalogue-writing demonstration. The system can rewrite every description, translate it into several languages and produce thousands of polished words before lunch. The proposed business case counts the hours required to write the same volume by hand.
That measures production rather than value. A description changes the economics only when it gives a buyer information that was missing from the decision. A fluent rewrite of an already complete page can be cheap to make and commercially irrelevant.
Let AI fill the product-information gap before it rewrites the catalogue. Start with listings that have missing, thin or unusable detail. Generate from verified attributes, keep complete listings as a control and measure whether the additional information changes purchase and post-purchase outcomes.
The gap is where the value can appear
The useful evidence comes from a series of randomised field experiments at a large cross-border retail platform. One experiment added AI-generated product descriptions to about 45,000 selected products. Across five language groups, 4,772,937 consumers who reached those product pages were assigned to see either the existing description or the generated description added above it. Fang and colleagues, 2026
The treatment increased orders by 1.1 per cent, sales by 2.1 per cent and conversion by 1.3 per cent. It produced no statistically significant change in cart value or the number of orders among people who purchased. The economic mechanism was therefore purchase incidence: more visitors crossed from consideration to an order, rather than existing buyers spending more. Generative AI and Sales Productivity, Table C5
The average hides the decision that matters. In the English-language experiment, the researchers separated products whose original descriptions held no more than fifty words from those with more. The first group recorded a 6.5 per cent sales increase after augmentation. The estimated increase for the second group was 0.03 per cent and was not statistically significant. The test comparing the two percentage effects returned a p-value of 0.0536, so the difference is suggestive rather than a settled threshold. Generative AI and Sales Productivity, Table C6
The commercial reading is narrower than “AI copy sells”. The intervention had room to add information because nearly half of the platform’s self-sold products had no description or only limited text. Many suppliers presented product detail inside images in Chinese, while the international pages needed structured text in the buyer’s language. The system was filling a visible information deficit.
| A volume-led rollout | A gap-led rollout |
|---|---|
| Rewrite every listing that can be processed | Rank listings by missing decision-relevant information |
| Treat fluent copy as the completed output | Require every generated claim to resolve to a verified product attribute |
| Report descriptions generated and time saved | Measure incremental conversion, returns and information-related contacts |
| Apply one review rule to the whole catalogue | Set review depth from product stakes, source quality and the consequence of an error |
The first column optimises the supply of words. The second invests in the buyer’s decision. That distinction determines which catalogue records enter the pilot and which evidence can justify extending it.
Build the queue from product data
A useful rollout begins with a completeness register. For each listing, record which attributes a buyer needs, which are present, where each value came from and whether the value is consistent across the catalogue, packaging and supplier record. Word count is a cheap screening signal. It is not a definition of quality.
Research with 3,544 UK online DIY shoppers shows why the register needs more than length. Participants were randomly shown scenarios varying eight aspects of product-information quality. Poor information increased the perceived severity of a service failure, which in turn reduced perceived service quality. Product images, title readability and attribute consistency had the strongest effects among the criteria tested. The experiment measured stated attitudes and willingness to continue shopping rather than completed purchases, and it covered two DIY products, so it identifies useful audit fields rather than a forecast. Weber and colleagues, 2023
The generation boundary follows that register. The model can turn verified attributes into readable text, adapt the structure to a market and surface missing fields for a person to resolve. It should not infer a material, safety property, compatibility claim or warranty from a product image because the paragraph sounds incomplete without one. Where sources disagree, the output is a data-quality task rather than publishable copy.
- 01 Find Score listings for missing, unreadable or inconsistent decision attributes.
- 02 Ground Generate only from approved catalogue, supplier and packaging records.
- 03 Review Check risky claims and sample lower-stakes output before release.
- 04 Measure Compare purchase and post-purchase outcomes with a credible control.
- 05 Decide Expand only into catalogue segments where the measured value exceeds the operating burden.
Keep the original listing and every source value with the generated version. That makes corrections possible and lets the team distinguish a poor generation from poor source data. Without that lineage, a later return or complaint becomes a copy problem with no recoverable cause.
Measure the decision, then protect its quality
Choose the randomisation unit before release. Assigning visitors to variants can measure the effect of a description while keeping the product fixed. Assigning products can measure the operating effect of changing whole listings, but it needs enough products and careful treatment of seasonality, category and traffic. Whichever design is used, predeclare the primary outcome and the smallest change that would alter the investment decision.
Conversion and orders show whether more visitors purchased. Read them with returns, cancellations, complaints, product ratings and contacts caused by missing or inaccurate information. A generated description that raises initial conversion and creates avoidable returns has shifted friction beyond checkout. Measure the human review rate, rejected-claim rate and correction time as the cost of maintaining the workflow.
Segment the result by the baseline gap, product category, language and purchase stakes. The hypothesis is strongest for listings where a buyer lacks usable information. A single catalogue-wide average can dilute that result with pages that had no problem to solve, then make a useful intervention look weak.
What the Inference Institute can help decide
The useful engagement starts before a model or copy platform is selected. It defines the product-information gap, maps the source records, chooses the safe generation boundary and designs an experiment that can separate additional conversion from ordinary traffic variation. The resulting specification gives an e-commerce director something concrete to procure: eligible catalogue segments, required inputs, review rules, measurements and a stopping condition.
This also establishes where generative AI is unnecessary. Some listings need a missing attribute collected from the supplier. Some need consistent structured data across channels. Some already answer the buyer’s questions. The programme creates value by routing each record to the right remedy, not by ensuring every record passes through a model.
What this does not tell you
The field evidence comes from one cross-border platform, its self-sold products and short experiments conducted in 2023 and 2024. The generated descriptions were added to the existing ones, so the result does not show what would happen if a retailer replaced its current copy. The sparse-description comparison used one language, and its between-group result sits just outside the conventional five per cent significance level. Sales are also not profit.
The study did not expose the descriptions or their source attributes, so it cannot establish a general accuracy method. The separate UK experiment used hypothetical shopping scenarios and measured reported responses. Neither source provides a universal completeness score, review rate or return threshold.
The e-commerce director’s next decision is which part of the catalogue has a buyer-information problem large enough to test. Begin there, preserve a control and make post-purchase quality part of the result. If the completed listings convert more buyers without creating later friction, the organisation has a case for expansion. If they do not, it has avoided turning inexpensive prose into an expensive catalogue-wide habit.