Integrating Generative AI into Ecommerce: Building an Intelligent Product Recommendation Engine
How I built an AI Product Recommender using Generative AI APIs, prompt engineering, structured response schemas, and interactive shopping UI in React.
By Uttam Thapa · · AI/ML
⚡ Executive Summary (TL;DR)
Keyword search cannot answer "lightweight jacket for cool mountain evenings under $80" — the words that carry the intent never appear in the product
title. The AI Product Recommender puts a language model between the shopper and the catalog, but only under strict
conditions: a schema-constrained JSON response, a candidate set pre-filtered in SQL so the model ranks rather than recalls, and a loading experience designed
around a two-second inference.
Figure 1: Each result carries a match score and a plain-language reason — the reason is what turns a guess into a recommendation.
Where Keyword Search Gives Up
Traditional ecommerce search matches strings or walks a taxonomy. Both work well when the shopper already knows the vocabulary of the catalog, and both fail the
moment the query describes a situation instead of a product.
| Shopper types |
Keyword search does |
Intent search should do |
| "cool mountain evenings" |
Matches nothing — no product says "evening" |
Infers mid-weight insulation, wind resistance |
| "under $80" |
Treats it as three search words |
Becomes a hard price <= 80 filter |
| "something smart but not stuffy" |
Zero results |
Maps to a style axis and ranks along it |
Built with React, TypeScript, and a generative model API, the recommender interprets buyer intent, budget constraints, and aesthetic preference, then returns
scored matches with an explanation attached to each one.
Constrain the Model Before You Trust It
A model that answers in prose is unusable in a render function. The response has to be a contract, and the schema is that contract — declared in the system
prompt and, where the provider supports it, enforced in structured-output mode so malformed JSON is impossible rather than merely unlikely.
const SYSTEM_PROMPT =
"You are an expert e-commerce shopping consultant.\n" +
"Analyze the user's intent, budget limit, and style preferences.\n" +
"Score ONLY products from the supplied candidate list. Never invent a productId.\n" +
"Return ONLY valid JSON matching this schema:\n" +
"{\n" +
" \"matches\": [\n" +
" {\n" +
" \"productId\": \"string\",\n" +
" \"matchScore\": number (0 to 100),\n" +
" \"reasoning\": \"string explaining why this item fits criteria\"\n" +
" }\n" +
" ]\n" +
"}";
🚨 The failure mode that will reach production
A model asked to recommend products will happily return a productId that does not exist, complete with a persuasive reason. Always reconcile the
returned ids against the candidate set you supplied and drop anything unmatched before it reaches the UI. Never render a price, stock level, or product
name the model produced — look those up from your own database using the id.
The Three-Stage Pipeline
Stage 1
Intent parsing
Extract budget ceiling, usage context, colour and style preference, material constraints — as typed fields, not free text.
Stage 2
Candidate scoring
Hard constraints run in SQL first. Only the surviving shortlist is described to the model for scoring.
Stage 3
Ranking and explanation
Sort by confidence and surface the reason: "94%: breathable polyester shell suits mountain weather, $12 under budget."
Stage two is the one that decides whether this is a product or a demo. Budget, stock, and shipping region are facts, and facts belong in a
WHERE clause. Handing the model the whole catalog and hoping it respects "$80" wastes tokens and produces out-of-stock recommendations. Filter first,
rank second — the model's job is judgement about fit, not arithmetic about price.
💡 Production tip
Cache aggressively on the parsed intent, not the raw query. "jacket under 80 for cold evenings" and "under $80 jacket, cold evenings" normalise to the same
structured intent, so one inference can serve both. On a busy catalog this is the difference between a per-search cost and a per-distinct-need cost.
Designing for Two Seconds of Latency
Generative calls take between 1.2 and 2.5 seconds. That is far too slow for a spinner and far too fast to justify a background job, so the interface has to spend
the time rather than hide it. The frontend shows skeleton results alongside step-by-step status copy that mirrors the actual pipeline:
"Parsing budget parameters…", "Analysing material durability…", "Ranking top recommendations…".
This is not decoration. Narrating real stages sets the expectation that thinking is happening, and a reader who understands why they are waiting reports the same
delay as significantly shorter. Stream the results as they rank if your provider supports it — first match on screen at 600 ms beats all matches at 2.4 s.
✅ Key takeaways
- ✓Enforce a JSON schema. A render function needs a contract, not prose.
- ✓Filter in SQL, rank with the model. Price and stock are facts; fit is judgement. Give each to the system that is good at it.
- ✓Reconcile every returned id. Treat model output as an untrusted reference into your own data.
- ✓Ship the reasoning. An explained recommendation converts; an unexplained score reads as an advert.
- ✓Narrate the wait. Honest stage-by-stage copy makes two seconds feel deliberate.
If you are weighing where the embeddings behind this kind of search should live, Pinecone vs pgvector vs in-browser embeddings
covers the trade-off, and autonomous AI workflows in TypeScript extends the
same schema-first discipline to multi-step agents.
Frequently asked questions
How do you stop an LLM recommending products that do not exist?
Supply an explicit candidate list, instruct the model to score only from it, and reconcile every returned id against your own database before rendering. Never display a price or name the model produced — look those up by id.
Should the model apply the budget filter?
No. Price and stock are facts and belong in a SQL WHERE clause. Filter to a candidate set first, then let the model rank the survivors on fit. Models are good at judgement and unreliable at arithmetic.
How do you make a two-second AI response feel acceptable?
Narrate the real pipeline stages while it runs and stream results as they rank. A reader who understands why they are waiting reports the same delay as substantially shorter.
Home · Projects · Blog · Services · Résumé · Contact