Visual search is not a storefront-wide upgrade. It is a per-category routing decision, and the variable that decides it is attribute nameability: whether the attribute driving the purchase can be typed in about three words. When it can — SKU, model number, size, spec, ingredient — an image query adds latency and a fallback risk without adding precision. When it cannot — a tile glaze, a jacket silhouette, a trim profile, a fabric weave the shopper has no vocabulary for — the image is the only query the shopper is capable of forming.
Most rollouts skip that question. The feature is switched on catalogue-wide because it is a platform setting rather than a per-category one, and the expectation is uniform lift. What actually happens is uneven: a few categories carry a real conversion delta, most are flat, and a handful are worse than the text path they partially displaced. The site-wide average hides all three.
Which categories reward image input?
The rubric below is the one we use when scoping a discovery surface. It sorts by nameability of the discriminating attribute, not by how photogenic the product is — a common confusion, since almost everything in a catalogue photographs well.
| Category pattern | Discriminating attribute | Nameable in ~3 words? | Route |
|---|---|---|---|
| Apparel, patterned textiles, rugs | Pattern, silhouette, drape, colourway | No | Visual-first, text refine |
| Tile, stone, laminate, paint finishes | Glaze, veining, sheen, texture | No | Visual-first |
| Furniture, lighting, decor | Form, proportion, finish | Rarely | Visual-first, text refine |
| Hardware, fasteners, trim, mouldings | Profile and thread geometry | Sometimes (if the shopper knows the standard) | Mixed — visual for profile, text for size |
| Consumer electronics, appliances | Model number, spec, compatibility | Yes | Text-first |
| Grocery, pharma, consumables | Brand, variant, ingredient, pack size | Yes | Text-first |
| Auto and industrial parts | Part number, fitment | Yes | Text-first, with fitment filters |
| Books, media, software | Title, author, edition | Yes | Text-first |
Two things about this table are worth stating explicitly. First, “text-first” does not mean “no visual entry point” — a shopper standing in front of an unlabelled appliance still needs the camera. It means the visual path is a recovery route for a minority of sessions in that category, and should be scoped and budgeted as one rather than as the primary surface. Second, the mixed rows are where most of the interesting engineering sits: hardware and trim are categories where the shopper can photograph the profile but must type the dimension, and a surface that forces a choice between the two loses.
The signals already sitting in your search logs
You do not need a storefront rollout to get a first answer. The category-level evidence is already in the existing text-search logs, and reading it before building anything is the cheapest step in the whole exercise.
- Null-result rate per category. High null rates on text queries usually mean shoppers are reaching for words the catalogue does not carry — a nameability failure, and the strongest single indicator of a visual-search candidate.
- Refinement depth. Categories where sessions run four, five, six query reformulations before a click are categories where the shopper is circling an attribute they cannot name.
- Query abandonment. Searches with no click and no reformulation are the sessions a visual path is meant to recover. Size that population per category before promising a lift.
- Descriptive-adjective density. Queries loaded with hedged visual language (“light wood grain looking”, “sort of a herringbone”) mark the boundary of the shopper’s vocabulary directly.
- Near-duplicate density in the catalogue. This one cuts the other way: categories where SKUs differ only by size, colourway or pack count are categories where image matching will confidently return the wrong variant. The failure modes of CV product-matching in retail visual search are worth reading against your own catalogue before committing a category.
The first four signals argue for visual search. The fifth argues against it in the same breath, and a category can score high on both — patterned apparel is often exactly that. When that happens, the decision is not visual-or-text; it is visual retrieval with a mandatory text or facet disambiguation step layered on top.
Combining the two on one surface
The framing that survives contact with production is one discovery surface with two query modes, not two competing surfaces. Our broader argument for that is in how AI visual search changes product discovery for retailers, and the reasons a replacement framing breaks down are covered in why visual search does not replace text search.
The practical pattern: the image narrows the candidate set to a visual neighbourhood, and text or facets resolve the nameable attributes inside it. Shopper photographs a jacket, then filters to size M and linen. Neither query mode could have done that alone. In our engagements on retail discovery surfaces, this hybrid pattern is where the per-category deltas actually show up — a pure image-only path tends to strand the shopper at the point where their intent becomes verbal.
Fallback behaviour is the other half of the surface design. When the matching model returns no confident result, the honest response is a named fallback — visually similar items with the confidence stated, or a handoff into text search with any extracted attributes pre-populated — not a padded result list. Fallback rate is a per-category number, it belongs in the rubric alongside conversion, and a category with a high fallback rate is a category the visual path is not ready for regardless of how well it photographs. What counts as acceptable depends on how the fallback is designed: a pre-populated text handoff tolerates a much higher rate than a bare “no results” page.
What to measure before you commit
The measurement that matters is the per-category delta between image-search-to-cart and text-search-to-cart on comparable traffic — not a site-wide average, which will report success for the storefront while three categories quietly underperform. Alongside it, track null-result rate for text queries in that category, the fallback rate above, and the abandonment population the visual path was supposed to recover. Isolating a real delta from seasonality and self-selection is its own discipline; see measuring visual search conversion lift against noise for the test design.
Routing matters downstream too. The per-category decision determines which slices of the catalogue the image index has to serve at latency, which is a sizing and throughput question rather than a modelling one. Deciding the rubric first means the index is scoped to the categories that earn it, and the computer vision engineering side of the work — matching confidence, fallback thresholds, index freshness — is scoped to the same set. The wider set of product-discovery problems this sits inside is covered on our retail AI page.
One thing the rubric does not settle: it degrades. Catalogue image quality and churn move the answer over time. A category that scored well on a professionally shot, slow-turning catalogue scores differently after a supplier-image migration or a weekly seasonal refresh. Re-running the nameability and null-result read once a quarter, per category, is cheaper than discovering the drift through a flat cart rate — and the question worth asking your own team is which of your categories would fail that re-read today.y.y.
Frequently Asked Questions
Which product categories reward image input, and which ones are better served by structured text and filters?
Categories where the purchase-driving attribute is visual and hard to name — pattern, silhouette, finish, tile glaze, trim profile — reward image input. Categories where the shopper already holds a precise identifier (model number, part number, pack size, ingredient) are better served by structured text and filters, because typing the identifier is faster and more precise than photographing it. The table above sorts the common patterns by attribute nameability.
How do we test visual-vs-text efficacy per category without running a full storefront rollout?
Start in the existing text-search logs: null-result rate, refinement depth, query abandonment and descriptive-adjective density are all per-category and already collected. Then pilot the two or three highest-scoring categories only, measuring image-search-to-cart against text-search-to-cart on comparable traffic. A catalogue-wide switch-on tells you less, later, and at higher cost.
What signals in existing search logs indicate a category is a visual-search candidate?
High null-result rates, deep query reformulation chains, high no-click abandonment, and queries full of hedged visual language all point the same way: the shopper cannot name the attribute they are searching for. Read them against near-duplicate density in the catalogue, which is the counter-signal — high near-duplicate density means image matching will return confidently wrong variants.
How should visual and text search be combined on one surface rather than treated as alternatives?
Treat them as two query modes in one discovery surface: the image narrows to a visual neighbourhood, then text or facets resolve the nameable attributes inside it — size, price band, material. This is also the fallback route, since an unconfident image match can hand off to text search with extracted attributes pre-populated rather than showing a dead end.
What happens to the shopper journey when the image match returns no confident result?
The acceptable outcome is a named, honest fallback — visually similar items with the confidence stated, or a handoff into text search — never a padded result list presented as a match. Fallback rate is a per-category metric that belongs in the rubric, and what counts as acceptable depends on how graceful the fallback is: a pre-populated text handoff tolerates a much higher rate than a bare no-results page.
How do catalogue image quality and churn change the visual-vs-text answer over time?
Both degrade the rubric’s verdict. A category scored on professionally shot, slow-turning imagery can score differently after a supplier-image migration or a weekly seasonal refresh, because match confidence falls as index staleness and image inconsistency rise. Re-reading nameability and null-result signals per category each quarter keeps the routing decision current.
Five product categories where images outperform keywords
Fashion, furniture, home décor, consumer electronics, and automotive parts all share one trait: shoppers struggle to name what they recognize instantly.