A size recommendation engine is software that tells an online shopper which labeled size to buy, and through the mid-2020s the engines that worked best made that call mainly from cohort behavior, meaning what shoppers with similar bodies previously bought and kept, rather than from size charts alone. Published vendor claims and retailer case accounts commonly associated deployed fit engines with reductions in size-driven returns in the twenty to thirty percent band, per the vendors' own reporting, a figure worth treating as promotional until replicated. The mechanics behind the recommendation, however, are straightforward enough to describe honestly.
What inputs does a size engine actually use?
Engines draw on four input classes, and their weight varies by engine design and by how much data a shopper has surrendered.
- Stated dimensions: height, weight, age, and bra or waist size entered by the shopper at the recommendation prompt, the fastest input and the least reliable.
- Purchase and return history: the account's past orders, what was kept, what was returned, and the stated return reason, which is the richest behavioral signal most retailers own.
- Cohort data: aggregated keep and return patterns of shoppers with similar measured dimensions across the same brand's garments, smoothing individual noise.
- Garment measurement data: the actual produced dimensions of each size in each style, which most engines extract from tech packs or by measuring samples, and which anchor the recommendation to a real object.
The garment side deserves emphasis. Recommendation quality collapses when the brand's size data describes the chart rather than the product, because the engine then recommends against intentions instead of reality. Retailers that measured their actual stock before deploying engines reported cleaner improvements, an unglamorous data hygiene finding that recurs across case discussions.
How does the engine reach a recommendation?
Designs differ, but the dominant pipelines resolve to one of three logics, and many commercial engines blend them.
- Rule-based matching: the engine compares the shopper's dimensions to the garment's measured dimensions and applies brand-specific ease rules, recommending the size whose intended body measurements the shopper most closely matches. Transparent and fast, but blind to how the shopper likes clothes to fit.
- Collaborative filtering: the engine finds shoppers with similar bodies or purchase histories and recommends what they kept, a nearest-neighbors logic that transfers fit preferences without anyone articulating them.
- Learned fit models: a statistical or machine-learned model trained on the brand's own outcome data, where the label is keep-or-return, scores each candidate size and picks the highest. This design learns idiosyncrasies, such as a brand whose sleeves run short, that rules never encode.
The learned approach is the mid-2020s default among major vendors, and its known weakness is cold start: a shopper with no history and a garment with no outcome data leaves the model with stated dimensions plus cohort priors, which is where recommendation accuracy is lowest and where most shopper skepticism is earned.
What does the evidence say about results?
Published numbers divide into two reliability classes. Vendor case studies, the majority of available evidence, report return-rate reductions and conversion lifts from named deployments, but they are selected successes reported by the selling party. Independent academic work, including studies through operations and data science venues with participation from universities such as Cornell, has examined recommendation effects on returns with more caution, generally finding real but smaller effects that depend heavily on data volume and on whether the engine actually changed the shopper's choice rather than merely adding friction.
One structural finding recurs across both literatures: engines reduce fit-motivated returns, meaning returns where the shopper cites size or fit, more reliably than they reduce total returns, because shoppers return items for many reasons a size engine cannot touch. Retailers evaluating engines on total return rate alone systematically under-measure or over-measure them.
Related stories: Body Scanning Gave Apparel Real Population Data, and Brands Still Barely Use It · PLM Software Tracks a Garment From Sketch to Invoice, Line by Auditable Line.
What do these engines cost, and who deploys them?
Commercial deployment has clustered into three models, compared below.
| Model | Typical economics | Who it suits |
|---|---|---|
| SaaS fit engine, third-party | Annual subscription, often banded by traffic | Mid-market brands without data science teams |
| Platform-native tools | Included or add-on within a major e-commerce platform | Brands standardizing on one commerce stack |
| In-house engines | Data science staffing plus data infrastructure | Large retailers with dense outcome data |
Vendor pricing has generally run from modest annual subscriptions for small catalogs to substantial enterprise agreements, per vendor communications, and the implementation effort is mostly measurement: capturing real garment dimensions across the catalog consumes more project time than integrating the widget.
Do shoppers trust the recommendation, and does it matter?
Trust is the adoption constraint the industry discusses least. A shopper who has been wrong by an engine once discounts the next recommendation heavily, and survey work on fit tool usage during the early 2020s repeatedly found significant shares of shoppers ignoring size tools entirely, relying instead on reviews and their own history with the brand. Engines respond with explanations, showing the dimensional reasoning behind a call, which measurably improves follow rates in vendor reporting. Letting the shopper override the call without penalty also matters, because an engine that fights a shopper who knows her size in a brand loses the account's trust faster than one that documents its reasoning and steps aside.
The practical read for 2026: size engines are worth deploying where return rates are high and garment measurement data is real, worth evaluating on fit-driven returns rather than totals, and worth treating as an input to the shopper's judgment rather than a verdict. They compress the size decision problem; they do not dissolve the deeper issue that labeled sizes vary across brands more than any engine can apologize for. Returns data closes the loop over time: every keep-or-return outcome a shopper generates improves the model's next call, which is why the engines' accuracy, unlike that of a size chart, compounds with use rather than aging toward irrelevance.
How should a brand evaluate engines before committing?
An evaluation protocol protects a brand from both vendor enthusiasm and its own impatience, and the structure is straightforward. Run the candidate engine silently on live traffic, recording its recommendation without showing it to shoppers, then compare recommended sizes against what those shoppers actually kept. This offline shadow period, run for a statistically meaningful stretch of orders, reveals whether the engine's calls correlate with real outcomes in the brand's own catalog before any customer faces the consequences.
The second test is measurement coverage. Ask what fraction of the catalog has verified produced dimensions, because the engine's effective scope is that fraction, whatever the contract says. A vendor promising a full-catalog solution for a brand that has measured a third of its styles is selling confidence, and the honest implementations start with the measured catalog and expand as measurement catches up.
Finally, agree the metric before launch. Fit-driven return rate on covered styles, measured against a matched control group, is the defensible primary metric; conversion lift and tool engagement are secondary and more gameable. Contracts written with performance fees tied to that primary metric align incentives, and vendors confident in their models have accepted such structures in the market. Brands that skip the shadow test and the metric agreement generally learn the engine's quality from their own return data a season later, which is the most expensive tuition in the fit technology field.
