Selected Work
PromptCloud / 42Signals

Automated Product Matching Pipeline

Fashion Retail Client · Digital Shelf Analytics

The same sneaker has a different title, a different photo, and a different SKU on every retailer's site. Comparing prices across them starts with solving that mismatch.

ContextPromptCloud / 42Signals
DomainFashion Retail
StackPython · Elasticsearch
StatusClient Work

Digital shelf analytics sounds simple until you try to compare "the same product" across two retailers and realize there's no shared ID connecting them. One site calls it "Men's Classic Leather Sneaker - White," another calls it "White Leather Trainers, Men"; the photos are shot differently, and the size and color attributes are structured differently too. A price comparison built on that mismatch isn't comparing the same product at all.

The matching layer is the unglamorous, load-bearing piece underneath every other kind of retail intelligence: pricing, assortment gaps, competitor benchmarking are all only as trustworthy as the matching beneath them. So I built a pipeline that reconciles listings across competing fashion retailers despite inconsistent titles, images, and attributes, turning approximate comparisons into genuinely apples-to-apples ones.

A false match silently corrupts everything built on top of it. A missed match is just a gap.

That asymmetry shaped nearly every decision in the matching logic. It's tempting to tune a matcher for the highest overall accuracy, but the cost of a false positive and a false negative aren't the same, so the system deliberately errs toward caution: when it isn't confident two listings are the same product, it says so, rather than guessing.

The tradeoff is a familiar one to anyone who's built entity resolution: optimizing for accuracy on paper isn't the same as optimizing for the outcomes that actually matter downstream.

PythonData ValidationElasticsearch

Pipeline & Results

The stage-by-stage shape of the matching pipeline and the aggregate outcomes it produced — matching 3,300 SKUs across 9 Indian e-commerce platforms.

Architecture

Input · SKU Master
Search & Discovery
Extraction & Matching
Retry & Error Handling
AI Diagnostics
Output & Monitoring

Tech: Python · Playwright · BeautifulSoup · Groq API · ScraperAPI · Google Sheets API

3,300Total SKUs
2,872URLs Matched
87.0%Overall Match Rate
92.3%Automation Rate
68%Reduction in Manual Effort

Match Rate by Platform

Flipkart100.0%
Amazon.in99.1%
Luxury TataCliq92.1%
TataCliq90.3%
Nykaa Fashion88.2%
Shoppers Stop87.7%
AJIO82.6%
Lifestyle81.5%
Myntra68.4%

Live Monitoring — Status Distribution

Matched — 2,872 In Progress — 241 Failed — 187

AJIO — Improvement After Secondary Search

Primary Search Only 651 19.7% matched
+348 URLs
+ Google Search Fallback 999 30.3% matched

Working through a product matching problem?

Happy to talk through matching strategy, attribute extraction, or where precision and recall trade off against each other.