Context
A complex industrial market where product knowledge lived in PDFs, supplier silos, and tribal memory. Buyers could not find, compare, or trust the information their decisions depended on.
The problem
Product data arrived in thousands of inconsistent formats, from structured feeds to scanned datasheets, with no shared taxonomy across suppliers.
Generic search treated industrial attributes as text. Relevance required understanding chemistry, specifications, units, and application context.
The system had to keep improving as new suppliers, categories, and data sources arrived — enrichment could not be a one-time migration.
The system
Critical decisions
Build an acquisition and enrichment pipeline, not a content team
Manual curation cannot keep pace with an industry catalog. Machine extraction with human review where confidence is low scales; pure human effort does not.
Treat taxonomy and entity resolution as the core asset
Search quality, comparison, and later AI features all depend on knowing that two differently-described products are the same thing.
Deterministic structure first, ML on top
Where a value can be parsed, parse it. Models fill the gaps and rank ambiguity — they do not replace structure that can be computed reliably.
Execution
Iterative rollout by category: ingest, structure, verify with domain experts, expose in search, measure engagement, expand. Relevance tuned against real buyer queries rather than synthetic benchmarks.
Outcome
- A durable structured-data and search layer for a fragmented industry.
- Enrichment throughput that scales with suppliers instead of headcount.
- [Quantified outcomes to confirm before publication]
Lessons
- In vertical AI, the moat is the structured data and workflow, not the model.
- Entity resolution is unglamorous and decisive.
- Relevance is a product decision expressed in engineering.
Have a similar problem?
Discuss an AI initiative