/ Selected work

Making industrial data usable by machines

Industrial / chemical distribution Founder-operator — platform build

Context

A complex industrial market where product knowledge lived in PDFs, supplier silos, and tribal memory. Buyers could not find, compare, or trust the information their decisions depended on.

The problem

Product data arrived in thousands of inconsistent formats, from structured feeds to scanned datasheets, with no shared taxonomy across suppliers.

Generic search treated industrial attributes as text. Relevance required understanding chemistry, specifications, units, and application context.

The system had to keep improving as new suppliers, categories, and data sources arrived — enrichment could not be a one-time migration.

The system

Supplier sources feeds · PDFs · sites Acquisition ingest + extract Enrichment taxonomy · entities Human review low-confidence only Search & API relevance layer
Architecture — making industrial data usable by machines

Critical decisions

Build an acquisition and enrichment pipeline, not a content team

Manual curation cannot keep pace with an industry catalog. Machine extraction with human review where confidence is low scales; pure human effort does not.

Treat taxonomy and entity resolution as the core asset

Search quality, comparison, and later AI features all depend on knowing that two differently-described products are the same thing.

Deterministic structure first, ML on top

Where a value can be parsed, parse it. Models fill the gaps and rank ambiguity — they do not replace structure that can be computed reliably.

Execution

Iterative rollout by category: ingest, structure, verify with domain experts, expose in search, measure engagement, expand. Relevance tuned against real buyer queries rather than synthetic benchmarks.

Outcome

  • A durable structured-data and search layer for a fragmented industry.
  • Enrichment throughput that scales with suppliers instead of headcount.
  • [Quantified outcomes to confirm before publication]

Lessons

  1. In vertical AI, the moat is the structured data and workflow, not the model.
  2. Entity resolution is unglamorous and decisive.
  3. Relevance is a product decision expressed in engineering.

Have a similar problem?

Discuss an AI initiative