Built for a manufacturer of custom production-line equipment. The client's name and industry-specific details are anonymized. Screenshots show the live platform running against its real market dataset; personal contact details are blurred.

  • Python 3.12
  • FastAPI
  • MySQL + SQLAlchemy
  • Playwright
  • OCR (Tesseract)
  • LLM Enrichment
  • Contact Enrichment APIs
  • OpenStreetMap / Overture
  • News RSS Monitoring
  • Leaflet Maps
  • Background Job System
13,700+ Companies Tracked

Discovered across directories, OCR'd trade publications, map data, and enrichment APIs — deduplicated into one canonical record each.

1,782 Qualified Leads

Companies scored into A/B/C tiers by product fit, revenue, contacts, and ideal-customer profile.

5,200+ Valid Contacts

Decision-maker contacts pulled through credit-gated enrichment and validated.

3,000+ Facilities Mapped

Headquarters and plants geocoded, footprint-measured, and plotted on interactive maps.

100% Cited Sources

Every AI-researched company profile stores the exact source URLs behind every claim.

The Brief

The client manufactures custom equipment for food-production lines. Their buyers are industrial manufacturers — but nobody sells a clean list of them. The companies hide behind parent corporations, holding groups, and brand names; the facilities that would actually house the equipment aren't listed anywhere; and the people who sign purchase orders don't answer info@ addresses. Sales prospecting meant weeks of manual research per territory, and the picture was stale before it was finished.

Project Type
Custom AI market research and lead intelligence platform.
Domain
Industrial B2B — finding every manufacturer that can use the client's custom equipment.
Data Sources
Industry directories, OCR'd trade publications, OpenStreetMap, enrichment APIs, company websites, news feeds, job postings.
Architecture
Scraper fleet → canonical deduplicated store → AI enrichment → scoring → web portal, all driven by one background job system.
Freshness
Recurring news sweeps, hiring scans, domain health checks, and re-scoring keep the dataset current.
Deliverable
A living web portal — scored leads, company dossiers, maps, and a prioritized call queue — plus CSV exports.

The Problem

Purchased lead lists fail industrial sales teams in three predictable ways. They're organized by legal entity, not by who actually buys — a subsidiary running two plants under a foreign parent shows up as noise, or not at all. They say nothing about facilities — and for equipment sales, the facility is the opportunity. And they're static — a list bought in January says nothing about the company that announced a plant expansion in March.

The client needed something different: a system that could find every company in the market from scattered, messy sources; untangle the corporate structures to reveal who owns what; research each facility down to the building footprint; identify the decision-makers with valid contact details; and keep watching — news, hiring, expansions — so the sales team always knows who to call next, and why.

The Solution

ThinkGenius built Breadcrumb — a market intelligence platform that treats an entire industry as one continuously-refreshed dataset. Scrapers and discovery jobs pull raw company sightings from industry directories, OCR'd trade publications, map data, and enrichment APIs. A deduplication pipeline collapses those sightings into one canonical record per real-world company while preserving full provenance — every source a company was ever seen in stays on the record. AI deep-search enrichment then builds a cited dossier for each company, a scoring model ranks them against the client's ideal customer profile, and a web portal turns the result into a daily sales instrument.

Breadcrumb market intelligence platform leads dashboard showing 1,782 qualified leads scored into tiers, with filters for state, source, email coverage, equipment fit, hiring signals, and news signals.
// The leads dashboard — every company in the market, deduplicated, scored, tiered, and filterable by fit, geography, contacts, hiring, and news signals.

How the Platform Works

Six subsystems, one canonical database. Each stage runs as a background job that the CLI and the web portal both call — so the whole pipeline is operable from one screen.

Discovery

Multi-Source Deep Search

Directory scrapers crawl industry listing sites; an OCR pipeline reads companies out of digital trade-publication editions page by page; OpenStreetMap queries surface facilities that never appear in any directory; and enrichment APIs sweep the market state by state. Four independent nets, because no single source sees the whole industry.

Canonical Store

Dedup With Full Provenance

Sightings match on normalized domain, then name + location, then fuzzy rules flagged for review. Merging never destroys evidence: every source sighting stays attached to the canonical company, so any field can be traced to where it was seen. Same-name corporate/brand splits are detected instead of blindly merged.

AI Enrichment

Cited Company Dossiers

An LLM research pass reads each company's website and public footprint and writes a structured profile — what they make, manufacturing footprint, parent company, revenue estimate, buying signals, and equipment fit — with every claim backed by stored source URLs a human can audit in seconds.

Corporate Structure

Entity & Brand Mapping

A discovery job untangles who owns what: parent corporations, sister companies, brands, and historic names — including acquisition years. A subsidiary stops hiding behind its holding group, and the platform recommends related entities worth adding to the system.

Facilities

Plant-Level Research

Every headquarters and plant is geocoded and plotted on interactive and satellite maps. Building footprints from OpenStreetMap and Overture Maps estimate facility square footage — a direct proxy for production capacity and equipment budget — with a per-facility fit score.

Decision-Makers

Contact Discovery, Credit-Safe

Contacts — plant managers, operations VPs, owners — are pulled through credit-gated enrichment only for companies that score into qualified tiers, then validated. Companies that already have contacts under a related domain are skipped automatically: zero wasted credits.

Company Dossiers With Verified Sources

The center of the platform is the company page: an AI-researched dossier that reads like an analyst wrote it. What the company makes, where it manufactures, who owns it, estimated revenue and headcount, and — line by line — which of the client's machines fit which of its production lines. Below every profile sits the sources list: the exact URLs the research pass drew from. When the AI says a company runs 28 plants across three continents under a private-capital parent, every one of those claims links back to where it was found.

AI-generated company dossier in the Breadcrumb platform showing an eleven-brand portfolio, revenue, plant count, parent company, and a line-item equipment-fit list mapping each machine type to specific production lines.
// A company dossier — an 11-brand, 28-plant manufacturer profiled by AI, with each machine type mapped to the specific production lines it fits.
Breadcrumb company dossier showing per-machine equipment-fit reasoning, an AI-written sales-fit assessment, and a numbered list of cited source URLs backing every claim.
// Line-item fit reasoning, the AI's overall fit assessment, and the sources list — every claim traceable to the URL it came from.

Breaking Down Corporate Structures

Industrial markets are full of camouflage: a billion-dollar manufacturer selling through eleven brand names, operating nineteen companies across three continents, held by a private-capital parent. Breadcrumb's entity discovery maps the whole family — parent, sisters, operating companies, brands, and historic names — and cross-references which relatives are already in the system. Above it, the facilities table lists every plant worldwide with its role, per-facility fit score, and estimated square footage from building-footprint data.

Breadcrumb platform showing an international facilities table with per-plant fit scores and estimated square footage across Canada, Austria, Belgium, Germany, and the UK, above a 33-entity corporate family tree with parent company, sister company, nineteen operating companies, and eleven brands.
// A 33-entity corporate family — private-capital parent, 19 operating companies, 11 brands — beneath the worldwide facilities table with per-plant fit scores and square footage.

The Market on a Map

Territory planning happens visually. The pin map plots every qualified company headquarters and plant across the continent — clustered, tier-filterable, and toggleable between leads and all facilities. A sales trip through a region starts with one look at where the density is.

Interactive North America pin map in the Breadcrumb platform plotting over 3,000 company headquarters and plants, clustered by region and filterable by lead tier.
// 3,000+ headquarters and plants plotted and clustered — the entire addressable market on one map.

Always Watching: News, Hiring, and Freshness

A lead list is a snapshot; a market intelligence platform is a feed. Recurring jobs keep the dataset current and re-rank the call queue as the market moves.

News Signals

Company News Sweeps

Recurring news monitoring across every tracked company and its related entities — expansions, acquisitions, product launches, leadership changes. Headlines are classified as growth or risk signals and attached to the company record, so reps open a call already knowing the news.

Hiring Signals

Job-Posting Intelligence

The platform scans hiring activity for target roles — a company hiring plant managers, maintenance techs, or line operators is a company investing in production capacity. Hiring flags feed the lead score and the call queue directly.

Prioritization

The Hot List

News, hiring, and tier blend into a single ranked answer to the only question that matters at 8 a.m.: who do we call today, and why? Each row shows the reason, the best decision-maker contact, and the latest headline.

Data Hygiene

Domain Health & Dead-Link Repair

A monthly probe checks every company website at the DNS/connection/SSL level. Failures — often the fingerprint of a rebrand or acquisition — get flagged, and a repair job finds the new domain and recovers contacts that would otherwise be lost.

Dedup Ops

Duplicate Merging & Re-Scoring

Preview-first merge tooling collapses duplicate rows that accumulate across sources — moving every contact, listing, news item, and score history to the surviving record — followed by a full re-score so tiers stay honest.

Operations

One-Click Pipeline Jobs

Every pipeline stage — discovery, ingest, enrichment, contacts, geocoding, hiring, news, scoring, export — runs as a one-click background job with live streaming logs, API budget meters, and resume-safe re-runs that never re-spend credits.

Breadcrumb Hot List call queue blending news signals, hiring activity, and lead tier into a ranked list, each row showing why to call now, the best decision-maker contact, and the latest company headline.
// The Hot List — news + hiring + tier blended into one ranked call queue, with the best contact and latest headline per company. (Contact details blurred.)
Breadcrumb pipeline actions page showing one-click background jobs for discovery, domain health checks, dead-domain repair, duplicate merging, re-scoring, and facility footprint measurement, with live API budget meters.
// Pipeline operations — every stage a one-click, resume-safe background job, with live budget meters for every external API.

Adaptable to Any Industry

Nothing in the architecture knows about food production. The platform is three swappable layers on top of industry-agnostic machinery:

  • Sources — which directories, publications, map queries, and APIs define the market. Swapping industries means swapping this list.
  • Ideal customer profile — the scoring model's definition of a great buyer: product categories, revenue bands, facility types, target roles.
  • Fit logic — what the AI enrichment looks for when it reads a company: which of your products map to which of their operations.

The same system that maps industrial manufacturers for an equipment maker could map logistics operators for a warehouse-automation vendor, medical practices for a device company, or contractors for a building-materials brand. Discovery, deduplication with provenance, cited AI enrichment, corporate structure mapping, facility research, contact discovery, and freshness pipelines carry over unchanged.

The Stack

  • Python 3.12
  • FastAPI + Uvicorn
  • MySQL 8 · SQLAlchemy · Alembic
  • Playwright (async)
  • Tesseract OCR + PyMuPDF
  • LLM research & structuring passes
  • Contact enrichment APIs (credit-gated)
  • OpenStreetMap Overpass + Overture Maps
  • News RSS monitoring
  • Leaflet interactive & satellite maps
  • Server-rendered portal + background job runner
  • CSV exports

Why It Matters

Most companies buy market data; very few own their market's data. The difference compounds. A purchased list decays from the day it arrives — companies rebrand, get acquired, open plants, change managers. A platform like Breadcrumb watches all of it: every record traceable to its sources, every score explainable, every signal feeding tomorrow's call queue. The sales team stopped researching and started calling — and the dataset underneath them gets better every week instead of worse.

Want a Market Intelligence Platform for Your Industry?

ThinkGenius builds custom AI research platforms that map your entire target market — companies, corporate structures, facilities, decision-makers — and keep it fresh automatically. Tell me who you sell to, and we'll scope the system that finds them.