Guides

How to Scrape Thomasnet Supplier and Manufacturer Data

Thomasnet is the leading US industrial sourcing platform, which makes it a rich public source of supplier and manufacturer data: company, location, capabilities, certifications, and firmographics across every industrial category. This guide covers what a listing holds, where the data lives, the challenges of collecting it cleanly, and how to turn it into a supplier or lead list. The worked example uses a real Thomasnet search.

DADataScrape TeamSeptember 2, 2026Updated September 28, 20264 min read

What data a Thomasnet listing holds

Each supplier record is dense with the firmographics buyers actually screen on. On a real search, one company read: Custom Equipment Company in Charleston, SC, described as an ISO 9001:2000 certified manufacturer, designer, and distributor, with annual sales of 5 to 9.9 million dollars, 1 to 9 employees, founded in 1978, listed under the Domes heading. The company, location, description, capabilities and certifications, sales band, employee count, and year founded are the fields that matter for sourcing and lead work, and they come back per supplier.

Where the data actually lives

Thomasnet runs on a Next.js frontend, so the supplier data comes back in the app's data JSON, for example a suppliers search.json endpoint, whose pageProps carry a companies array. Each company object holds the name, an address with city and state, the description, the capabilities and headings, the annual sales band, the employee count, and the year founded, all as structured fields. Reading that JSON is cleaner than parsing rendered cards, and the company profile is the stable key you deduplicate on.

Skip the build, get the feed

Send us your Thomasnet target and we will scrape a real sample and send it back, no signup.

Get a free sample

The practical challenges

Firmographics are self-reported and vary in completeness, capabilities and certifications are nested lists rather than flat fields, and results page through many companies per category. Some detail is gated to a signed-in account, and the same supplier can appear across several categories, so deduplication matters. A scraper that works today can break after a build or endpoint change on the Next.js data path.

Turning it into a supplier or lead list

For procurement or outreach you want one row per company, deduplicated, captured per category and region, keeping the company, location, capabilities, certifications, sales band, employees, and year founded. Store each run with a date so you can track new suppliers entering a category, and keep the fields your workflow uses so the list stays workable rather than a dump of raw firmographics.

Build it yourself or have it delivered

For one category, reading the data JSON from a few searches works. At scale the paging, the nested capabilities, the gated detail, and the upkeep turn it into a standing maintenance job. A done-for-you feed is usually cheaper than the maintenance once you need more than a category or two. You tell us the categories and regions, and we deliver a clean, deduplicated supplier list and keep it running.

Where each field lives in the data JSON

Because Thomasnet is a Next.js app, its supplier data comes back under pageProps in the data JSON, in a companies array. Each company object carries name, an address object with city, state, and stateName, a description, a capabilities list and a headings list that hold the categories and specialties, annualSales as a band (for example $5 - 9.9 Mil), numberEmployees as a band, and yearFounded. The capabilities and certifications are nested lists rather than flat fields, so a clean record flattens the ones you care about. One practical note: the URL carries a build-id that rotates when Thomasnet redeploys, so a durable scraper reads the current build-id from the page rather than hard-coding it. The company profile is the stable key you deduplicate on.

Step by step: reading a Thomasnet search

Thomasnet runs on Next.js, so its supplier data comes back in the app's data JSON rather than the rendered page. Here it is on a real search, where each company came back with its location, certifications, and firmographics.

python
import requests

# Thomasnet's Next.js frontend loads suppliers from a data JSON endpoint.
url = "https://www.thomasnet.com/_sda/_next/data/<build-id>/suppliers/search.json"
params = {"searchterm": "Stamped Domes for Pressure Vessel"}
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"}

data = requests.get(url, params=params, headers=headers).json()
for c in data["pageProps"]["companies"]:
    print(c["name"], c["address"]["city"], c["address"]["state"],
          c["annualSales"], c["yearFounded"])
CompanyCustom Equipment Company
LocationCharleston, SC
CertificationISO 9001:2000 certified
Annual sales$5 - 9.9 Mil
Employees1-9
Founded1978

A sample of the clean data we deliver for one product.

The parse is straightforward, though the build-id in the URL rotates and the capabilities and certifications come as nested lists. The friction is that firmographics are self-reported and uneven, some detail is gated to a signed-in account, results page through many companies, and the same supplier spans categories, so you deduplicate on the company. We run all of that and hand back a clean, deduplicated supplier list with contact fields kept or stripped to match your needs.

Fields worth capturing from Thomasnet

  • Company name
  • City / state
  • Description
  • Capabilities
  • Certifications
  • Annual sales band
  • Employee count
  • Year founded
  • Category
  • Company profile URL

Frequently asked questions

Want a Thomasnet feed without the build?

Send us the products or pages you need, and we will deliver a clean feed and maintain it. Start with a free sample.

Get a Free Sample