Madhav Sharma

Topic

Web data pipelines and LLM classification

Collecting structured data from the open web at scale, then classifying and extracting from it with LLMs, so that a team can work from a clean table instead of a browser.

Overview

A large share of useful business data is public but unusable as it stands: spread across maps listings, company websites, registers and job boards, in inconsistent formats and languages. Turning it into a table someone can act on takes three kinds of work, and I’ve done all three for clients and for my own products.

Acquisition. Running collection at scale through tools such as Apify and BeautifulSoup, with retries, quota handling and checkpoints, so that a run of tens of thousands of records can be resumed rather than restarted.

Resolution and cleaning. Deduplicating by URL and normalised name, deciding when two records are the same business, and recording why a row was dropped. Most of the quality of the final table is decided here.

Classification and extraction. Using an LLM to answer one narrow question per record, such as which category a company belongs to or who its decision-makers are, at temperature zero, in JSON, validated against a schema. The model never decides the structure of the output, only the values inside it.

The same discipline applies to structured data. My financial reporting engine detects the schema of any CSV or Excel file and converts it into one canonical long format, so that the variance analysis and charts downstream never depend on column names.

Projects

Where I did this

Questions

When should you use an LLM to classify websites instead of rules?

When the categories depend on meaning rather than keywords, for example whether a company actually manufactures agricultural machinery or only sells parts for it. Rules are cheaper and should handle everything they can; the model handles the remainder, and its output should be constrained to a fixed set of labels.

What breaks first in a large scraping pipeline?

Usually not the scraper. It is duplicates, inconsistent company names, pages that render differently for bots, and sources that change their output format without notice. Deduplication and validation deserve as much engineering as collection.

Can a reporting pipeline really be schema-agnostic?

Partly. It can detect time columns, categories and measures automatically and normalise any table into one long format. What it cannot infer reliably is meaning, such as whether a numeric column is an ID, what the right reporting period is, and how accounting signs work. Those need explicit guards.

Related topics