Topic
Web data pipelines and LLM classification
Collecting structured data from the open web at scale, then classifying and extracting from it with LLMs, so that a team can work from a clean table instead of a browser.
Overview
A large share of useful business data is public but unusable as it stands: spread across maps listings, company websites, registers and job boards, in inconsistent formats and languages. Turning it into a table someone can act on takes three kinds of work, and I’ve done all three for clients and for my own products.
Acquisition. Running collection at scale through tools such as Apify and BeautifulSoup, with retries, quota handling and checkpoints, so that a run of tens of thousands of records can be resumed rather than restarted.
Resolution and cleaning. Deduplicating by URL and normalised name, deciding when two records are the same business, and recording why a row was dropped. Most of the quality of the final table is decided here.
Classification and extraction. Using an LLM to answer one narrow question per record, such as which category a company belongs to or who its decision-makers are, at temperature zero, in JSON, validated against a schema. The model never decides the structure of the output, only the values inside it.
The same discipline applies to structured data. My financial reporting engine detects the schema of any CSV or Excel file and converts it into one canonical long format, so that the variance analysis and charts downstream never depend on column names.
Projects
- 483signals
- 65scored
- 60companies
- 48leads
Morpheus
483 public signals in, 48 ranked buyers out.
- parents book classes
- 66vendor listings imported
BabyBrain
A two-sided marketplace with bookings, payments and payouts, built from the first commit.
- any wide export
- period, entity, metric, value
Universal Financial Reporting Engine
Built deliberately without AI, so every number is computed and none is generated.
Where I did this
- COO and executive business partner Apr 2024 to May 2025
Functional COO during my first year of university. I owned operations, risk frameworks and strategic planning, and prepared the data for the company's ML training pipeline.
- Technical partner 2025 to present
Technical partner on an FP&A automation product built with the firm's partners, for CFOs, finance directors and scale-up founders who still report out of spreadsheets. I'm responsible for the technical architecture and delivery.
- Lead engineer Jun 2026 to present
I built BabyBrain's marketplace platform from its first commit and was its only engineer for the first ten weeks.
Questions
When should you use an LLM to classify websites instead of rules?
When the categories depend on meaning rather than keywords, for example whether a company actually manufactures agricultural machinery or only sells parts for it. Rules are cheaper and should handle everything they can; the model handles the remainder, and its output should be constrained to a fixed set of labels.
What breaks first in a large scraping pipeline?
Usually not the scraper. It is duplicates, inconsistent company names, pages that render differently for bots, and sources that change their output format without notice. Deduplication and validation deserve as much engineering as collection.
Can a reporting pipeline really be schema-agnostic?
Partly. It can detect time columns, categories and measures automatically and normalise any table into one long format. What it cannot infer reliably is meaning, such as whether a numeric column is an ID, what the right reporting period is, and how accounting signs work. Those need explicit guards.