Madhav Sharma
Side
BuildGrow
Type
Product
My role
Designed and built it
When
Oct 2026 to present
Status
In progress

Raven

An agent for SEO and AI-search work that can only report what it can cite. Every finding has to point at a source a tool fetched in the same run, or it is rejected before it reaches the report.

skills the agent chooses from, including AI-search analysis
7
Checked against the data
modelled cost of a standard run, before caching changes
$0.55 to $0.95
Checked against the data
hard cost caps for quick, standard and deep runs
$0.60 / $1.50 / $4
Checked against the data

Stack

  • TypeScript
  • Pi agent runtime
  • Next.js
  • Supabase Postgres
  • Modal
  • Playwright
raven / evidence gateSEO and AI search
  1. tool calls store sources
  2. only cited findings pass
Schematic of the rule Raven enforces: a finding can only be recorded if it points at a source a tool fetched in the same run. Anything uncited is rejected.

Search work is moving from ten blue links to answers written by AI assistants. The tools for auditing it mostly produce confident recommendations with nothing behind them, and an AI agent doing the same work has an obvious failure mode: it will invent a problem, or a fix, and describe it fluently.

Raven is built around refusing to do that. You give it a task, such as “audit this site for AI search” or “compare us with these three competitors”, and it chooses the skills it needs, inspects real pages, runs searches, and returns findings. Each finding is split into what it observed, what it infers and what it recommends, and each observation links to a stored source.

The evidence gate

It is the rule that runs through most of my AI work, enforced in code at the one point where it matters: the moment something becomes a claim.

What it knows how to do

The agent chooses from seven skills: a site audit, technical SEO, AI-search analysis, content analysis, web research, competitor analysis and a persona prompt generator. The AI-search skill separates the crawlers that retrieve pages for live answers from the ones that collect training data, checks whether pages answer their own question directly under each heading, and is explicitly forbidden from promising that a page “will be cited”, because nobody can know that.

The persona generator takes a profile of a buyer, with role, industry, stage, pains, objections and vocabulary, and writes the Google searches and AI-assistant prompts that person would actually type. When live search data is available the prompts are grounded in it; otherwise they are labelled as hypotheses.

Charts in the reports are computed in code from the captured pages and findings. The model never produces a number.

How it is tested

An agent that is never measured drifts. Raven ships with a test site full of planted defects: a page missing its title and description, a wrong canonical with a noindex, broken structured data, a page that renders only in JavaScript, an article that never answers its own question, a robots file that blocks AI crawlers, and a broken link in the navigation. An evaluation runner scores how many of them a run finds.

Built to stay inside a budget

Every run has a hard cost cap by depth, and the modelled cost of a standard run is $0.55 to $0.95. The web app runs on Vercel, data lives in Postgres with row-level security, and long runs execute on a separate worker. Live progress is streamed from database rows rather than from memory, so it keeps working when the web app is spread across several serverless instances. I learned that one the hard way on an earlier product.

Where it stands

Raven is deployed with its full back end. It has not yet been scored against the planted-defect site with a production model, so there are no accuracy figures to publish yet. Those come next.

More work on the same problems