An audit engine and an autonomous agent that make a site legible to AI, not just to search engines.
The system has two halves that run in sequence. The audit engine crawls a live site with a real browser and produces two scores per page — an SEO score and an AI Readiness score. The fix agents then take that report, clone the site's repository, rewrite the flagged signals directly in code, and push the result back to GitHub.
01 Why AEO/GEO, not just SEO
This project is not a traditional SEO tool. Traditional SEO optimizes for search engine crawlers ranking a page in a list of blue links. This tool optimizes for a broader and newer audience: the answer engines and generative models that read a page once, extract a claim from it, and present that claim directly to a user — often without a click-through.
Foundational crawlability and indexing — canonical tags, sitemaps, robots rules, hreflang. Required for a page to be found at all, by any engine.
Answer Engine Optimization — structuring content so a system like an AI Overview or voice assistant can lift a direct answer: FAQ schema, a speakable snippet, a clear heading hierarchy.
Generative Engine Optimization — making a page easy for an LLM to cite: HowTo schema, well-scoped keywords, clean semantic structure that survives being summarized.
02 The two‑part system
These are two separate codebases that hand off a single CSV file between them.
Audit Engine
Crawls the live, rendered site with a headless browser. Measures real technical, on-page, and AI-readiness signals per page. Scores each page two ways and exports a CSV report.
seo_report_<uuid>.csv — consumed by →Fix Agents
Clones the site's GitHub repository. Reads the CSV row for each URL, generates the corrected title/meta/schema/canonical/etc. per page, writes it into the actual source files, and commits + pushes.
The two parts are independently runnable — you can audit a site without ever touching its repo, and you can run the fix agents against any CSV in the right shape, not only one this audit engine produced. The CSV handoff contract defines that shape exactly.
03 Discovery & fetchingPART 1
robots.txt once per run (disallow rules, listed sitemaps), then locates the sitemap via those listed URLs or common paths like sitemap.xml / sitemap_index.xml.X-Robots-Tag header — and to compare word counts against the rendered version.04 What gets measuredPART 1
Every number the scoring layer uses is either directly observed by the browser or computed from what it observed — nothing is estimated or fabricated when a real measurement isn't available.
- Core Web Vitals: LCP and CLS are captured via
PerformanceObserverinjected before page scripts run — these are genuinely measurable in a headless browser. INP is always reported asnot_measured, since it requires a real user interaction a headless crawl never produces. - JS-dependency ratio: raw (non-JS) word count divided by rendered word count — a low ratio means the page's content is mostly injected by JavaScript, a real risk for crawlers that don't execute it.
- Status codes & redirects: taken from the actual navigation response chain Playwright observes, cross-checked against a plain HTTP fetch.
- Timing: TTFB and load time come from the browser's own Navigation Timing API, not a synthetic stopwatch.
05 Scoring: SEO ScorePART 1
A weighted composite across three categories, each built from individually-verifiable checks with a real measured_value attached, not just a pass/fail label.
| Category | Weight | Checks |
|---|---|---|
| Technical | 40% | Canonical, robots.txt, sitemap, noindex, redirects, HTTP status, HTTPS, page speed, Core Web Vitals, mobile/viewport, JS rendering dependency, hreflang, structured data |
| On-page | 40% | Title, meta description, headings, content length (thresholds vary by page type), readability, keywords, internal links, alt text, JSON-LD, Open Graph, Twitter tags |
| Off-page | 20% | Reserved — not populated by the audit engine; this is the category the manual off-page process (see <a href="#offpage-agent" class="text-blue-600">Off-page</a>) is meant to eventually feed. |
Content-length thresholds are page-type-aware — a homepage, an article, and a standard page are held to different word-count bars rather than one universal number.
06 Scoring: AI Readiness ScorePART 1
A deliberately separate scoring system from the SEO score. It estimates how prepared a page is to be understood, summarized, and cited by AI systems — not how it would rank in a search results page.
actual_ai_visibility: not_measured. Real citation and mention data from AI systems requires live query access this crawler doesn't have. This score reflects readiness proxies only, and says so explicitly rather than implying otherwise.| Category | What it measures |
|---|---|
| Semantic understanding | Topic clarity, title/H1 coherence, content completeness and depth |
| Content answerability | Question-style headings actually answered nearby, direct-answer presence, definitions, FAQ coverage |
| Entity understanding | Organization / person / product / place / brand entities detected across text, title, and headings |
| Trust & authority | Author expertise and credentials (only where relevant to the page type), about/contact pages, HTTPS |
| Citation potential | Unique data points, quoted material, source attribution, first-hand research language |
| Retrieval & crawlability | Indexability, canonical validity, and a genuine raw-vs-rendered HTML comparison to catch JS-gated content |
| Structured data | Schema types relevant to this specific page type, not every schema type on every page |
| Freshness | Content age parsed from meta tags, <code><time></code> elements, or visible text across a wide range of date formats |
Two design rules run through all eight categories:
- Weights shift by page type. A homepage is weighted toward entity and retrieval; an article is weighted toward answerability and citation. 14 page types are supported, each with its own weight profile merged over a shared default.
- Unknown is never zero. If a signal genuinely can't be measured — no date on the page, no target query supplied, no author byline where one isn't expected — it's excluded from the weighted average and its weight is redistributed across what was actually measured, rather than counted against the page.
An optional LLM pass (Claude, via the same Anthropic client the fix agents' orchestrator can use) can enhance a handful of the semantic sub-scores by reading a content excerpt. That result is blended with the deterministic score rather than replacing it, and the entire pipeline works correctly with use_ai=False.
07 Page type classificationPART 1
Before any AI Readiness scoring happens, each page is classified into one of 14 types (homepage, article, product, documentation, forum, review, organization, and so on) using a deterministic vote across URL structure, schema.org types already present, and content-structure signals like link density and heading patterns — no LLM call, so it runs on every page instantly.
This classification is what lets the scoring system avoid two common mistakes: penalizing a homepage for not answering FAQ-style questions, and penalizing a documentation page for missing Product schema it was never going to have.
08 InterfacesPART 1
The audit engine is exposed through a Gradio app with three primary actions:
| Action | api_name | Returns |
|---|---|---|
| Analyze SEO | analyze_seo_async | Formatted markdown SEO report |
| Analyze AI Visibility | analyze_ai_async | Formatted markdown AI Readiness report |
| Structured SEO JSON | analyze_seo_json_async | The full structured response object, including per-page checks and the CSV path — this is the programmatic entry point |
All three share the same underlying async functions (run_seo_analysis_fastapi,run_full_seo_audit, run_ai_visibility_analysis), so nothing is fetched or rendered twice per run.
09 How a fix run worksPART 2
A fix run is triggered from a single form (or a direct POST) and executes as one sequential pipeline per repository, starting from the CSV the audit engine produced.
BRANCH.package.json and repo layout to identify Next.js (App or Pages Router), Jekyll, or static — this decides whether per-page fixes are possible.origin/BRANCH, so a technical-only run never depends on on-page having run first.🟢 On-page agent
autonomousOwns everything that lives inside the page and its content.
In scope
- Title tag & meta description
- Heading structure notes (H1/H2/H3)
- Keyword set, per page
- Internal link suggestions (keyword overlap)
- Alt-text generation for flagged images
- WebPage / FAQPage / HowTo JSON-LD
- Speakable snippet, Open Graph, Twitter tags
Explicitly out of scope
- robots.txt, sitemap
- Canonical tags, noindex
- Redirects, hreflang
🔵 Technical agent
autonomousOwns everything that makes the site crawlable, renderable, and trustworthy to an engine — the layer AEO/GEO signals depend on.
In scope
- robots.txt
- sitemap.xml and native Next.js sitemap route
- Canonical tag & noindex handling
- hreflang (when locales are supplied)
- Redirect scaffold check
- Diagnostic report for what can't be blind-fixed
Flagged, not fixed
- Core Web Vitals (already measured in Part 1; needs a live re-check post-deploy)
- Page speed
- 404 / 5xx status codes
- JS rendering, mobile rendering
🟠 Off-page
manual, plannedOff-page signals — citations, backlinks, brand mentions, third-party listings — are not something a repo edit can produce, and are not automated in this version.
In scope
Out of scope for automation
- Citations, backlinks, brand mentions, third-party listings
- This is handled manually until an automated approach is deliberately designed.
- The SEO score's reserved 20% off-page weight is meant to eventually reflect this.
10 Fix Agents APIPART 2
| Endpoint | Method | Purpose |
|---|---|---|
/ | GET | Serves the operator form used to trigger a run. |
/health | GET | Liveness check. Returns <code>{"status":"healthy"}</code>. |
/run | POST | Executes a full optimization run synchronously and returns stdout/stderr from the orchestrator. |
POST /run — form fields
| Field | Type | Default | |
|---|---|---|---|
github_token | string | — | required |
repo_url | string | — | required |
site_base | string | — | required |
csv_file | file (.csv) | — | required |
branch | string | main | optional |
git_username | string | SEO-Auto-Fix-Bot | optional |
git_email | string | seo-bot@example.com | optional |
agents | string | onpage,technical | optional |
locales | string, comma-separated | "" | optional |
The request blocks for the duration of the run (10-minute server-side timeout) and returns the orchestrator's return code, stdout, and stderr as JSON.
// response shape
{
"returncode": 0,
"stdout": "...",
"stderr": "...",
"success": true
}11 CSV handoff contract
This is the exact shape Part 1 exports and Part 2 expects — one row per URL. If you're feeding the fix agents a CSV from somewhere other than the audit engine, matching this schema is what makes it work.
| Column | Used by |
|---|---|
url | Both — required, resolves the target file |
title, meta_description | On-page agent |
h1_count, heading_order | On-page agent |
missing_alt_tags, total_images | On-page agent |
top_keywords, seo_suggestions | On-page agent |
canonical_tag, robots_meta | Technical agent |
viewport_present | Technical agent (diagnostic report) |
schema_types, opengraph_tags, twitter_tags | On-page agent |
word_count, readability_score, grammar_errors, text_to_html_ratio, seo_score | On-page agent (context for the model) |
The audit engine's export includes many additional columns (Core Web Vitals, AI Readiness sub-scores, redirect chains, and so on) for reporting purposes — the fix agents simply ignore any column they don't need.
12 Environment variables
main.Part 2onpage, technical, or both.Part 213 Write safety & idempotencyPART 2
Every file write goes through the same guard rails, regardless of which agent is writing:
- Marker-scoped blocks. Generated content lives between explicit start/end markers (
ONPAGE-AUTO-FIX,TECHNICAL-CANONICAL) so a second run replaces its own output instead of duplicating it. - Truncation guard.
safe_writerefuses to overwrite a file if the model's output is empty or suspiciously shorter than the original — the original file is kept and the failure is logged, not silently swallowed. - Single JSX root. Schema injection locates the component's actual return statement by scanning for the largest balanced-parenthesis span, rather than assuming the first or last
return (in the file — this avoids corrupting early guard clauses or helper functions. - Client component awareness. Files starting with
"use client"are skipped for server-only metadata exports, since that syntax is invalid there.
14 Supported stacks
| Stack | Audit (Part 1) | Per-page fixes (Part 2) | Site-level fixes (Part 2) |
|---|---|---|---|
| Any live site | Full — the crawler only needs a rendered URL | — | — |
| Next.js — App Router | Full | Full | Full, incl. native <code>sitemap.js</code> |
| Next.js — Pages Router | Full | Full | Full |
| Jekyll | Full | Not wired up | robots.txt, sitemap.xml |
| Static | Full | Not wired up | robots.txt, sitemap.xml |
The audit engine (Part 1) works against any live, publicly reachable site regardless of stack — it never touches source code. The per-page fix agents (Part 2) need a repo it recognizes to know where to write.
15 Roadmap
Shopify support
On-page and technical signals mapped to theme/metafield equivalents rather than JSX — likely via the Admin API rather than a cloned Liquid repo.
WordPress support
Same agent logic, applied through the REST API (and existing SEO plugin fields where present) instead of direct theme file edits.
Off-page automation
Remains a manual process until there's a reliable, verifiable way to automate citation and mention building — and to responsibly populate the SEO score's off-page category.
