Logo
Hashfor
✦ Documentation

An audit engine and an autonomous agent that make a site legible to AI, not just to search engines.

The system has two halves that run in sequence. The audit engine crawls a live site with a real browser and produces two scores per page — an SEO score and an AI Readiness score. The fix agents then take that report, clone the site's repository, rewrite the flagged signals directly in code, and push the result back to GitHub.

Audit runtime Playwright + GradioFix runtime FastAPI + subprocess orchestratorStack support Next.js (App & Pages Router)

01 Why AEO/GEO, not just SEO

This project is not a traditional SEO tool. Traditional SEO optimizes for search engine crawlers ranking a page in a list of blue links. This tool optimizes for a broader and newer audience: the answer engines and generative models that read a page once, extract a claim from it, and present that claim directly to a user — often without a click-through.

SEO

Foundational crawlability and indexing — canonical tags, sitemaps, robots rules, hreflang. Required for a page to be found at all, by any engine.

AEO

Answer Engine Optimization — structuring content so a system like an AI Overview or voice assistant can lift a direct answer: FAQ schema, a speakable snippet, a clear heading hierarchy.

GEO

Generative Engine Optimization — making a page easy for an LLM to cite: HowTo schema, well-scoped keywords, clean semantic structure that survives being summarized.

In practice: the on-page agent's FAQ schema, HowTo schema, and speakable snippet exist specifically for AEO/GEO. The technical agent's canonical, sitemap, and robots work exist so those signals are actually reachable and trusted in the first place. Neither layer works well without the other.

02 The two‑part system

These are two separate codebases that hand off a single CSV file between them.

PART 1

Audit Engine

Crawls the live, rendered site with a headless browser. Measures real technical, on-page, and AI-readiness signals per page. Scores each page two ways and exports a CSV report.

produces seo_report_<uuid>.csv — consumed by →
PART 2

Fix Agents

Clones the site's GitHub repository. Reads the CSV row for each URL, generates the corrected title/meta/schema/canonical/etc. per page, writes it into the actual source files, and commits + pushes.

The two parts are independently runnable — you can audit a site without ever touching its repo, and you can run the fix agents against any CSV in the right shape, not only one this audit engine produced. The CSV handoff contract defines that shape exactly.

03 Discovery & fetchingPART 1

01
Site-level checks
Fetches and parses robots.txt once per run (disallow rules, listed sitemaps), then locates the sitemap via those listed URLs or common paths like sitemap.xml / sitemap_index.xml.
02
URL discovery
Tries every likely sitemap path in parallel, filters out XML/RSS/feed files and non-HTML assets, and keeps only real pages. If no sitemap yields anything, falls back to a light Playwright crawl of internal links from the homepage.
03
Rendered fetch
Each URL is opened in a real headless Chromium context (images/CSS blocked for speed). Captures the full rendered DOM: title, all meta tags, headings, images with alt text, links, JSON-LD schemas, hreflang alternates, and real navigation timing.
04
Raw HTTP cross-check
Separately fetches each URL with no JavaScript execution to capture the true HTTP status code, redirect chain, and X-Robots-Tag header — and to compare word counts against the rendered version.
05
Analysis
Parses the rendered HTML for structure, computes keyword density, readability (Flesch reading ease), and — where the optional dependency is installed — a grammar-error count.

04 What gets measuredPART 1

Every number the scoring layer uses is either directly observed by the browser or computed from what it observed — nothing is estimated or fabricated when a real measurement isn't available.

  • Core Web Vitals: LCP and CLS are captured via PerformanceObserver injected before page scripts run — these are genuinely measurable in a headless browser. INP is always reported as not_measured, since it requires a real user interaction a headless crawl never produces.
  • JS-dependency ratio: raw (non-JS) word count divided by rendered word count — a low ratio means the page's content is mostly injected by JavaScript, a real risk for crawlers that don't execute it.
  • Status codes & redirects: taken from the actual navigation response chain Playwright observes, cross-checked against a plain HTTP fetch.
  • Timing: TTFB and load time come from the browser's own Navigation Timing API, not a synthetic stopwatch.

05 Scoring: SEO ScorePART 1

A weighted composite across three categories, each built from individually-verifiable checks with a real measured_value attached, not just a pass/fail label.

CategoryWeightChecks
Technical40%Canonical, robots.txt, sitemap, noindex, redirects, HTTP status, HTTPS, page speed, Core Web Vitals, mobile/viewport, JS rendering dependency, hreflang, structured data
On-page40%Title, meta description, headings, content length (thresholds vary by page type), readability, keywords, internal links, alt text, JSON-LD, Open Graph, Twitter tags
Off-page20%Reserved — not populated by the audit engine; this is the category the manual off-page process (see <a href="#offpage-agent" class="text-blue-600">Off-page</a>) is meant to eventually feed.

Content-length thresholds are page-type-aware — a homepage, an article, and a standard page are held to different word-count bars rather than one universal number.

06 Scoring: AI Readiness ScorePART 1

A deliberately separate scoring system from the SEO score. It estimates how prepared a page is to be understood, summarized, and cited by AI systems — not how it would rank in a search results page.

What this is not: the response always includes actual_ai_visibility: not_measured. Real citation and mention data from AI systems requires live query access this crawler doesn't have. This score reflects readiness proxies only, and says so explicitly rather than implying otherwise.
CategoryWhat it measures
Semantic understandingTopic clarity, title/H1 coherence, content completeness and depth
Content answerabilityQuestion-style headings actually answered nearby, direct-answer presence, definitions, FAQ coverage
Entity understandingOrganization / person / product / place / brand entities detected across text, title, and headings
Trust & authorityAuthor expertise and credentials (only where relevant to the page type), about/contact pages, HTTPS
Citation potentialUnique data points, quoted material, source attribution, first-hand research language
Retrieval & crawlabilityIndexability, canonical validity, and a genuine raw-vs-rendered HTML comparison to catch JS-gated content
Structured dataSchema types relevant to this specific page type, not every schema type on every page
FreshnessContent age parsed from meta tags, <code>&lt;time&gt;</code> elements, or visible text across a wide range of date formats

Two design rules run through all eight categories:

  • Weights shift by page type. A homepage is weighted toward entity and retrieval; an article is weighted toward answerability and citation. 14 page types are supported, each with its own weight profile merged over a shared default.
  • Unknown is never zero. If a signal genuinely can't be measured — no date on the page, no target query supplied, no author byline where one isn't expected — it's excluded from the weighted average and its weight is redistributed across what was actually measured, rather than counted against the page.

An optional LLM pass (Claude, via the same Anthropic client the fix agents' orchestrator can use) can enhance a handful of the semantic sub-scores by reading a content excerpt. That result is blended with the deterministic score rather than replacing it, and the entire pipeline works correctly with use_ai=False.

07 Page type classificationPART 1

Before any AI Readiness scoring happens, each page is classified into one of 14 types (homepage, article, product, documentation, forum, review, organization, and so on) using a deterministic vote across URL structure, schema.org types already present, and content-structure signals like link density and heading patterns — no LLM call, so it runs on every page instantly.

This classification is what lets the scoring system avoid two common mistakes: penalizing a homepage for not answering FAQ-style questions, and penalizing a documentation page for missing Product schema it was never going to have.

08 InterfacesPART 1

The audit engine is exposed through a Gradio app with three primary actions:

Actionapi_nameReturns
Analyze SEOanalyze_seo_asyncFormatted markdown SEO report
Analyze AI Visibilityanalyze_ai_asyncFormatted markdown AI Readiness report
Structured SEO JSONanalyze_seo_json_asyncThe full structured response object, including per-page checks and the CSV path — this is the programmatic entry point

All three share the same underlying async functions (run_seo_analysis_fastapi,run_full_seo_audit, run_ai_visibility_analysis), so nothing is fetched or rendered twice per run.

09 How a fix run worksPART 2

A fix run is triggered from a single form (or a direct POST) and executes as one sequential pipeline per repository, starting from the CSV the audit engine produced.

01
Clone
The target repo is cloned fresh into a temp working directory using a token-authenticated URL, on the requested BRANCH.
02
Detect stack
Inspects package.json and repo layout to identify Next.js (App or Pages Router), Jekyll, or static — this decides whether per-page fixes are possible.
03
Parse the audit CSV
Every row from Part 1's output becomes one page's current-state record and, for Next.js sites, is resolved to a real file path on disk.
04
Run the on-page agent
Generates and writes title, meta description, schema, OG/Twitter tags, and alt-text markers per page, then commits.
05
Run the technical agent
Writes site-level files (robots.txt, sitemaps) and per-page canonical / robots metadata, then commits.
06
Push
Each agent commits and pushes independently to origin/BRANCH, so a technical-only run never depends on on-page having run first.

🟢 On-page agent

autonomous

Owns everything that lives inside the page and its content.

In scope

  • Title tag & meta description
  • Heading structure notes (H1/H2/H3)
  • Keyword set, per page
  • Internal link suggestions (keyword overlap)
  • Alt-text generation for flagged images
  • WebPage / FAQPage / HowTo JSON-LD
  • Speakable snippet, Open Graph, Twitter tags

Explicitly out of scope

  • robots.txt, sitemap
  • Canonical tags, noindex
  • Redirects, hreflang

🔵 Technical agent

autonomous

Owns everything that makes the site crawlable, renderable, and trustworthy to an engine — the layer AEO/GEO signals depend on.

In scope

  • robots.txt
  • sitemap.xml and native Next.js sitemap route
  • Canonical tag & noindex handling
  • hreflang (when locales are supplied)
  • Redirect scaffold check
  • Diagnostic report for what can't be blind-fixed

Flagged, not fixed

  • Core Web Vitals (already measured in Part 1; needs a live re-check post-deploy)
  • Page speed
  • 404 / 5xx status codes
  • JS rendering, mobile rendering

🟠 Off-page

manual, planned

Off-page signals — citations, backlinks, brand mentions, third-party listings — are not something a repo edit can produce, and are not automated in this version.

In scope

    Out of scope for automation

    • Citations, backlinks, brand mentions, third-party listings
    • This is handled manually until an automated approach is deliberately designed.
    • The SEO score's reserved 20% off-page weight is meant to eventually reflect this.

    10 Fix Agents APIPART 2

    EndpointMethodPurpose
    /GETServes the operator form used to trigger a run.
    /healthGETLiveness check. Returns <code>{"status":"healthy"}</code>.
    /runPOSTExecutes a full optimization run synchronously and returns stdout/stderr from the orchestrator.

    POST /run — form fields

    FieldTypeDefault
    github_tokenstring—required
    repo_urlstring—required
    site_basestring—required
    csv_filefile (.csv)—required
    branchstringmainoptional
    git_usernamestringSEO-Auto-Fix-Botoptional
    git_emailstringseo-bot@example.comoptional
    agentsstringonpage,technicaloptional
    localesstring, comma-separated""optional

    The request blocks for the duration of the run (10-minute server-side timeout) and returns the orchestrator's return code, stdout, and stderr as JSON.

    // response shape
    {
      "returncode": 0,
      "stdout": "...",
      "stderr": "...",
      "success": true
    }

    11 CSV handoff contract

    This is the exact shape Part 1 exports and Part 2 expects — one row per URL. If you're feeding the fix agents a CSV from somewhere other than the audit engine, matching this schema is what makes it work.

    ColumnUsed by
    urlBoth — required, resolves the target file
    title, meta_descriptionOn-page agent
    h1_count, heading_orderOn-page agent
    missing_alt_tags, total_imagesOn-page agent
    top_keywords, seo_suggestionsOn-page agent
    canonical_tag, robots_metaTechnical agent
    viewport_presentTechnical agent (diagnostic report)
    schema_types, opengraph_tags, twitter_tagsOn-page agent
    word_count, readability_score, grammar_errors, text_to_html_ratio, seo_scoreOn-page agent (context for the model)

    The audit engine's export includes many additional columns (Core Web Vitals, AI Readiness sub-scores, redirect chains, and so on) for reporting purposes — the fix agents simply ignore any column they don't need.

    12 Environment variables

    ANTHROPIC_API_KEY
    Optional. Enables AI suggestions in the audit report and the LLM semantic-enhancement pass in the AI Readiness score. Both work without it.Part 1
    OPENAI_API_KEY
    Required by the fix agents. The FastAPI server refuses to start a run without it.Part 2
    GITHUB_TOKEN
    Required per run. Embedded into the clone URL, never logged.Part 2
    REPO_URL
    Required per run. HTTPS GitHub URL of the target repository.Part 2
    BRANCH
    Defaults to main.Part 2
    SITE_BASE
    Required. Canonical domain used for schema, sitemap, and OG/Twitter image URLs.Part 2
    CSV_PATH
    Set automatically by the server from the uploaded file.Part 2
    AGENTS
    Comma-separated: onpage, technical, or both.Part 2
    LOCALES
    Comma-separated locale codes, enables hreflang output.Part 2
    Note: the two parts currently call different LLM providers — Part 1 (audit) uses Anthropic, Part 2 (fix agents) uses OpenAI. Both run independently of that choice; worth knowing if you're managing API budgets across the two.

    13 Write safety & idempotencyPART 2

    Every file write goes through the same guard rails, regardless of which agent is writing:

    • Marker-scoped blocks. Generated content lives between explicit start/end markers (ONPAGE-AUTO-FIX, TECHNICAL-CANONICAL) so a second run replaces its own output instead of duplicating it.
    • Truncation guard. safe_write refuses to overwrite a file if the model's output is empty or suspiciously shorter than the original — the original file is kept and the failure is logged, not silently swallowed.
    • Single JSX root. Schema injection locates the component's actual return statement by scanning for the largest balanced-parenthesis span, rather than assuming the first or last return ( in the file — this avoids corrupting early guard clauses or helper functions.
    • Client component awareness. Files starting with "use client" are skipped for server-only metadata exports, since that syntax is invalid there.
    Note: schema injection is regex-based JSX parsing, not a full AST parse. It handles the common shapes reliably but is worth watching on repos with unusual JSX structure.

    14 Supported stacks

    StackAudit (Part 1)Per-page fixes (Part 2)Site-level fixes (Part 2)
    Any live siteFull — the crawler only needs a rendered URL——
    Next.js — App RouterFullFullFull, incl. native <code>sitemap.js</code>
    Next.js — Pages RouterFullFullFull
    JekyllFullNot wired uprobots.txt, sitemap.xml
    StaticFullNot wired uprobots.txt, sitemap.xml

    The audit engine (Part 1) works against any live, publicly reachable site regardless of stack — it never touches source code. The per-page fix agents (Part 2) need a repo it recognizes to know where to write.

    15 Roadmap

    next

    Shopify support

    On-page and technical signals mapped to theme/metafield equivalents rather than JSX — likely via the Admin API rather than a cloned Liquid repo.

    next

    WordPress support

    Same agent logic, applied through the REST API (and existing SEO plugin fields where present) instead of direct theme file edits.

    later

    Off-page automation

    Remains a manual process until there's a reliable, verifiable way to automate citation and mention building — and to responsibly populate the SEO score's off-page category.

    Internal documentation — Autonomous AI Visibility Optimizer (SEO / AEO / GEO). Generated for engineering reference.