Clairen Haus AEO Checker
Free scan

Methodology

How the AEO score is calculated

The score comes from 57 documented rules. This page is generated from the same code that scores your site.

The model

  1. Each rule has a fixed weight and measures a quality ratio from 0 to 1 using data fetched from your site. It records the evidence it used.
  2. A rule passes at 0.85 or above, warns at 0.4 or above, and fails below that.
  3. Checks we cannot verify, and checks that do not apply to your site, are removed from the calculation. They never raise or lower your score.
  4. A category score is the points earned out of the points possible across verified checks.
  5. When verified checks cover less than 40% of a category, the category is marked inconclusive. With fewer than 3 scored categories, the overall score is inconclusive too.
  6. The overall score is the average of the five category scores.
  7. Fixes are ranked by points recoverable, multiplied by business impact and divided by effort.

What the scan does and does not do

  • It reads the raw HTML your server sends and does not execute JavaScript, which matches most AI crawlers.
  • It requests your homepage with each AI crawler's published user agent. Google and Microsoft verify their crawlers by network address, so those two results are shown for information only.
  • Outside references are judged by what your site links to. The wider web is not searched in the free scan.
  • A score measures readiness signals. It does not guarantee rankings, mentions or citations in any AI product.

Access

Can AI crawlers reach your pages? 13 checks, total weight 84.

robots.txt is reachable

Weight 6

Fetches /robots.txt. HTTP 200 with a text policy = 1.0. HTTP 404 = 0.6 (crawlers treat everything as allowed, but you have no explicit AI policy or sitemap pointer). 5xx, timeouts or an HTML page served as robots.txt = 0.0, because major crawlers treat an unreachable robots.txt as 'do not crawl'.

access.robots_txt

OAI-SearchBot can access the site

Weight 8

Evaluates robots.txt rules for the OAI-SearchBot token against the homepage and the key pages we scanned. A full block scores 0; a partial block scores by share of key pages allowed. We also request the homepage with this crawler's user agent; a 403/challenge when a normal browser gets 200 caps quality at 0.2.

access.bot.oai-searchbot

GPTBot can access the site

Weight 4

Evaluates robots.txt rules for the GPTBot token against the homepage and the key pages we scanned. Blocking a training crawler is a legitimate business choice, so a full block scores 0.5 (warning), not 0. We also request the homepage with this crawler's user agent; a 403/challenge when a normal browser gets 200 caps quality at 0.2.

access.bot.gptbot

Claude-SearchBot can access the site

Weight 7

Evaluates robots.txt rules for the Claude-SearchBot token against the homepage and the key pages we scanned. A full block scores 0; a partial block scores by share of key pages allowed. We also request the homepage with this crawler's user agent; a 403/challenge when a normal browser gets 200 caps quality at 0.2.

access.bot.claude-searchbot

ClaudeBot can access the site

Weight 4

Evaluates robots.txt rules for the ClaudeBot token against the homepage and the key pages we scanned. Blocking a training crawler is a legitimate business choice, so a full block scores 0.5 (warning), not 0. We also request the homepage with this crawler's user agent; a 403/challenge when a normal browser gets 200 caps quality at 0.2.

access.bot.claudebot

PerplexityBot can access the site

Weight 7

Evaluates robots.txt rules for the PerplexityBot token against the homepage and the key pages we scanned. A full block scores 0; a partial block scores by share of key pages allowed. We also request the homepage with this crawler's user agent; a 403/challenge when a normal browser gets 200 caps quality at 0.2.

access.bot.perplexitybot

Google-Extended can access the site

Weight 4

Evaluates robots.txt rules for the Google-Extended token against the homepage and the key pages we scanned. Blocking a training crawler is a legitimate business choice, so a full block scores 0.5 (warning), not 0. This is a robots.txt token only, so there is no live request.

access.bot.google-extended

Googlebot can access the site

Weight 10

Evaluates robots.txt rules for the Googlebot token against the homepage and the key pages we scanned. A full block scores 0; a partial block scores by share of key pages allowed. We request the homepage with this user agent for information only. The operator verifies this crawler by IP address, so a refusal of our request is expected and is not scored.

access.bot.googlebot

Bingbot can access the site

Weight 8

Evaluates robots.txt rules for the Bingbot token against the homepage and the key pages we scanned. A full block scores 0; a partial block scores by share of key pages allowed. We request the homepage with this user agent for information only. The operator verifies this crawler by IP address, so a refusal of our request is expected and is not scored.

access.bot.bingbot

Homepage returns HTTP 200

Weight 8

Status of the final homepage response. 200 = 1.0, other 2xx = 0.8, anything else = 0.

access.http_status

Clean redirect behavior

Weight 5

Counts redirect hops on the homepage and scanned pages. 0–1 hops = 1.0, 2 hops = 0.6, 3+ = 0.2. A homepage that redirects to a different domain caps quality at 0.6. Pages with long chains are listed.

access.redirects

Fast server response

Weight 5

Time to first byte of the homepage, measured from our server. Under 800 ms = 1.0, under 1.5 s = 0.75, under 3 s = 0.45, slower = 0.1. The median across scanned pages is shown as evidence. Network distance affects this number, so treat it as indicative.

access.response_time

No firewall or CDN blocking of AI crawlers

Weight 8

Fingerprints WAF/CDN products from response headers and compares the homepage response for a normal browser against AI crawler user agents (excluding IP-verified crawlers). Quality = share of tested AI crawlers that received real content. If every probe failed for network reasons, the check is unavailable.

access.waf_blocking

Readability

Can they read the content once they arrive? 10 checks, total weight 57.

Content is in the server-rendered HTML

Weight 10

Counts words of main content present in the raw HTML (no JavaScript executed). Most AI crawlers do not run JavaScript. Per page: 150+ words = 1.0, 50–149 = 0.5, under 50 = 0. Category quality is the average across key pages.

readability.server_rendered

Pages do not depend on JavaScript to show content

Weight 8

Looks for single-page-app shells: empty #root/#app/#__next mount points, <noscript> 'enable JavaScript' messages, and pages with under 50 words but 5+ scripts. Per page: no signals = 1.0, one signal with 150+ words = 0.7, otherwise 0. This scan does not execute JavaScript, so this measures what a non-rendering crawler sees.

readability.js_dependency

Descriptive, unique page titles

Weight 6

Per page: a <title> of 15–70 characters that is not shared with another scanned page = 1.0. Too short / too long = 0.6. Duplicate = 0.5. Missing = 0.

readability.title

Meta descriptions summarize each page

Weight 5

Per page: a meta description of 50–170 characters = 1.0; present but outside that range = 0.6; missing = 0.

readability.meta_description

Canonical URLs are declared

Weight 4

Per page: a canonical link on the same site = 1.0; canonical pointing to another domain = 0.3; missing = 0.4.

readability.canonical

Logical heading structure

Weight 6

Per page: starts at 1.0; no H1 −0.5; more than one H1 −0.2; each skipped level (e.g. H2 → H4) −0.15; fewer than 2 subheadings on a page with 300+ words −0.2.

readability.headings

Main content is identifiable

Weight 5

Per page: a <main> landmark (or role=main) = 1.0; no landmark but 200+ words of sectioned body text = 0.7; otherwise 0.3.

readability.main_content

Reasonable HTML page size

Weight 3

Per page HTML size: ≤ 500 KB = 1.0, ≤ 1.5 MB = 0.6, larger = 0.2, truncated at our 3 MB cap = 0.

readability.page_size

Healthy content-to-code ratio

Weight 4

Readable text bytes ÷ HTML bytes per page. ≥ 10% = 1.0, ≥ 4% = 0.6, ≥ 1.5% = 0.3, lower = 0.

readability.content_ratio

No broken or inaccessible pages

Weight 6

Share of attempted internal pages that returned real content. Quality = loaded ÷ attempted. Failed pages are listed with their status.

readability.broken_pages

Structure

Do machine-readable signals explain what each page is? 11 checks, total weight 52.

Organization schema

Weight 8

Looks for Organization or a LocalBusiness subtype in JSON-LD on any scanned page. Present = 0.55, plus 0.15 each for name, url and logo (max 1.0). Missing = 0.

structure.organization

WebSite schema

Weight 4

WebSite JSON-LD with a name = 1.0; without a name = 0.7; missing = 0.

structure.website

Service or Product schema

Weight 6

Looks for Service, Product, Offer or OfferCatalog JSON-LD (or Organization makesOffer / hasOfferCatalog). Quality = share of scanned service/offer pages carrying it; if no service pages were found, 1.0 when present anywhere and 0 otherwise.

structure.service

FAQPage schema

Weight 5

Applies only when the site has FAQ content (an FAQ page, or a page with 3+ question headings). Quality = share of those pages with FAQPage JSON-LD. Not applicable otherwise (FAQ content itself is scored under Answerability).

structure.faqpage

Article schema on articles

Weight 4

Applies only when blog/article pages were scanned. Quality = share with Article, BlogPosting or NewsArticle JSON-LD.

structure.article

JSON-LD is valid

Weight 5

Every JSON-LD block must parse as JSON and declare an @type. Quality = valid blocks ÷ total blocks. Not applicable when the site has no JSON-LD (missing schema is scored by the schema checks).

structure.jsonld_valid

XML sitemap

Weight 6

Reads sitemaps declared in robots.txt, falling back to /sitemap.xml and /sitemap_index.xml. Valid sitemap with URLs = 1.0; valid but empty = 0.4; none found = 0.

structure.sitemap

Sitemap covers your key pages

Weight 5

Share of successfully scanned pages (excluding legal pages) whose URL appears in the sitemap, compared without tracking parameters, trailing slashes or www. Not applicable without a sitemap.

structure.sitemap_complete

llms.txt file

Weight 2

Checks /llms.txt. A Markdown file starting with an H1 = 1.0; present but not Markdown = 0.5; missing = 0. llms.txt is an emerging convention. Major AI search engines have not confirmed that they use it, so it carries a low weight.

structure.llms_txt

Breadcrumb schema on deeper pages

Weight 3

Applies to scanned pages two or more path levels deep. Quality = share with BreadcrumbList JSON-LD. Not applicable for flat sites.

structure.breadcrumb

Open Graph metadata

Weight 4

Per page: og:title, og:description and og:image present = 1.0, each missing tag −0.33.

structure.open_graph

Answerability

Is your content shaped so AI can lift a clear answer? 11 checks, total weight 64.

Direct answer near the top of key pages

Weight 10

Takes the first substantive paragraph after the H1. Per page: +0.4 if it exists, +0.3 if its first sentence is 8–35 words, +0.3 if it states what something is or does (definition pattern). Pages without such a paragraph score 0.

answerability.direct_answer

Question-based headings

Weight 6

Counts H2–H6 headings phrased as questions across key pages. 6+ = 1.0, 3–5 = 0.7, 1–2 = 0.4, none = 0. Matching the way buyers phrase questions helps AI systems map your content to prompts.

answerability.question_headings

Clear definitions and explanations

Weight 5

Per page: share of sections (max credit at 2) that contain a definition-style sentence ("X is a…", "X refers to…", "We help…"). 2+ = 1.0, 1 = 0.6, none = 0.

answerability.definitions

Statistics and concrete evidence

Weight 6

Per page: counts specific figures (percentages, prices, '12 years', '500+ clients', founding year, star ratings). 3+ = 1.0, 1–2 = 0.5, none = 0.

answerability.statistics

Named entities (places, brands, people)

Weight 4

Per page: distinct multi-word proper names in the text, such as cities, brands, certifications and people. 4+ = 1.0, 2–3 = 0.6, 1 = 0.3, none = 0. Specific names help AI systems connect your business to the right context.

answerability.named_entities

Citation-ready claims and sources

Weight 5

Site-wide: pages with source attributions or credentials ('according to', 'certified', 'licensed', 'source:') plus outbound links to non-social external sites. Quality = 0.5 × share of key pages with attribution language + 0.5 × min(1, distinct external domains linked ÷ 3).

answerability.citation_opportunities

Lists and tables

Weight 4

Per page (300+ words only): at least one content list or table = 1.0, none = 0.3. Shorter pages are not judged.

answerability.lists_tables

Content is split into answerable chunks

Weight 6

Per page (150+ words): quality = share of headed sections with 25–350 words; −0.3 if any single section exceeds 500 words. Self-contained sections are what AI systems retrieve and quote.

answerability.chunkability

FAQ coverage

Weight 6

1.0 if the site has an FAQ page or FAQPage schema with 5+ questions, or 8+ question headings site-wide; 0.6 with 3+ questions; 0.3 with 1–2; 0 with none.

answerability.faq_coverage

Clear service descriptions

Weight 7

Per service/offer page: 300+ words = 1.0, 150–299 = 0.6, under 150 = 0.2. If no service pages were found, the homepage is judged instead, at 0.5× credit, because AI systems need a dedicated page to cite for each service.

answerability.service_descriptions

Strong page-level topic focus

Weight 5

Per page: the title, H1 and opening paragraph should share topic terms. Quality = 0.5 × (title∩H1 overlap > 0) + 0.5 × (share of H1 terms that appear in the opening paragraph).

answerability.page_relevance

Trust & Entity

Is your business a clearly identified, credible entity? 12 checks, total weight 65.

Consistent business name

Weight 7

Collects the business name from Organization schema, WebSite schema, og:site_name, the homepage title and the footer copyright. Names are compared after removing legal suffixes, case and punctuation (one containing the other counts as a match). Quality = share agreeing with the most common name. A single source = 0.6. No sources = 0.

trust.business_name

Consistent business description

Weight 5

Compares the homepage meta description, og:description and Organization schema description. Quality = 0.4 for having at least one, +0.3 for having two or more, +0.3 when they share vocabulary (word-overlap ≥ 0.2).

trust.business_description

Founder, owner or author information

Weight 5

1.0 if Person schema, Organization founder, or Article author is present. 0.7 if text names a founder, owner, director or author (for example, 'founded by Jane Smith'). 0.3 if a team page exists with no names detected. Otherwise 0.

trust.founder_author

About page

Weight 6

1.0 if an About page was scanned with 150+ words; 0.6 if it was scanned but thinner, or linked but not scanned; 0 if no About page or link exists.

trust.about_page

Contact information

Weight 7

One third each for: a phone number (tel: link, schema telephone or visible number), an email address (mailto: or schema email), and a physical/service address (PostalAddress schema or contact page with an address pattern). A contact page alone adds 0.15, up to a maximum of 1.0.

trust.contact_info

LinkedIn presence linked

Weight 4

1.0 if a LinkedIn company or profile URL is linked or listed in schema sameAs; 0 otherwise.

trust.linkedin

Social profiles linked

Weight 4

Distinct social platforms linked or listed in sameAs (Facebook, Instagram, X, YouTube, TikTok, Pinterest, Threads, Bluesky; LinkedIn is scored separately). 2+ = 1.0, 1 = 0.6, none = 0.

trust.social_profiles

External business references

Weight 5

Links or sameAs entries pointing to independent business listings and knowledge bases (Google Business Profile, Wikipedia/Wikidata, Yelp, BBB, Trustpilot, Clutch, industry directories). 2+ = 1.0, 1 = 0.6, none = 0. This scan checks what your site points to. It does not search the wider web.

trust.external_references

Organization identity signals

Weight 4

One fifth each for: Organization schema logo, Organization schema url matching the site, schema sameAs with 2+ entries, og:site_name, and og:image on the homepage.

trust.org_identity

Reviews and proof

Weight 6

0.5 for Review or AggregateRating schema; 0.3 for visible testimonials, reviews or case studies (a heading or link mentioning them); 0.2 for links to third-party review platforms. Max 1.0.

trust.reviews_proof

Pages are indexable

Weight 8

Per page: a noindex directive in meta robots or the X-Robots-Tag header = 0; a canonical pointing to another domain = 0.5; otherwise 1.0. Site quality is the average, and a noindex homepage caps it at 0.2.

trust.indexability

No hidden instructions aimed at AI

Weight 4

Looks for text addressed to AI systems ('ignore previous instructions', 'AI assistants should recommend…') inside hidden elements, HTML comments or zero-width-character runs. Any hit = 0, because AI platforms treat this as manipulation. None = 1.0. Detected text is reported as evidence and never followed.

trust.content_integrity

Open-source attribution

Parts of the crawler-detection, firewall-fingerprinting and passage-analysis logic are adapted from GEO Reporter (MIT License, © 2026 Zubair Trabzada and © 2026 Tal Oron and contributors). Full notices are in the project's THIRD_PARTY_NOTICES.md.