Check "can AI get the content" first (crawlable, renderable, discoverable), then check "is AI willing to cite it" (structured facts, freshness signals, source credibility). The order is not interchangeable — if the first layer fails, perfect work on the second layer produces no result at all — that is why 3 of the 6 causes below sit in layer one.
Layer 1: AI cannot get your content (three technical causes)
Cause 1: robots.txt shuts AI crawlers out
What you see: site traffic looks normal, but server logs contain almost no visits from AI crawlers. Why: either a wildcard rule blocked everything at once, or a historical anti-scraping rule disabled crawler user agents and took GPTBot, ClaudeBot, PerplexityBot and Google-Extended down with it. Fix: allow the target AI crawlers explicitly, one by one, in robots.txt, and declare the sitemap location there. This is step one of GEO and also the step with the highest failure rate.
Cause 2: the body text is rendered by JavaScript, so the crawler gets a shell
What you see: the page carries full content in a browser, but when the same URL is requested the way a crawler requests it, the HTML contains only a framework and loading placeholders. Why: most AI crawlers do not execute JavaScript. If content is injected by client-side rendering, the crawler reads an empty shell. Fix: move to server-side rendering or pre-rendering so the body text is delivered as static HTML. A purely static site has this problem by construction — all 32 pages on this site ship hard-coded body text — which is one reason static sites tend to do better at GEO.
Cause 3: no sitemap and no llms.txt, so AI has to guess
What you see: the home page is crawled frequently while deep article pages receive almost no visits. Why: the site provides no content index, so engines do not know which pages exist or which ones matter. Fix: publish sitemap.xml listing every content page (this site lists 32 URLs), publish llms.txt giving an authoritative site summary, and declare the llms.txt location in the <head> of every page.
Layer 2: AI can get the content but will not cite it (three content causes)
Cause 4: opinions and positioning language, but no structured facts
What you see: the content is indexed, but when AI answers a related question it quotes somebody else. Why: engines extract facts, figures, specifications, dates and scope. If the whole page consists of unverifiable phrases such as "industry-leading" or "outstanding results", there is nothing to lift. Public research points the same way: verified and structured data accounts for 54.53% of independent citation sources. Fix: rewrite descriptive claims as verifiable factual statements — specific figures, explicit dates, clear scope — and add JSON-LD structured data (Organization / Service / Product / FAQPage).
Cause 5: no freshness signal
What you see: the content is not bad, yet engines that favour recent material cite it rarely or never. Why: with no datePublished or dateModified anywhere on the site, an engine cannot judge how old the material is and tends to treat it as unusable. Some engines are strongly time-sensitive — Perplexity in particular favours recent sources. Fix: emit publication and modification dates and establish a monthly refresh routine, refreshing at least the 4 article pages and the core service pages on a 30-day cycle. Content needs maintenance; publishing once is not the end of the job.
Cause 6: unverifiable sourcing and a missing operating entity
What you see: when AI mentions you, the information is wrong — or it never mentions you at all. Why: an engine has to judge whether a source is credible. If the page carries no legal company name, no scope of business and no checkable contact route, the entity profile is incomplete and the page is hard to treat as a trustworthy source. Fix: complete the entity profile — legal name, scope of business, contact channels and service commitments — and mark it up with Organization structured data. These 5 fact types are the foundation of a GEO architecture and are listed in full in this page's footer.
This table compresses the six causes into three columns — Symptom → Most likely cause → Fix first. (1) No AI crawler in the logs ⇒ robots rules block them ⇒ allow GPTBot, PerplexityBot and peers explicitly (this site whitelists 52 user agents in robots.txt); (2) the page looks fine in a browser but a crawler sees an empty shell ⇒ body text is rendered by JavaScript ⇒ move to server-side rendering or pre-rendering; (3) only the home page is crawled ⇒ the site has no content index ⇒ add sitemap.xml and llms.txt plus a head declaration; (4) indexed but never cited ⇒ there is no liftable fact ⇒ use specific figures and explicit scope, add structured data; (5) even new content is not cited ⇒ no freshness fields ⇒ emit datePublished / dateModified and refresh on a cycle; (6) AI states wrong facts about you ⇒ incomplete entity profile ⇒ complete the entity information and use Organization markup. Rows 1–3 belong to "AI cannot get it"; rows 4–6 to "AI will not cite it".
Self-check list: symptom → cause → fix
| Symptom | Most likely cause | Fix first |
|---|---|---|
| No AI crawler in the logs | Robots rules block them or the site is blocked wholesale | Allow GPTBot, PerplexityBot and peers explicitly |
| Page looks fine in a browser, crawler sees nothing | Body text is rendered by JavaScript | Move to server-side rendering or pre-rendering |
| Only the home page is crawled | No content index | Add sitemap.xml and llms.txt plus a head declaration |
| Indexed but never cited | No liftable facts | Use specific figures and explicit scope, add structured data |
| Even new content is not cited | No date fields | Emit datePublished / dateModified and refresh on a cycle |
| AI states wrong facts about you | Incomplete entity profile | Complete the entity information and use Organization markup |
| Finding | Published figure | Source |
|---|---|---|
| Structured, verifiable data is cited independently more often | Verified and structured data accounts for 54.53% of independent citation sources | Yext, 2026 17.2 million citations |
| Ranking on Google is not the same as being cited by AI | Only about 12% of pages cited by AI assistants also rank in Google's top 10 for that question | Ahrefs, 2025 15,000 long-tail questions |
| Citations come mostly from sources you neither own nor pay for | About 84%–94% of AI citations come from sources the brand does not own or pay for | Muck Rack Generative Pulse |
| A single test is noise, not signal | Variance across repeated runs of the same question on one day can reach 38%; answer drift is about 40.5% on Perplexity and 59.3% on Google AI Overviews | Industry reporting, 2026 |
The figures above are quoted from public reporting and research. Third-party data, not our own measurement, and not a prediction of results for any specific project. Samples and definitions differ between studies, so the numbers should not be read as fixed algorithmic ratios. Checkable sources: Writesonic / Ahrefs reporting (2026), Yext / Muck Rack reporting (2026), University of Toronto and related papers (2026), GEO measurement benchmarks (2026).
Frequently asked questions
robots.txt or a blanket block on the whole site. If crawlers do arrive but you are still not cited, the problem has moved to the content layer.FAQPage structured data; keep the two in sync, as this page does.This article describes method and observable process indicators. We do not guarantee citation by any AI engine, we do not write or publish encyclopedia entries on your behalf (we prepare a sourced evidence pack instead), and we do not disparage competitors. Capability and price statements follow the 13-section terms of service and the latest display inside the WeChat mini program.