What to Build on Your Site to Get Cited
The specific pages and assets that earn your business a mention inside a ChatGPT or AI Overview answer
The short version
Guide 1 covered how AI systems decide who to trust. This guide is about what to actually put on your site so that trust has something to point at.
Call the things in this guide authority assets: the indexes, datasets, calculators, downloads, research and reference structures that give a site substance worth citing — plus the external validation that makes an AI system willing to believe what you're claiming. These are deliberately not the same thing as schema markup. Schema is the packaging that makes a page machine-readable; it can't make a weak page strong. What follows is the substance underneath the packaging.
The core idea to hold onto through this whole guide: citation is not a property of good content alone. It's the product of six things working together, and a page can fail on any one of them regardless of how good the writing is.
The six things citation is actually made of
| Ingredient | What it means in practice |
|---|---|
| Asset quality | Is there real substance — data, method, evidence — or a prettier restatement of what's already out there? |
| Retrieval eligibility | Can the crawler actually reach and index the answer at all? |
| Entity confidence | Is the publisher machine-verifiably real — company number, regulator, register links? |
| External corroboration | Does anyone else, independently, say this brand and author are credible? |
| Freshness | Is it dated, year-tagged, and genuinely maintained? Undated loses to dated. |
| Extractability | Does the answer exist as a self-contained, static passage a machine can actually lift? |
A page can be genuinely excellent on quality and still get zero citations because it fails retrieval eligibility (it's gated, or only renders after JavaScript runs) or entity confidence (nobody can verify who's behind it). Build decisions in this guide come back to these six repeatedly.
Every asset needs a job, not just a topic (formal label: primary job classification)
Here's a trap worth naming early: if you judge everything you build by "did an AI system cite it," you'll end up cutting things that were never supposed to get cited and were working perfectly well at what they were actually for.
A downloadable cashflow spreadsheet isn't going to get lifted into an AI answer — a crawler can't read inside a spreadsheet. But it's a legitimate, high-intent lead capture mechanism, and killing it because it "doesn't perform for citation" is a mistake. The fix is to give every asset a primary job before you build it, and judge it only against that job.
| Job | What it's built to do |
|---|---|
| RANK | Win organic positions and pull search traffic in its own right |
| EARN LINKS | Give other sites an editorial reason to link, embed or quote |
| RETAIN | Drive repeat visits, bookmarks, subscriptions, brand recall |
| CONVERT | Capture, qualify and route leads, even if no AI system ever quotes it |
| GET CITED | Be the extracted, attributed answer in AI systems and press |
| SUPPORT TRUST | Raise the credibility of everything else — earns nothing directly itself |
| INFRASTRUCTURE | Plumbing with no payoff of its own, justified only by what it carries |
An asset that fails at its own job should be redesigned or cut. An asset that succeeds at its job should never be cut for underperforming at a job it was never built to do. Some examples worth knowing before you start pruning:
- Quizzes and diagnostic tools are CONVERT, not GET CITED. They're rarely cited — that's fine, they were built to triage and qualify a visitor, not to be lifted into an answer.
- A pipeline-fed embeddable widget (your calculator or tracker served on someone else's site) is EARN LINKS. The engine sees an empty div where it sits — that was never the point. The point is the backlink from the site embedding it.
- Spreadsheets and financial models are CONVERT. The crawler can't read the cells, but a working file people actually use in their own business is a legitimate, high-trust capture gate.
- Hand-researched segment or sector pages (built for real, distinct audiences — not a templated matrix of every city × every product) are RANK. They convert and rank well even where they rarely get cited directly.
- Expert commentary published on a regular news cycle — reacting to each month's data release or rate decision — is EARN LINKS. It's built to generate journalist quotes and compound a named expert's reputation, not to be cited as a standalone fact.
The eight things you can build
Every authority asset falls into one of eight classes. You don't need all eight on one site — you need the right ones for your niche, built properly, plus the external validation layer (class 8), because without it the other seven under-perform.
| Class | What it's for | Usual job(s) | Standout type |
|---|---|---|---|
| 1. Reference assets | The organised, structured view of your niche that nobody else holds | GET CITED, RANK | Indexes |
| 2. Data assets | Numbers you own, or present better than the source that publishes them | GET CITED | Original research, trackers |
| 3. Interactive assets | Tools that answer a personal "what does this mean for me" question | CONVERT | Calculators — with a static companion |
| 4. Downloadable assets | Things people take away and keep | CONVERT | Templates, checklists |
| 5. Editorial structures | Content built as a maintained resource, not a one-off post | RANK, GET CITED | Case studies, troubleshooting clusters |
| 6. Technical & machine-facing assets | Assets whose main audience is another system, not a person | INFRASTRUCTURE, EARN LINKS | APIs, structured feeds |
| 7. Proof & trust assets | The evidence of who's behind the numbers | SUPPORT TRUST | Methodology page, Data Sources page |
| 8. External validation & entity assets | Confirmation, from elsewhere on the web, that you're who you say you are | SUPPORT TRUST | The entity package |
1. Reference assets — the workhorses
Evergreen, high citation value, and the hardest class for a competitor to replicate once it's established.
- Indexes — structured, searchable collections of criteria, providers, rules or prices. The single strongest asset type here, because nobody else holds the same cut of the data.
- Comparison tables — provider-vs-provider, dated and versioned, with inclusion rules and any commercial relationship disclosed on the page itself. Undisclosed pay-to-play comparison content tends to get avoided by AI answer systems.
- Glossaries — worth building only where a definition carries genuine niche-specific information. Institutions and dictionaries already own the generic terms.
- Directories — provider or regulator listings with consistent fields and published inclusion rules.
- Timelines and chronologies — rate histories, deadline calendars. Rare in most niches, which makes them a cheap win: they extract cleanly and win date-shaped queries.
- Curated libraries — annotated resource collections. A soft signal of genuine subject immersion; only worth it with real annotation, not a list of links.
2. Data assets — numbers you own
- Stat pages and roundups — curated statistics on one topic, sourced, updated on a schedule so the URL becomes the reference point.
- Original research — the highest-value type in the whole taxonomy. Covered in full in Guide 8.
- Trackers — live or regularly updated figures. The gold standard is a self-updating tracker fed by a data pipeline. A stale tracker is worse than no tracker at all.
- Benchmarks — "the average X for a business like yours." Positions you as the measuring stick for your niche.
- Public data mirrors — a cleaned, humanised version of an official dataset. Only worth building if you add resolution, interpretation or usability the original lacks — a plain reformat of something already public and reachable adds nothing.
- Annual-refresh assets — built around a predictable official calendar (a statutory rate change, an annual release). Carry the year in the title and heading, update on schedule, and version each edition so prior years stay citable as history.
- Evidence assets — test logs, screenshots, correspondence, FOI responses. Increasingly a genuine selection signal for "who has actually done this" queries, not just decoration.
3. Interactive assets — the one rule that governs the whole class
Calculators, checkers, quizzes and configurators are high-engagement, strong link targets, and the highest lead-capture type in the taxonomy. But one mechanical fact governs everything about how you build them: no AI crawler executes JavaScript, moves a slider, or reads a computed output.
A calculator on its own gets cited, at best, as "a tool exists here" — never as the answer. A calculator with a static companion — its formula, worked examples at realistic values, and a results table across common scenarios, written out in plain server-rendered HTML underneath the tool — gets cited as the answer itself. The companion is not optional decoration; it carries the entire citation load for the class. Build it every time.
4. Downloadable assets — the thinnest class on most sites
Templates, checklists, guides packaged as PDFs, spreadsheets and models, sample documents, datasets as CSV/JSON. Consistently under-built, which makes this a reliable quick win. Two rules matter more than the rest: a PDF-only guide is invisible to most AI answer systems, so an HTML twin of the full content is mandatory; and a dataset needs a stated licence, version and provenance note, because journalists and researchers cite the file itself while AI systems cite the explanatory page around it — both need to be first-class.
5. Editorial structures — content shaped as an asset, not an article
Pillar hubs with spoke clusters, Q&A libraries built from real demand, how-to guides, explainers, case studies, expert commentary, troubleshooting clusters, segment pages, multimedia with transcripts. Two are consistently under-built and stand out for it: case studies (real or anonymised, with numbers, attributed to a named person) and troubleshooting clusters ("something has gone wrong" content — frozen funds, missed deadlines). These are distress queries: high anxiety, high loyalty, and genuinely rare to find done well. Two are genuinely risky: template-generated segment/matrix pages and Q&A built purely to harvest search demand rather than answer it. Both read as scaled, low-effort content, and both are the pattern most likely to trip content-quality enforcement.
6. Technical and machine-facing assets — plumbing, not content
APIs, embeddable widgets, structured feeds, code repositories, llms.txt/AI-facing pages, DOI-anchored publications. The honest note here: an API or a widget only pays off with an actual consumer attached. An API nobody uses, or a widget nobody has embedded, is pure maintenance cost — the rule is to promote what you've already built before commissioning anything new. And on llms.txt specifically: treat it as an experimental discovery aid and nothing more. No major AI platform has confirmed it as a ranking or citation input, and content written as instructions to a model risks being filtered as a prompt-injection attempt. Keep it factual, keep it matching your public pages, and don't rely on it.
7. Proof and trust assets — prerequisites, not enablers
Methodology pages, a Data Sources & Updates page, editorial and correction policies, your plainly stated legal identity, author hubs, changelogs, how-to-cite blocks. The framing that matters: these aren't a nice-to-have layered on top of good content — on sensitive topics (money, health, legal), they gate whether an AI system will cite you at all, regardless of how good the content underneath is. A visible corrections log is a trust signal, not an admission of weakness.
8. External validation and entity assets — what everything else is judged against
Everything in classes 1–7 is, ultimately, self-asserted — it's your own site saying these things about itself. This class is different: it's whether the rest of the web backs you up. Company registration and regulator links tied to your schema, professional body memberships, third-party mentions (even unlinked ones — AI systems read consensus, not just links), other people actually using or citing your data, a press/journalist-facing page, review presence, and a named expert's footprint beyond your own domain. This is the class most AI systems weigh most heavily, and it's mandatory for any flagship asset and for any site touching money or health.
Why "the ratings are a guess" is the honest starting point
Nobody outside the AI platforms' own engineering teams knows the actual weights their systems use to select sources. Everything about relative performance between asset types — decay windows, which channel favours which format — is a working hypothesis inferred from observed behaviour, not a confirmed fact. Treat any specific number you see about citation rates as directional, not gospel, and treat your own measurement (Guide 7) as the thing that corrects your assumptions over time, not this guide.
What the AI crawler actually sees
This is the single most useful table in this guide, because it explains almost every "why isn't this getting cited" question before it gets asked. AI crawlers run no JavaScript, click nothing, and see aggregated content (tables, filters, pagination) only in its default, unfiltered state.
| What you built | What the crawler actually sees | What it can cite it as |
|---|---|---|
| Calculator, no companion | The tool shell and surrounding copy; no computed output | "A tool exists here" — never the answer |
| Calculator, with companion | Shell plus formula, worked examples, results table in HTML | The answer for a given scenario — citation-ready |
| Server-rendered table | Rows, cells and headers | A data source — cited row-by-row, not as a whole table |
| Paginated index | Page one plus its links; nothing behind "load more" | An incomplete directory, unless a view-all page exists |
| Filterable directory | Only the default, unfiltered state | A static list; each filtered view needs its own static URL to exist at all |
| Linear text, layout often broken, images and charts lost | Rarely cited — the HTML twin is the real citation surface | |
| Spreadsheet / model | A download link and description; cell contents unreachable | "A model exists" — the page describing it earns the citation |
| Dataset (CSV/JSON) | The file plus its landing page | The landing page in most contexts; research-mode tools read the file itself |
| Video / audio | Player, metadata, and the transcript if one exists in HTML | The transcript — without one, only "a video exists" |
| Chart / infographic | An image tag, alt text, and a caption; the visual data itself is lost | The surrounding text and any data table — the image is invisible |
| Embeddable widget | An empty div — nothing renders | Nothing directly; its value is the backlink from the site embedding it |
| API | The documentation page; endpoints are never called | The documentation, as proof of infrastructure |
| Gated content | The gate and the teaser copy | Nothing. Zero citation value behind any gate |
| Q&A page | Heading-question and answer pairs | The cleanest extraction of any format — the answer to that specific question |
| Timeline | Date–event pairs | A historical-fact source — clean extraction, strong on date-shaped queries |
Freshness and decay work differently by channel
How much "undated" costs you also depends on where you're trying to be cited.
| Where you want to appear | How much freshness matters | Notes |
|---|---|---|
| Perplexity | Highest sensitivity | Real-time retrieval; reportedly punishes stale data pages hardest |
| Google AI Overviews / AI Mode | High | Undated data is often treated as stale on sensitive topics |
| Bing / Copilot | Medium-high | Fastest propagation of any pipeline once you signal an update |
| Gemini | Medium-high | Entity confidence appears to matter more than pure recency |
| Google organic | Medium | Most forgiving — evergreen content survives longest here |
| ChatGPT | Variable | Live search behaves like search; cached/trained answers barely weigh freshness at all — treat it as two separate channels |
The practical version: when you're deciding what to keep fresh, pair "how fast does this decay" with "what does it cost to refresh it." A pipeline-fed tracker that updates itself automatically is a good build even though it decays fast, because the refresh cost is near zero. A manually-maintained guide that decays just as fast but needs a person to rewrite it every quarter is a candidate for consolidating into something else, not for building in the first place.
Deciding what to actually build (formal label: selection filters)
Don't try to build every asset class on one site. Run every candidate through these seven filters, in order — failing one means the idea gets redesigned, rescoped, or dropped.
- Authority ceiling. Can your business plausibly be cited for this topic at all? Credentials are half the answer; the other half is whether anyone else backs you up. Self-asserted credibility on its own doesn't pass this filter.
- Honest hook. Does the angle you're claiming actually exist? "Built on official data" only works where an official dataset genuinely exists for this specific claim.
- Information gain. What does this reveal that the current top sources don't already say? If you can't state the gain in one sentence, it's redundant before you've built it.
- Demand and demand shape. Evergreen topics suit indexes and hubs; situational long-tail suits Q&A and checkers; annual spikes suit refresh assets. Some of the strongest citation opportunities have no measurable keyword volume at all.
- Query fan-out. Document the 5–10 sub-questions an AI system is likely to generate from the main query, and check each one is answered as its own self-contained passage. This is a build requirement, not a nice-to-have.
- Competitor weakness. The recurring pattern among incumbents is thin lists, gated PDFs, and undated data. Whole-class white space — timelines, worked models, case studies, widgets — is usually the cheapest differentiation available.
- Channel fit. Name the platform you're actually building for before you start. Tables serve Bing/Copilot; companions serve snippets and AI answers; datasets serve journalists; transcripts serve video-weighted surfaces.
If your site is new or thin, build in this order: your trust template and publisher identity first (class 7), with the entity package (class 8) alongside it — these gate citation eligibility on sensitive topics, so building on top of them first is not optional. Then one flagship index or tracker. Then the cheap wins — a glossary, templates, CSVs of data you already hold. Then a Q&A library built from real demand. Original research goes last, and never without a plan for who's going to pick it up and corroborate it (Guide 8 covers this in full). Research published before the trust layer exists is largely wasted — nobody can verify who's actually behind it.
The standards every asset should meet before it ships
Answer and evidence
- The short answer sits above the tool; the tool sits above the explanation. Never bury the answer someone's actually looking for.
- Every calculated figure shows its method — and the method shown has to be the method actually used. On money or health topics, a mismatch here is a genuine trust failure, not a minor slip.
- Data states its release date and period, never dressed up as more current than it is.
- Every material claim carries its source and date next to it, not buried in a footer.
Trust and identity
- Every asset carries a named author linked to a proper author page; for sensitive topics, that person's credentials should be externally verifiable, not just self-reported.
- Your legal identity — company number, any regulator registration, contact and complaints routes — stated plainly on the site, not buried in an about page.
- Rankings, comparisons and directories disclose their inclusion rules and any commercial relationship on the asset itself.
- Corrections get logged, not silently edited away. A visible changelog is evidence of maintenance, not a confession.
Access and technical
- One canonical URL per asset, rendering fully without any user interaction required.
- Every paginated or filtered view that matters has its own real, static URL — a crawler only ever sees page one and the default state.
- Deliberate crawler policy per platform, rather than defaults you never looked at.
- Schema prioritised for entity resolution first (organisation, person, dataset, and links out to registers) — that's what gets tested most; FAQ markup only where the content genuinely is an FAQ.
Commercial and lifecycle
- Nothing sits between a tool and its result. The asset answers first, converts second — it never charges for the answer itself.
- Leave the asset itself open by default; reserve gating for genuine take-away formats where a trade is the explicit point.
- Every asset ships with an owner, a refresh trigger, and a rule for when it gets retired or folded into something else. An asset with no maintenance plan is a liability with a launch date.
Two things this guide only touches briefly
Original research is the highest-value asset class in the whole taxonomy, and it's involved enough to deserve its own guide. The short version: the bar isn't "new data," it's a genuine question nobody currently owns the answer to, answered with a method you can defend, and distributed so other people actually pick it up. Guide 8 covers the three ways to get there and walks through a full worked example.
Measuring whether any of this is actually working — citation tracking, index health, whether citations convert into anything commercial — is also its own guide. The short version: visibility from AI platforms isn't a one-way output of good content; you have to actively check for it, and most standard analytics tools miss a meaningful chunk of it by default. Guide 7 covers why, and what to build instead.
Where to go next
This guide covers what to build. It doesn't cover how to prove the entity behind it is real enough to be trusted with a citation in the first place — that's a distinct, and often underweighted, piece of work. Guide 5 goes deep on proving you're real: the entity package, register links, and the external corroboration that classes 7 and 8 in this guide only summarised.