Guide 8 of 9 · Draft Preview

Research Nobody Else Has

How to create the one kind of content AI search engines have no choice but to cite

Guide 8 of [series name TBD] — content draft, no design applied.

The short version

Guide 4 listed original research as one of eight asset classes, and flagged it as the strongest one. This guide is the full version of that: what actually qualifies as original research, three practical routes to producing it without needing a research department, a complete worked example, and the publication standards that decide whether it gets picked up or ignored.

The bar to clear isn't "new data." It's this: a question with real demand that nobody currently owns the answer to, answered with a method you can defend, and distributed so that third parties actually corroborate it. All three parts matter. A question nobody's asking isn't worth answering. A method you can't defend won't survive scrutiny. And research nobody distributes is just a page sitting on your site — not a publication, and not something an AI system has any independent reason to trust.


Three ways to get there

You don't need a survey team or a data science department. Almost all genuinely original research comes from one of three routes.

Route 1 — Primary observation

Generate data that doesn't currently exist: your own operational data, a survey you run or commission (with a stated sample and consent standard), freedom of information requests, or systematic collection from public records.

This is the highest-effort route, and the only one that makes you the primary source outright — nobody can dispute where the number came from, because you're the one who produced it. The discipline that matters here: the question comes first, the collection method second. Don't collect data and then go looking for a question to attach it to; that produces research that answers something nobody was asking.

Route 2 — Derivative narrowing

Take existing public data and cut it by a dimension nobody else has published — by region, by sector, by company size, by time of year. A national statistic becomes a regional index. A UK-wide rate becomes a sector-by-sector league table.

The insight here isn't the underlying data — that's public. It's the resolution: the specific cut exists only on your page. This is the fastest route to genuinely ownable research, and it's the right default for most sites without the resources for primary observation. One important check: the cut has to reveal something, not just subdivide the same number into smaller pieces. "Late payment by region" is a finding if regions genuinely differ in an interesting way; if they don't, you've made a table, not research.

Route 3 — Recombination

Join two or more public datasets that have never been put together before — insolvency rates against payment terms by sector, adviser density against regional demand, price against effective dose. The joined view itself is the original work, even though every individual input is public.

Recombination also underpins combination tools, where a user supplies one dataset themselves (their own numbers) and you join it against a public dataset you hold — the output is original to that specific combination even though neither input alone is.


Worked example — invoice finance

This is a full walk-through of the process, using a genuinely commodity topic to show the method still works where the subject looks saturated.

The starting problem. The topic given is invoice finance. Generic content on this topic is everywhere and earns nothing — every provider's site says roughly the same thing about how invoice finance works.

  1. Find the unowned question. A candidate: what day of the week do UK businesses actually pay invoices? Nobody publishes this. It's concrete, quotable, and it's exactly the kind of specific line a finance journalist writing about late payment would want to use.

  2. Choose the route. If payment-timing data is available operationally or through a survey, that's primary observation (Route 1). If not, derivative narrowing (Route 2) works too: take existing UK late-payment statistics and cut them to a region or sector, so the national story becomes "late payment in the East of England" or "which UK sector pays invoices fastest."

  3. Verify nobody owns it. Search the exact question, and its likely sub-questions, across Google and the AI platforms directly. If an answer already exists, narrow the cut further until it's genuinely unclaimed — and be able to state the specific gain in one sentence before moving on.

  4. Source and document as you go. Named datasets, retrieval dates, sample sizes, exclusions, the exact calculation used, licences on any inputs, and an evidence trail (collection logs, code, working notes) captured during the work — not reconstructed afterwards from memory. This step is what makes the work defensible later; skipping it is the single most common way research gets picked apart.

  5. Publish in a citation-ready format. A short capsule answer first, then supporting tables, then the method — with the source and date sitting next to every claim, not buried in a footer. The dataset itself ships as a downloadable CSV with a stated licence and version. A how-to-cite block on the page, for anyone who wants to reference it properly. The named, qualified author's page links to the piece.

  6. Distribute it. This is the step that's most commonly skipped, and it's the one that decides whether any of the above was worth doing. Prepare a short pack for journalists (the finding, a quote, a chart, ready to use), pitch it to relevant trade press, deposit the dataset somewhere researchers actually look, and send the specific cut to anyone it's directly relevant to (a sector body, a commentator who covers this). Research without distribution is a page. Third-party pickup is the actual point of doing this, because outside corroboration is what an AI system is checking for before it trusts a claim (see Guide 4's point on external validation).

  7. Build the refresh cycle. Turn the finding into an annual, year-tagged piece: same question, new period, each edition versioned with prior years archived and still citable in their own right. The real compounding effect starts at the second release — at that point you have a trend, which is a second story in its own right, and any citations the first edition earned carry forward into the credibility of the second.


Publication standards for research

These are the things that separate research an AI system (or a journalist) will trust from research that reads as marketing dressed up as data.

  • Keep official data and your own analysis clearly separate. Never blend them without labelling which is which. A projection is always labelled as a projection.
  • Every figure states the period it covers and the date it was retrieved or calculated, with the source sitting next to the claim itself.
  • Method transparency should be complete enough that someone else could reproduce your number. Publish code and collection logs where they exist.
  • The dataset ships alongside the page, as CSV or JSON, with a licence, a version, and a provenance note. Both the explanatory page and the dataset file are first-class — journalists and research-mode AI tools tend to cite the file directly; standard search-mode answers tend to cite the page.
  • If you want the work formally citable, mint a DOI for it, link it from the author's own page and external profiles, and actually run the distribution plan from step 6 above. All three together, or the piece isn't really finished — a DOI on its own doesn't make undistributed work authoritative.

Where to go next

This guide covers how to build the single highest-value asset in the framework. It's also, on its own, the hardest one to scale — which is exactly the problem the next guide addresses. If you're running more than one site or brand, the same research methodology, trust signals and distribution habits need to work across all of them without collapsing into the kind of templated, low-effort repetition that AI systems (and readers) can spot immediately. Guide 9 covers building authority across multiple brands without falling into that trap.


Guide 8 of 9 · internal draft preview, not for search engines.