Guide

How to Do Demand Research Before You Have Any Search Data

The short answer

Where do you get questions to answer when your site has no search data?

On day 0 a new site has no Search Console data to mine, so the questions come from research you run by hand: what people already ask you, parallel single-lens research documents, communities where buyers describe the problem, and platform demand tools used for topics only. Our auto-finance rebuild categorized 280 questions that way on its first day, and Search Console took over by day 6 [our data].

Every guide to query mining assumes a Search Console property with rows in it. On day 0 there are no rows: the report only describes queries your pages already earn impressions for, so a domain with nothing published has nothing to mine. The questions have to come from somewhere else for the first week or two.

This page is the cold-start method our three production builds actually ran, and the dated point at which each one stopped needing it [our data].

Why can't you mine Search Console on day 0?

Because Search Console reports demand your site already touches, and a new site touches none. That is a property of the instrument rather than a delay you can shorten with effort: no impressions means no queries, and no queries means no export.

The corollary is the useful half. The report becomes readable as soon as pages are live and collecting impressions, which happens in days rather than months if you publish something substantial. So the honest framing is not "research versus mining" — it is a short hand-built bridge to a permanent instrument. The mining method that takes over is how to mine Search Console for AI-shaped queries, and this page is only about what happens before it can run.

Where do the first questions actually come from?

From four sources that exist before traffic does, each with a different failure mode:

SourceWhat it gives youWhat it cannot give you
Questions you are already askedReal buyer phrasing, free, immediatelyVolume, or any sense of relative demand
Structured research documentsCoverage of a whole vertical in a fixed orderEvidence that anyone searches for it
Communities where buyers describe the problemProblem phrasing in the words people typeA representative sample of anything
Platform topic-suggestion toolsGenuine demand signal inside that platform's nicheTrustworthy wording, and any cross-platform read

None of them is a substitute for query data, and saying so is the point. Together they produce a ranked list of plausible questions; Search Console later tells you which ones were real.

What does parallel single-lens research look like?

One document per lens, written at the same time, then synthesized into one plan. On our auto-finance rebuild that meant 10 research documents — roughly 3,000 lines across demand, buyer journey, competitors, engine-by-engine citation mechanics, compliance, architecture, E-E-A-T and conversion — synthesized into a single playbook on the same day the first code was committed [our data]. The demand lens alone produced 280 categorized questions.

Splitting by lens is what makes the volume possible and the output usable. A single "research the market" document goes wide and stops; a lens with one question to answer goes deep, and the categories are already separated when the synthesis pass starts. What the synthesis feeds — clusters, URL assignment, build order — is a different decision entirely, worked through in how to structure a content library.

One lens deserves separate handling. Our auto-finance research logged 20 questions where reputable sources contradict each other, and kept them as their own list rather than folding them into the general inventory [our data]. A reconciling page has something to offer that the existing answers do not, and the method for writing one is the answer-variance method.

How long does cold-start research have to last?

Days, on the record we have. Each of our builds reached mineable Search Console data quickly enough that the hand-built inventory was a bridge rather than a foundation [our data]:

BuildCold-start demand inputDay it landedFirst mineable GSC data
Greenfield insuranceA community demand-research briefDay 1Day 4
Auto-finance migration10 single-lens research documents, 280 categorized questionsDay 0Day 6
This libraryA 92-question inventory in its research documentBefore launchNot recorded

Day numbers count from each repo's first commit; dates come from git history [our data]. The migration build's day 6 is the weaker of the two GSC figures for comparison purposes, because that domain inherited years of search history — the greenfield row is the clean read on how fast a new domain produces signal.

Two limits on reading this table. It describes builds that published dozens of pages in the first days, so the query surface existed for impressions to land on; a site that ships five pages will wait longer. And it says nothing about outcomes — these are dates on which a report became readable, not evidence that any page ranked, got cited, or earned traffic.

Which cold-start sources can you actually trust?

Trust them for topics and phrasing; verify every claim independently. The distinction matters most with generated suggestions: a video platform's topic-suggestion tab gave us real, validated demand for its niche, and the AI-written copy attached to those topics fabricated a coverage claim that contradicted the controlling source [our data]. We take the topics and write the copy from primary sources, and the incident is documented in when AI-suggested copy fabricated a claim.

The same split applies to community reading. A forum thread is excellent evidence of how a problem is phrased and poor evidence of how common it is. Collect the wording; get the facts elsewhere. Google's people-first guidance is the standard the resulting pages have to meet either way — content written for people, demonstrating real knowledge, rather than content assembled to cover a keyword list (Google, creating helpful content).

Does cold-start research still matter once the data arrives?

Less than you would expect, and that is worth planning for. Once query data exists it is a better input than any inventory built from guesses, because it reports what people typed rather than what a researcher predicted. The cold-start inventory's remaining job is coverage of questions your pages have not yet earned impressions for — the demand you have not touched, which the report is structurally blind to.

There is one durable argument for keeping the research documents. Kevin Indig's 2026 state-of-AI-search analysis reports that content published within the last three months is cited about three times as often as older content (Growth Memo, 2026) — a correlation from one analyst's dataset, not a causal finding, and not something we have reproduced. If recency has any bearing on retrieval, a standing inventory of unanswered questions is what lets you publish deliberately rather than reactively.

When is this too much research?

When the site will have twenty pages. Against our own interest, since research is a phase we bill for: a small site does not need ten single-lens documents or a categorized inventory. It needs the ten questions its owner already answers on the phone every week, written as ten pages, and a Search Console property to read in a fortnight.

The scale where structured demand research earns its cost is the scale where nobody can hold the question list in their head and several people are writing at once. The rest of the system that inventory feeds — contract, gates, publishing, measurement — is described in our generative engine optimization guide.

Frequently asked questions

How do you do keyword research for a brand-new site?

By hand, from sources that exist before traffic does: the questions you already get asked, structured research documents, communities where buyers describe the problem, and platform demand tools. Our auto-finance rebuild produced 280 categorized questions this way on day 0 [our data].

How long before Search Console has data worth mining?

Sooner than most plans assume, if pages are live. Search Console returned mineable query data on day 4 of our greenfield build and on day 6 of our migration build, which inherited domain history [our data]. One build's record is not a promise about yours.

Are paid keyword tools worth it before launch?

They estimate demand you have not touched yet, which is exactly what a new site lacks, so they answer a real question. They also cost money and describe a vendor's panel rather than your site. Our own cold-start research ran on structured documents, community reading and platform demand tools instead [our data].

Can you use an AI tool to generate the question list?

Use it for topics, never for wording. A video platform's topic-suggestion feature gave us genuine demand signal while its AI-written copy fabricated a coverage claim that the source material contradicted [our data]. Topic input is cheap to verify; generated claims are not.

What do you do with questions where sources contradict each other?

Record them in their own list. Contradictions are the most valuable rows in a cold-start inventory because a page that reconciles them has something the existing answers do not. Our auto-finance research logged 20 such questions before any page was written [our data].

Sources

  1. Creating helpful, reliable, people-first contentGoogle
  2. State of AI Search Optimization 2026Growth Memo