# Prompt Research vs Keyword Research

> Source: https://rankxai.com/blog/prompt-research · Last updated: 2026-09-08

Prompt research is the practice of finding the questions people put to AI assistants, then tracking how those answers name your brand. Keyword research measures typed demand against published volume figures; prompt research has no volume data at all, so it trades forecasting for a tracked set you run repeatedly and read as a rate.

## What is prompt research, and how is it different from keyword research?

Prompt research finds the questions people put to AI assistants, tracks how the assistants answer, and records whether the answer names your brand. Keyword research does the equivalent job for search engines and has anchored SEO planning for two decades. One difference changes everything downstream: keyword research is built on published keyword search volume, and no assistant publishes anything equivalent.

That difference is not a gap waiting to be filled. Google reports impressions for its own search queries in Search Console and volume ranges in Keyword Planner. The AI systems report nothing at all to the sites they cite: OpenAI, Anthropic and Perplexity publish no aggregate query data of any kind. Prompt research therefore produces a tracked set you run repeatedly, rather than a table you sort by volume.

| Dimension | Keyword research | Prompt research |
| --- | --- | --- |
| What it measures | How often a phrase is typed into a search engine | How an assistant answers a question, and whether you are named |
| Volume data | Published by Google and modelled by third parties | None published by any assistant |
| Typical length | A short phrase, three to four words | A conversational question, median 8 to 12 words |
| The result | A rank, stable enough to chart daily | A mention or no mention, which differs between identical runs |
| What you read | A results page anyone can inspect | Generated text that changes every time you ask |
| The reportable number | Position, and clicks from Search Console | A rate across a fixed set of prompts and runs |

The last row is where most prompt research goes wrong. A rank is a property of your page. A mention rate is a property of your page **and** of the questions somebody chose to ask, and the second half is usually invisible in the report.

## Does prompt research replace keyword research?

Prompt research does not replace keyword research, and treating it as a successor costs you the surface with the most traffic on it. Google AI Overviews and AI Mode are both assembled from Google's own index, so a page that cannot reach the search results is rarely a candidate for citation in either. Traditional SEO still governs that half, and traditional keyword research is how it gets planned.

What prompt research adds is the half no keyword tool reports. A keyword tool can tell you that 2,400 people a month search a phrase. It cannot tell you what the AI answer says when people ask AI the same thing, or that ChatGPT and Perplexity name four competitors and not you. That is a fact about your market you can only get by asking.

The practical division is that keyword research decides what a page targets and prompt research decides what an answer needs to contain. The mechanics of how one question becomes several searches behind the scenes belong to [query fan-out](/blog/query-fan-out), and the [platform-by-platform comparison](/blog/how-to-rank-in-ai-search) sits with the pillar.

## How long are the prompts people actually type?

Real prompts are short, and the figure most often repeated about them is not. The number in wide circulation is that the average ChatGPT prompt runs about 60 words, published on Similarweb's GEO blog in May 2026 and drawn from its 2025 Generative AI Landscape report. Every clickstream and panel measurement of prompt length lands between about 4 and 12 words.

Semrush clickstream data puts the average ChatGPT prompt at [4.2 to 8.7 words](https://searchengineland.com/real-people-actually-prompt-ai-geo-479731). Otterly.AI's analysis found real prompts run 71 per cent longer than the synthetic ones marketers write for them, and still land at a median of 12 words. Stella Rising surveyed 524 active assistant users through the Centiment panel in January 2026, at a stated margin of error of 4.3 per cent, and found two thirds writing 15 words or fewer.

Similarweb publishes no method for the 60-word figure, describing the report's numbers as estimates produced by proprietary algorithms, so a reader cannot tell whether it counts one prompt, a whole conversation or text pasted into the box. Read it as a disagreement between one modelled figure with an unpublished method and three measured ones, rather than as a debunk. For building a prompt set the practical consequence is the same either way.

Shape matters as much as length. In the same January 2026 survey 60 per cent of prompts were phrased as questions and 9 per cent as direct commands, 32 per cent carried a personal attribute, 28 per cent mentioned price or budget and 16 per cent named a location. Track conversational questions carrying one of those details, since that is the input the AI systems are actually given.

### Long prompts measure less than they look like they do

Tracking elaborate buyer scenarios feels more realistic and tests less. Go Fish Digital measured how much of a prompt survives into the searches an assistant actually runs. Prompts under 12 words were shortened by 11.2 per cent, prompts of 12 to 17 words by 27.1 per cent, and prompts of 18 words or more by 56.1 per cent.

More than half of a long prompt is discarded before retrieval happens, which means a 40-word scenario is partly testing text no [synthetic query](/glossary/synthetic-query) ever carries. Prompt shape also moves fast: the share of prompts classified as keyword-shaped in the same survey programme fell from about 50 per cent in August 2025 to about 30 per cent in January 2026.

## Where do you get real prompts to track?

Five sources produce prompts worth tracking, and only one of them costs money. Ranked by how well the questions predict revenue rather than by how many you get:

1. **Your own sales and support inbox.** The questions people already ask a human before buying, in their words. Highest value, smallest volume, and you already have it.
2. **Google Search Console.** Filter your search queries to question shapes and comparisons. This is typed demand rather than prompt demand, but it is your market's vocabulary, measured.
3. **Community threads.** Reddit, Quora and industry forums carry the specific prompts people write when nobody is selling to them, in the words they would use with an assistant.
4. **Public prompt corpora.** [OpenAI Signals](https://openai.com/signals/data/) publishes what people bring to ChatGPT, and WildChat is a freely licensed corpus of about a million real conversations.
5. **A prompt research tool.** What a keyword research tool is to search, pointed at the AI tools instead. Buys you scale and repeated measurement, which is the part you cannot assemble by hand.

The public corpora carry a limit worth naming. OpenAI Signals covers Free, Go, Plus and Pro accounts only, so it excludes enterprise and Codex usage and underrepresents business and technical questions by OpenAI's own description. WildChat holds about a million opt-in conversations and 2.5 million turns under an ODC-BY licence, recorded on gpt-3.5-turbo and gpt-4 and timestamped from 2023, so it is a record of how people prompted three years ago rather than now.

### There is no prompt volume, and the vendors say so

No tool can sell you the volume of an individual prompt. Semrush, which sells prompt research, states in its own documentation that [individual prompts are often too specific and unique to measure directly](https://www.semrush.com/kb/1607-semrush-ai-visibility-data). It reports volume for topics that group semantically related prompts instead, estimated from 317 million prompts of clickstream data and machine learning models.

Its blog puts the same point plainly: there are no years of historical search volume, cost per click or trend data for AI prompts, so prompt research works on direction and pattern rather than precise counts. Use a volume estimate to order a shortlist and to drop questions nobody asks, never to build a forecast. The [prompt volume glossary entry](/glossary/prompt-volume) sets out which proxies produce which numbers.

## How much does your choice of prompts change the number?

The prompt set changes the answer by more than most site changes do, which makes it the first thing to check when an AI visibility score moves. Go Fish Digital ran the cleanest published test of this: 1,554 responses across 73 prompt variations on ChatGPT, Google AI Overviews and Perplexity, seven runs each, between 20 and 26 July 2026.

Phrasing the same need three ways moved the brand's mention rate by a factor of 6.6. Prompts written in category language scored 23.6 per cent. The same need described as a problem, without industry labels, scored 7.6 per cent. Full buyer situations scored 3.6 per cent. On Perplexity the situational prompts scored zero.

| Mention rate (Go Fish Digital, 20 to 26 July 2026) | Category-led prompts | Full buyer situations |
| --- | --- | --- |
| ChatGPT | 31.6% | 3.6% |
| Google AI Overviews | 28.4% | 7.1% |
| Perplexity | 10.8% | 0.0% |
| All three platforms | 23.6% | 3.6% |

That spread is a warning as much as a finding. Category language scores highest, so a prompt set chosen to produce a good number will fill up with it, and the resulting score will describe questions your market does not actually ask. Choose prompts for how closely they match real buying language, then read the score knowing which kind you chose.

### What happened when we changed our own prompt set

RankX AI tracks its own site, and the panel changed size on 30 August 2026 when the set went from 8 prompts to 45 across five assistants. The pre-expansion reading was recorded that day: 61 mentions in 328 analysed checks over 30 days, a rate of 18.6 per cent. Read again on 3 September, the same 30-day window returned 61 mentions in 351 analysed checks, a rate of 17.4 per cent.

The mention count did not move at all, on any platform. ChatGPT 15, Claude 13, Grok 14, Gemini 10 and Perplexity 9, identical in both readings. Twenty-three further checks were analysed and none of them named the brand, so every per-platform rate fell between 0.2 and 1.9 points with the numerator untouched. Nothing about the site changed in those four days except that four blog posts were published.

Four days and 23 checks cannot tell you whether the 37 new questions are good ones, and this API does not break checks down by prompt. The growth is also small enough to be the original eight prompts still accruing, which is the likelier reading, because 45 prompts across five assistants would produce far more than 23 checks in four days.

That turns out to be the more useful finding. A prompt set expanded last week is still reporting last month's questions, because the new ones have no history and the old ones dominate the window for as long as the window is. An AI visibility rate is a fraction whose denominator is the set somebody chose, so a comparison across the boundary of a set change is two different panels with the difference written up as a result.

## How many prompts should you track, and how often?

Track enough prompts that one of them changing its mind cannot move your reported number, which means dozens rather than a handful. The reason is instability rather than sample size etiquette: assistants give different answers to identical questions, so a small set produces a number that swings on nothing.

| Decision | Starting point | What it rests on |
| --- | --- | --- |
| How many prompts | Dozens, not a handful | RankX AI tracks 45. Below that, one prompt flipping moves the reported rate |
| Prompt length | 8 to 15 words | Two thirds of real prompts, Stella Rising, January 2026 |
| Mix by buying stage | Roughly even across the three | Ours splits 16 ready-to-buy, 16 comparing, 13 researching |
| Prompts naming your brand | One or two at most | Judgement. The useful question is whether you appear unprompted |
| Repeats per prompt | A fixed schedule, never one check | Only 8 of 216 prompt and platform pairs were stable across seven runs |
| Comparing two periods | Equal windows, unchanged set | A set changed mid-window makes two panels look like one result |

Go Fish Digital's seven-run design measured that directly. Of 216 unique prompt and platform combinations, only 8 named the brand in all seven runs. A prompt checked once tells you what one run said. The same prompt checked thirty times tells you a rate, and the rate is the only figure worth putting in a report.

RankX AI's own set is 45 active prompts across five assistants, which produced 351 analysed checks in the 30 days to 3 September 2026. Below roughly 30 checks per platform the tool refuses to generalise from its own numbers and says so in the response, which is the behaviour to want: a rate computed from four checks is a real measurement of almost nothing.

Re-run the whole set on a fixed schedule and compare windows of equal length. Adding prompts mid-window makes two windows incomparable, which is what produced the drop above. The wider argument for measuring a panel rather than spot-checking answers sits with [AI share of voice](/blog/ai-share-of-voice).

## How should a prompt set be classified?

Classify prompts by buying stage and by whether they name a brand, because those two splits decide what a movement in the number means. A set weighted to one stage produces a score that answers one question and gets read as though it answered all of them.

RankX AI's 45 tracked prompts split 16 ready-to-buy, 16 comparing and 13 researching. Ready-to-buy prompts ask which tool to pick. Comparing prompts name rivals, including six of the form "X alternatives". Researching prompts ask what a thing is, and those are the ones a blog post can win.

Keep the set unbranded apart from a deliberate handful. A prompt containing your own brand name measures whether the assistant knows who you are, which is worth one or two slots and nothing more, because the interesting question is whether you appear when nobody mentioned you. Watch for homonyms: on this site's own tracking, every in-Overview brand mention on one query turned out to be about the RANKX function in Power BI.

One practical detail from the same set. Two prompts had identical text under different stage labels, and one is now switched off. Deduplicate by text before you count, since a duplicate quietly doubles one question's weight in the rate.

## What do you do with a prompt set once you have one?

A prompt set earns its keep when it changes your content strategy, and the route from one to the other is short. Group the prompts into the sub-questions they share, check which of those sub-questions your site answers in one self-contained passage, and write the ones it does not. Relevance to a specific sub-question is how pages get cited, and how a brand enters AI recommendations at all.

The prompts where competitors appear and you do not are the brief. Read the answer text rather than the score: an answer that lists four rivals is telling you which comparison the assistant thinks it is making, and which page you are missing. Answers that name you without linking are worth tracking separately, since an [unlinked brand mention](/glossary/unlinked-brand-mention) counts as visibility and sends no traffic, while an AI citation with a link does both.

Then measure the same set again on the same schedule. What the whole method rests on is comparability, and the surrounding measurement stack is set out in [how to measure AI search visibility](/blog/measure-ai-search-visibility). Our own first published reading, method included, is [the AI visibility baseline](/blog/our-ai-visibility-baseline).

## How RankX AI does prompt research

RankX AI tracks a fixed prompt set across ChatGPT, Claude, Gemini, Perplexity and Grok, classified by buying stage, and reports mention rate against the checks that returned a verdict rather than against every check attempted. Unanalysed checks are shown as unknown rather than counted as a miss.

The figures in this article are from that account, read on 3 September 2026, and the pre-expansion baseline they are compared against was committed to the repository on the day the set changed. That is our own data on our own site, which is the only prompt data on this page nobody had to model.

Choosing the prompt set is its own surface now. [AI prompt research](/features/ai-prompt-research) takes up to eight seed topics and returns the questions people ask assistants around them, each with an estimate of the demand behind it and a flag on anything close to a prompt already tracked. That estimate is modelled from Google's People Also Ask results rather than counted inside an assistant, which is the same caution this article applies to everybody else's prompt numbers, applied to ours.

## Does prompt research replace keyword research?

No. The two answer different questions and neither substitutes for the other. Keyword research tells you how many people type a phrase into a search engine, which is still the demand signal behind Google's own surfaces, and AI Overviews sit on top of those results. Prompt research tells you how an assistant answers a question and whether it names you, which no keyword tool reports. Run both, and expect prompt research to inform which pages you write while keyword research still decides which of them can rank.

## Can you get prompt volume data?

Not for an individual prompt, and the vendors selling prompt research say so. Semrush's own documentation states that individual prompts are often too specific and unique to measure directly, which is why it reports volume for topics that group semantically related prompts rather than for the prompts themselves. No assistant publishes query data, so every prompt volume figure you are shown is modelled from clickstream panels, browser extensions or classic keyword data. Use it to order a list, never to build a forecast.

## How long should a tracked prompt be?

Roughly 8 to 15 words, because that is what people type. Stella Rising's January 2026 survey of 524 active assistant users found two thirds writing prompts of 15 words or fewer, and Semrush clickstream puts the average ChatGPT prompt between 4.2 and 8.7 words. Long prompts also measure less than they appear to: Go Fish Digital found prompts of 18 words or more were shortened by 56.1 per cent before retrieval ran, so a 40-word buyer scenario is partly testing text the retrieval layer discards.

## Can you do prompt research without paying for a tool?

Yes, for the discovery half. Your sales and support inbox is the highest-value free source because those questions already convert. Google Search Console shows the question-shaped search queries your site already earns, and community threads on Reddit and Quora carry the same questions in unpolished language. OpenAI Signals publishes what people bring to ChatGPT, and WildChat is a freely licensed corpus of about a million real conversations. What you cannot get free is repeated measurement across several AI platforms over time, which is the half that produces a rate rather than a list.

## How many prompts should you track?

Enough that one prompt changing its mind cannot move your reported number, which in practice means dozens rather than a handful. Answers are unstable between identical runs: Go Fish Digital ran 73 prompt variations seven times each across three platforms and only 8 of 216 prompt and platform combinations named the brand in all seven runs. Track a fixed set, run it repeatedly, and report the rate with its denominator rather than reporting individual answers.

## Why did our AI visibility score drop when we added prompts?

Most likely because the score is a fraction and you changed its denominator. Adding questions you have never been cited on adds checks that count against you from the first run, so the rate falls even when nothing about your site or your existing answers changed. There is a slower version of the same problem worth knowing about: new prompts have no history, so for as long as your reporting window is, the rate you are shown still mostly describes the old set. Either way the fix is to compare like sets, or to segment old prompts from new ones.
