Methodology

The prompt universe

To measure how a brand appears in AI answers, you first need to define the questions you're measuring against. This methodology explains how we build and classify those questions, how often we run them, and how to interpret the resulting metrics, including what they can and cannot tell you.

What it is
The set of questions we send to AI engines on your behalf, week after week. We call it the prompt universe.
Typical size
A few thousand prompts per brand topic or line of business, determined by the relevant prompt universe and observed behavior.
Built from
Five data sources, none of which is sufficient alone.
Labeling
Seven attributes recorded on every prompt at the time it's written.
How it's run
Full set weekly, with options for a smaller subset to run daily. Each prompt runs once against each enabled engine per cycle. The reported figure is calculated across every prompt × engine observation in the set.
Unit of measurement
The intent cluster, not the individual prompt. A single prompt's result is one observation and isn't meant to be read on its own.
Who can change it
You can see all of it, exclude anything from it, and set the priorities that drive the weighting.

The process, end to end

Six stages

1 2 3 4 5 6 Category & journeys Demand universe Prompt clusters Representative coverage Localize by market Validate & refresh
1

Define the category and user journeys

We map the major intents that matter across the customer journey (discovery, education, comparison, evaluation, recommendation, purchase and post-purchase) for each brand, product, category and market.

2

Build the underlying demand universe

We combine multiple demand signals (search behavior, clickstream and panel data, AI interaction patterns, your own inputs, and the category and product taxonomy) to work out the questions and needs consumers are likely to express.

3

Generate prompt clusters

Those intents are translated into natural-language prompts that reflect how people actually interact with AI systems. Each underlying intent can have dozens of linguistic and contextual permutations, rather than being represented by a single rigid query.

4

Ensure representative coverage

Prompts are balanced across funnel stage, topic, product and category, consumer need, brand and non-brand intent, and relevant market characteristics, so the dataset doesn't become disproportionately influenced by any one theme.

5

Localize by market (when needed)

Prompt sets are adapted for local language, terminology, products and market context rather than translating a global English prompt list.

6

Continuously validate and refresh

As consumer behavior, AI usage, products and categories evolve, the prompt universe is reviewed and refreshed. The objective is to keep the measurement framework stable while making sure it continues to represent real demand.

01 · Where the prompts come from

Five sources, each partial

The fundamental challenge is that there is no direct view into what people type into AI engines. Unlike traditional search, there is no equivalent of Search Console that shows the full universe of questions users are asking. So answering a seemingly simple question (“what are people asking about my category?”) requires us to infer demand from multiple data sources, each capturing a different part of the picture.

It's important to distinguish between those sources rather than treat them as interchangeable. Each contributes a different signal, with different strengths, limitations and levels of representativeness. The methodology works by combining those signals to build the most representative possible view of the underlying question universe.

Consumer panel and clickstream

Licensed data on what people browse, compare and buy around a topic. Good evidence for which needs are common and which are rare.

Doesn't show how anyone phrases a question to an AI engine.

Observed AI prompts

Hundreds of millions of real prompts entered into AI engines, sourced through licensed, statistically representative consumer panels. This is the only source that captures how people actually phrase their questions in AI environments, not just what they search for, or what they might ask. That direct behavioral signal makes it disproportionately valuable, even relative to much larger indirect datasets.

For a narrow or new category there may be very little of it, and we say so rather than pretending otherwise.

Search and keyword demand

Large, stable and reliable for working out how much interest a topic carries relative to others. We use it mainly for weighting.

Search queries are shorter and worded differently from prompts. Treating a keyword list as a prompt list is the most common way these sets go wrong.

Your Search Console and site search

First-party data showing what people search for when they are already finding or using your brand. Unlike panel data, it is specific to your customers and products rather than the broader category.

Focused on users who already reached you, and largely rooted in traditional search rather than AI-native behavior.

Your category and competitor list

The business priorities, products, competitors, markets and terms that have to be covered whether or not the data says they're high-volume. A new launch has no search history yet.

This is organizational priorities and judgement, not data, so it's the part you have most say over.

What happens to the different intent clusters and themes

The combined pool is deduplicated, grouped by intent (questions that are asking the same thing in different words end up together) and then weighted.

Weighting determines how much each topic contributes to the measurement. Because visibility is calculated across the full prompt set, the mix of prompts ultimately defines what the score represents. If a topic represents 2% of real demand but 20% of the prompts, it will disproportionately drive the result. Weighting prevents this by aligning the composition of the prompt universe as closely as possible with the underlying distribution of demand.

The categories and themes document, agreed before anything is written

That structure is documented before any prompts are written. We create a categories and themes framework defining what we propose to measure (the categories, themes, intents, and their relative importance) and send it for review and sign-off.

This is where you shape the measurement: add a missing theme, increase the weight of a business priority, remove an out-of-scope product line, or shift emphasis toward a particular part of the customer journey. We do this early because everything downstream follows from this structure.

This is also the most important approval in the process. Reviewing a prompt tells you whether a sentence is right; reviewing the framework tells you whether we are measuring the right things in the first place.

Sources Consumer panel & clickstream Observed AI prompts Search & keyword demand Your Search Console & site search Your category & competitor list Group by intent remove duplicates, roll up into themes Weight by the demand each topic carries Categories & themes doc you sign it off The prompt set a few thousand, labeled on 7 attributes Your priorities, objectives and exclusions
The dashed line is where brands have direct input. Priorities agreed at sign-off raise or lower the weight on a topic, and anything on the exclusion list never enters the set at all. Nothing is written until the document is signed.

02 · How each prompt is labeled

Seven attributes, recorded when the prompt is formalized

Every prompt is tagged across dimensions such as branded vs. unbranded, persona, category, intent, funnel stage, query type, and market. These labels are what make the data explorable: any dimension you want to analyze or filter later must be captured at the prompt level from the start.

If a dimension isn't recorded, it can't simply be reconstructed later without rebuilding and re-running the prompt set.

“What's the best electric shaver if I have really sensitive skin?” Market Language Persona Category Funnel stage Product line Query type United States English Sensitive skin Skin Consideration Electric shavers Unbranded Which makes these reportable Share of voice by market Visibility by funnel stage Branded vs unbranded gap
A consumer example because it reads quickly. The same seven attributes apply to a B2B prompt: “Which managed detection and response vendors work well for a mid-size bank?” is US / English / mid-size financial services / security / consideration / MDR / unbranded.

03 · What's in the prompt universe

Mostly unbranded questions

Most prompts do not mention a brand. They reflect a need, problem, or decision (“what should I use for X?”) rather than a company name.

The branded/unbranded mix is calibrated to your actual search behavior, so it varies by brand. An established brand may warrant more branded prompts; a challenger, fewer. Around 80/20 is common, but it is an output of the data, not a fixed rule.

Categories

Prompts are distributed across categories in proportion to how much of the real search demand each one carries. A core product line gets more of the set than an accessory line. The alternative, an equal split, would quietly overweight your smallest categories and drag the headline number toward parts of the business that few people ask about.

Tags

Prompts can also be tagged and grouped in multiple ways. The category and theme structure provides one business view of the data, shaped around the market and your priorities, but it is not the only one. Additional tags allow the same prompt universe to be sliced across dimensions such as intent, persona, funnel stage, query type, or emerging themes, making it possible to surface patterns and trends that may not be visible through the primary category structure alone.

The whole set No brand named Brand named about 80% about 20% Inside the branded fifth names a specific competitor brand on its own Share of the set, by category equal split would be 20% 23% 21.5% 20% 18.5% 17% Core line Second line Adjacent Accessories Niche
Illustrative figures for a five-category brand. Most of the branded fifth names a specific competitor, because that's how people ask once they're close to deciding; those prompts are reported separately rather than folded into the headline number. The category spread here is modest because demand in this example is fairly even. Where it isn't, the spread is wider.

04 · How the set is sampled and run

Breadth vs. repetition

AI answers are inherently non-deterministic. Ask the same question multiple times and the response can change: different brands may appear, rankings may shift, and the supporting sources may vary. That volatility is a property of the systems themselves, which means any measurement methodology has to decide how to sample through it rather than assume a single answer is definitive.

There's more than one defensible way to sample AI answers, and the right one depends on what you're trying to find out. A broad monitoring program measuring overall visibility has different sampling needs from a narrow research question about one prompt, one theme, or the effect of one change.

There are two primary ways to address the non-deterministic nature of AI answers, with different trade-offs: (1) run the same prompt repeatedly, or (2) run the same underlying intent across many permutations. Brandlight allocates to the second.

This method better reflects how modern AI engines actually work. A user prompt is often not treated as a single static query. Depending on the engine and task, the system may rewrite or fan out the original prompt into multiple underlying searches, retrieval queries, sub-questions, or reasoning steps, retrieve different sources or candidates, rank those results, and then synthesize them into a final response. Small differences in wording can therefore change the query decomposition, sources retrieved, products or brands considered, and ultimately the answer generated.

For example, “What is the best phone for photography?” and “Which smartphone has the best camera for travel?” express a very similar underlying intent, but may cause an AI engine to generate different supporting searches, retrieve different evidence, and construct a different recommendation set. Those variations are part of the real-world surface area we are trying to measure.

We could execute the exact same prompt 20 times, and that would help quantify the variance of that specific formulation. But our methodology generally allocates that compute toward more representative prompts and permutations of the same intent, because this captures both model-level non-determinism and the variability introduced by query rewriting, retrieval, fan-out, ranking, and synthesis.

Repeatedly running the same prompt n times does reduce volatility around that prompt's observed outcome, but it comes with a compute and cost trade-off, and beyond a certain point the incremental reduction in volatility may not justify the additional executions compared with using that capacity to broaden the representative query universe.

Run the same sentence again: variation starts here The wording Rewrite & fan-out Retrieve sources Rank candidates Synthesize answer Ask the same need a different way: variation starts here
The two strategies enter the chain at different points. Repeating a sentence re-rolls the stochastic steps from the same starting point. Changing the wording moves the starting point, so the decomposition, the sources retrieved and the shortlist can all differ.

They measure different things

This is the part worth being precise about. The two approaches aren't a better and a worse version of the same measurement. They estimate different quantities.

Running one sentence twenty times tells you how much that sentence varies. Running twenty phrasings of the same need tells you how much the need varies, which includes the model's own randomness plus everything the wording changes on the way through. For a brand deciding where to spend effort, the second is the useful number, because you can't choose which sentence your customers type.

What runs in production

Each prompt runs once against each enabled engine per cycle. The visibility figure is then calculated across every prompt × engine observation in the set, so the sample supporting the reported number is the whole query universe, not any single prompt. Repeat rates can be raised deliberately for a specific experiment, such as testing whether a change actually moved something. That's a research setup, not the standard one.

05 · What follows for optimization

Why we don't optimize for a single prompt

If no single sentence represents how people express a need, then improving your position on one exact prompt is a narrow result. It may disappear as soon as the wording changes, the engine takes a different retrieval path, or the fan-out produces a different set of underlying searches.

So the unit of optimization is the intent, not the individual prompt.

Across the many ways people express the same need, the useful question is what consistently shapes the answer: which concepts and entities recur, which sources engines repeatedly rely on, which claims are reinforced, and where your brand is simply absent from the information available to retrieve or cite. Those are the signals that generalize across prompt variations.

The practical difference is important. “We don't appear for prompt 4,127” is not a meaningful work item. “Across sixty ways people ask about durability, four sources dominate the answers and we appear in none of them” is.

That is where optimization becomes actionable: improving the underlying information environment around an intent so the brand is more likely to be retrieved, cited, and recommended across the full set of ways people ask.

06 · How the prompt universe is refreshed and changed

A stable core to maintain a baseline, and a layer that evolves with the market

There's a tension built into this. A set that never changes gradually stops matching the market and user behavior: new products, new competitors, new ways of asking. A set that changes constantly has no trend line, because month nine isn't measuring the same thing as month one, and any movement could just be the set moving.

So a core set stays fixed and carries the comparable trend. An evolving layer that's part of it absorbs new topics and retires prompts that have stopped being asked. When you look at a change over time, you're looking at the core.

Run against the engines Native-speaker review of phrasing Compared with search and clickstream data Quarterly audit and sign-off Core set fixed: carries the trend Evolving layer new and retired prompts
Each station can change the evolving layer. Only the quarterly audit changes the core, and when it does, the version is recorded so a break in the trend can be explained rather than discovered.

07 · Who controls the set

We deliver it; you can see, influence and change it

Some tools ask customers to provide the prompts themselves. We don't, because building a representative prompt universe is a data problem, not a brainstorming exercise. Starting from intuition or simple business categories tends to produce the questions we'd expect or hope people are asking, rather than a balanced view of actual demand.

That does not mean the process is a black box. You retain full visibility and control:

  • You can review and audit the complete prompt set at any time.
  • You can exclude topics, competitors, products, or questions that should not be included.
  • You define the business priorities that inform weighting, the input with the greatest influence on the headline metric.
  • Every version of the prompt set is retained, so changes in performance can be separated from changes in what is being measured.

The division of responsibility is simple: we are accountable for the methodology, representativeness, and coverage of the set; you are accountable for ensuring the business priorities are right and that nothing is being measured that should be out of scope.

08 · Limits

What this method can't tell you

There are limits worth stating up front.

It's a sample, not observation

We're not observing real user conversations. We're running our own set of questions and measuring what comes back. Structurally it's the same as a brand tracking study or a TV audience panel: nobody watches every living room, they construct a sample and reason about the population from it. That reasoning holds when the sample is built to match real demand and breaks when it isn't. That is what this entire process is for, and why the composition of the set is worth more of your attention than the headline number.

A single prompt's result isn't a measurement

Because each prompt runs once per engine per cycle, an individual prompt result is one observation. If a prompt shows you present one week and absent the next, that on its own means very little. It's inside the variation you'd expect from a non-deterministic system, and it's exactly the variation the breadth of the set is there to absorb.

The figures worth reading are the aggregates: an intent, a theme, a category, the whole set. If you genuinely need a prompt-level answer, usually to test whether a specific change worked, we can raise the repeat rate on that group of prompts, but that's a deliberate experiment rather than something the standard set supports.