Methodology
To measure how a brand appears in AI answers, you first need to define the questions you're measuring against. This methodology explains how we build and classify those questions, how often we run them, and how to interpret the resulting metrics, including what they can and cannot tell you.
The process, end to end
Define the category and user journeys
We map the major intents that matter across the customer journey (discovery, education, comparison, evaluation, recommendation, purchase and post-purchase) for each brand, product, category and market.
Build the underlying demand universe
We combine multiple demand signals (search behavior, clickstream and panel data, AI interaction patterns, your own inputs, and the category and product taxonomy) to work out the questions and needs consumers are likely to express.
Generate prompt clusters
Those intents are translated into natural-language prompts that reflect how people actually interact with AI systems. Each underlying intent can have dozens of linguistic and contextual permutations, rather than being represented by a single rigid query.
Ensure representative coverage
Prompts are balanced across funnel stage, topic, product and category, consumer need, brand and non-brand intent, and relevant market characteristics, so the dataset doesn't become disproportionately influenced by any one theme.
Localize by market (when needed)
Prompt sets are adapted for local language, terminology, products and market context rather than translating a global English prompt list.
Continuously validate and refresh
As consumer behavior, AI usage, products and categories evolve, the prompt universe is reviewed and refreshed. The objective is to keep the measurement framework stable while making sure it continues to represent real demand.
01 · Where the prompts come from
The fundamental challenge is that there is no direct view into what people type into AI engines. Unlike traditional search, there is no equivalent of Search Console that shows the full universe of questions users are asking. So answering a seemingly simple question (“what are people asking about my category?”) requires us to infer demand from multiple data sources, each capturing a different part of the picture.
It's important to distinguish between those sources rather than treat them as interchangeable. Each contributes a different signal, with different strengths, limitations and levels of representativeness. The methodology works by combining those signals to build the most representative possible view of the underlying question universe.
Consumer panel and clickstream
Licensed data on what people browse, compare and buy around a topic. Good evidence for which needs are common and which are rare.
Doesn't show how anyone phrases a question to an AI engine.
Observed AI prompts
Hundreds of millions of real prompts entered into AI engines, sourced through licensed, statistically representative consumer panels. This is the only source that captures how people actually phrase their questions in AI environments, not just what they search for, or what they might ask. That direct behavioral signal makes it disproportionately valuable, even relative to much larger indirect datasets.
For a narrow or new category there may be very little of it, and we say so rather than pretending otherwise.
Search and keyword demand
Large, stable and reliable for working out how much interest a topic carries relative to others. We use it mainly for weighting.
Search queries are shorter and worded differently from prompts. Treating a keyword list as a prompt list is the most common way these sets go wrong.
Your Search Console and site search
First-party data showing what people search for when they are already finding or using your brand. Unlike panel data, it is specific to your customers and products rather than the broader category.
Focused on users who already reached you, and largely rooted in traditional search rather than AI-native behavior.
Your category and competitor list
The business priorities, products, competitors, markets and terms that have to be covered whether or not the data says they're high-volume. A new launch has no search history yet.
This is organizational priorities and judgement, not data, so it's the part you have most say over.
The combined pool is deduplicated, grouped by intent (questions that are asking the same thing in different words end up together) and then weighted.
Weighting determines how much each topic contributes to the measurement. Because visibility is calculated across the full prompt set, the mix of prompts ultimately defines what the score represents. If a topic represents 2% of real demand but 20% of the prompts, it will disproportionately drive the result. Weighting prevents this by aligning the composition of the prompt universe as closely as possible with the underlying distribution of demand.
That structure is documented before any prompts are written. We create a categories and themes framework defining what we propose to measure (the categories, themes, intents, and their relative importance) and send it for review and sign-off.
This is where you shape the measurement: add a missing theme, increase the weight of a business priority, remove an out-of-scope product line, or shift emphasis toward a particular part of the customer journey. We do this early because everything downstream follows from this structure.
This is also the most important approval in the process. Reviewing a prompt tells you whether a sentence is right; reviewing the framework tells you whether we are measuring the right things in the first place.
02 · How each prompt is labeled
Every prompt is tagged across dimensions such as branded vs. unbranded, persona, category, intent, funnel stage, query type, and market. These labels are what make the data explorable: any dimension you want to analyze or filter later must be captured at the prompt level from the start.
If a dimension isn't recorded, it can't simply be reconstructed later without rebuilding and re-running the prompt set.
03 · What's in the prompt universe
Most prompts do not mention a brand. They reflect a need, problem, or decision (“what should I use for X?”) rather than a company name.
The branded/unbranded mix is calibrated to your actual search behavior, so it varies by brand. An established brand may warrant more branded prompts; a challenger, fewer. Around 80/20 is common, but it is an output of the data, not a fixed rule.
Prompts are distributed across categories in proportion to how much of the real search demand each one carries. A core product line gets more of the set than an accessory line. The alternative, an equal split, would quietly overweight your smallest categories and drag the headline number toward parts of the business that few people ask about.
Prompts can also be tagged and grouped in multiple ways. The category and theme structure provides one business view of the data, shaped around the market and your priorities, but it is not the only one. Additional tags allow the same prompt universe to be sliced across dimensions such as intent, persona, funnel stage, query type, or emerging themes, making it possible to surface patterns and trends that may not be visible through the primary category structure alone.
04 · How the set is sampled and run
AI answers are inherently non-deterministic. Ask the same question multiple times and the response can change: different brands may appear, rankings may shift, and the supporting sources may vary. That volatility is a property of the systems themselves, which means any measurement methodology has to decide how to sample through it rather than assume a single answer is definitive.
There's more than one defensible way to sample AI answers, and the right one depends on what you're trying to find out. A broad monitoring program measuring overall visibility has different sampling needs from a narrow research question about one prompt, one theme, or the effect of one change.
There are two primary ways to address the non-deterministic nature of AI answers, with different trade-offs: (1) run the same prompt repeatedly, or (2) run the same underlying intent across many permutations. Brandlight allocates to the second.
This method better reflects how modern AI engines actually work. A user prompt is often not treated as a single static query. Depending on the engine and task, the system may rewrite or fan out the original prompt into multiple underlying searches, retrieval queries, sub-questions, or reasoning steps, retrieve different sources or candidates, rank those results, and then synthesize them into a final response. Small differences in wording can therefore change the query decomposition, sources retrieved, products or brands considered, and ultimately the answer generated.
For example, “What is the best phone for photography?” and “Which smartphone has the best camera for travel?” express a very similar underlying intent, but may cause an AI engine to generate different supporting searches, retrieve different evidence, and construct a different recommendation set. Those variations are part of the real-world surface area we are trying to measure.
We could execute the exact same prompt 20 times, and that would help quantify the variance of that specific formulation. But our methodology generally allocates that compute toward more representative prompts and permutations of the same intent, because this captures both model-level non-determinism and the variability introduced by query rewriting, retrieval, fan-out, ranking, and synthesis.
Repeatedly running the same prompt n times does reduce volatility around that prompt's observed outcome, but it comes with a compute and cost trade-off, and beyond a certain point the incremental reduction in volatility may not justify the additional executions compared with using that capacity to broaden the representative query universe.
This is the part worth being precise about. The two approaches aren't a better and a worse version of the same measurement. They estimate different quantities.
Running one sentence twenty times tells you how much that sentence varies. Running twenty phrasings of the same need tells you how much the need varies, which includes the model's own randomness plus everything the wording changes on the way through. For a brand deciding where to spend effort, the second is the useful number, because you can't choose which sentence your customers type.
Each prompt runs once against each enabled engine per cycle. The visibility figure is then calculated across every prompt × engine observation in the set, so the sample supporting the reported number is the whole query universe, not any single prompt. Repeat rates can be raised deliberately for a specific experiment, such as testing whether a change actually moved something. That's a research setup, not the standard one.
05 · What follows for optimization
If no single sentence represents how people express a need, then improving your position on one exact prompt is a narrow result. It may disappear as soon as the wording changes, the engine takes a different retrieval path, or the fan-out produces a different set of underlying searches.
So the unit of optimization is the intent, not the individual prompt.
Across the many ways people express the same need, the useful question is what consistently shapes the answer: which concepts and entities recur, which sources engines repeatedly rely on, which claims are reinforced, and where your brand is simply absent from the information available to retrieve or cite. Those are the signals that generalize across prompt variations.
The practical difference is important. “We don't appear for prompt 4,127” is not a meaningful work item. “Across sixty ways people ask about durability, four sources dominate the answers and we appear in none of them” is.
That is where optimization becomes actionable: improving the underlying information environment around an intent so the brand is more likely to be retrieved, cited, and recommended across the full set of ways people ask.
06 · How the prompt universe is refreshed and changed
There's a tension built into this. A set that never changes gradually stops matching the market and user behavior: new products, new competitors, new ways of asking. A set that changes constantly has no trend line, because month nine isn't measuring the same thing as month one, and any movement could just be the set moving.
So a core set stays fixed and carries the comparable trend. An evolving layer that's part of it absorbs new topics and retires prompts that have stopped being asked. When you look at a change over time, you're looking at the core.
07 · Who controls the set
Some tools ask customers to provide the prompts themselves. We don't, because building a representative prompt universe is a data problem, not a brainstorming exercise. Starting from intuition or simple business categories tends to produce the questions we'd expect or hope people are asking, rather than a balanced view of actual demand.
That does not mean the process is a black box. You retain full visibility and control:
The division of responsibility is simple: we are accountable for the methodology, representativeness, and coverage of the set; you are accountable for ensuring the business priorities are right and that nothing is being measured that should be out of scope.
08 · Limits
There are limits worth stating up front.
We're not observing real user conversations. We're running our own set of questions and measuring what comes back. Structurally it's the same as a brand tracking study or a TV audience panel: nobody watches every living room, they construct a sample and reason about the population from it. That reasoning holds when the sample is built to match real demand and breaks when it isn't. That is what this entire process is for, and why the composition of the set is worth more of your attention than the headline number.
Because each prompt runs once per engine per cycle, an individual prompt result is one observation. If a prompt shows you present one week and absent the next, that on its own means very little. It's inside the variation you'd expect from a non-deterministic system, and it's exactly the variation the breadth of the set is there to absorb.
The figures worth reading are the aggregates: an intent, a theme, a category, the whole set. If you genuinely need a prompt-level answer, usually to test whether a specific change worked, we can raise the repeat rate on that group of prompts, but that's a deliberate experiment rather than something the standard set supports.