Get a demo Demo
Brandlight Research Lab · August 2026

The AI Engine
Volatility Study

July 2026 Edition
Published by ·

On one day in July, two of the largest AI engines quietly changed how they answer. Brandlight tracked 21 weekly snapshots across five industries, four markets and seven AI engines to separate what the engines did from what the market did.

5
Industries
7
AI engines
21
Weekly snapshots
BrandlightResearch
TL;DR

The whole study, in 60 seconds.

If you only read one page of this report, read this one. Every number below is unpacked in the sections that follow.

1 day
Two engines changed on the same date. Both were found in the measurement data first, before anyone looked at a vendor release calendar.
+21%
Google AI Mode started naming a fifth more brands per answer, and up to 58% in one market. A version-pinned Google model, asked the same questions that week, moved −2.4%.
−41%
ChatGPT cut the sources it cites by two fifths while naming the same brands at the same rate. All five industries in the panel landed between 35% and 43% down.
8x
One engine moved alone. AI Mode outmoved the next engine by up to eight times, in four industries that share no questions and no brands.
+11.5
Points of visibility gained by the average category leader. The bottom of the same category gained 0.13. This was concentration, not growth.
4 wks
The shift held and kept growing for a month. This was a permanent step, not a weekly wobble.
0
Of these changes were visible in a blended visibility score. Averaging your engines together hides exactly the thing that moved.
2
Promising findings we tested and threw out because the control moved as much as the engine did. That is the reason to believe the two that survived.
The Panel

Which industries is this actually based on?

Five industries across three markets and two languages. Here is exactly what each one contributes, including the two that could only test half the story, and the one we threw out.

BrandlightResearch

What each industry actually recorded

Both findings, per industry, in the two weekly snapshots either side of 21 July. Google AI Mode is the change in how many brands it names per answer. ChatGPT citations is the change in how many sources it points to. Sorted by the size of the Google effect.

Industry Market Engines Version-pinned control Google AI Mode ChatGPT citations
Consumer goods US, Canada 7 Yes +57.6% −35.3%
Automotive Italy, France 4 to 6 Yes, a pinned Google model +35.1% −41.2%
Health insurance United States 6 Yes +21.0% −41.1%
Banking US, Canada 5 to 7 Yes +20.7% −40.1%
Hospitality and resorts United States 5 Cross-engine only +9.8% −42.5%

Three things in that table are worth pausing on, because they are the difference between a study and a sales deck.

Every industry moved on Google AI Mode, and none of the others did. The range runs from +9.8% to +57.6%, and the ordering is not driven by anything the industries have in common, because they have nothing in common. They do not share questions, brands, buying cycles, seasonality or news.

Automotive carries the cleanest control in the study. Its Google measurement comes from an Italian account whose pinned engine is a fixed Google build. So a floating Google surface rose 35.1% in the same week that Google's own version-pinned model moved −2.4%, on identical questions. Same vendor, same week, one build frozen and one not. That is as close to a controlled experiment as this data gets.

Both findings appear in all five. Google AI Mode expanded in every industry, between +9.8% and +57.6%. ChatGPT contracted in every industry, between −35.3% and −42.5%. Two opposite movements, on two engines, in the same week, everywhere we looked.

5
Industries measured
Consumer goods, health insurance, banking, hospitality and resorts, automotive
4
Markets
United States, Canada, France, Italy
4
Carry a version-pinned control
A frozen build, asked the same questions in the same week
3
Languages
English, French and Italian, so the effect is not an artefact of one language
The Method

How do I know it was the engine and not my brand?

This is the question that kills most claims about AI search. Here is the design that answers it, and the reason you can trust everything that follows.

Suppose your visibility on one AI engine jumps six points in a week. Before you credit the engine, look at everything else that could have caused it. Your team launched a campaign. A journalist wrote something. A competitor had a recall. The category got busy. The measurement itself changed. Every one of those produces the same shape on a chart: a step in one line.

A single engine's history cannot separate them. But there is something better available, and it is free.

Every engine is asked the same questions in the same week. So whatever happened in the real world that week hits all of them at once. If one engine jumps and the others are flat on identical questions, that is the engine. If they all move together, it is the world, and there is no story.

The control that cannot change

There is a stronger version of this. Most AI engines are live consumer products that shift silently underneath you whenever the vendor ships. But some can be run as a version-pinned model: a fixed build that does not change unless someone deliberately changes it.

That gives us a placebo. A pinned model cannot respond to a vendor release, because it did not receive one. So if the pinned model moves, the cause is upstream of the engines entirely: the questions changed, the news changed, or the measurement pipeline changed. If the pinned model sits still while one live engine leaps, the engine did something.

The Brandlight Take · 01
A finding only counts when the moving engine moves alone and the pinned control stays flat. Every number in this study was tested that way, including the ones we ended up throwing out.
Finding One

What actually happened on one day in July?

Google AI Mode started naming meaningfully more brands in every answer. It did this in four industries at once, in two countries, in the same week, while the pinned control did not move.

The engine that moved Best-performing rival engine Version-pinned control
BrandlightResearch

Change in brands named per answer

Measured across the two weekly snapshots either side of the event date. Higher means the engine mentions more brands in a typical answer. Five industries that share no questions, no brands and no market conditions.

Google AI ModeConsumer goods · Canada
+57.6%
Google AI ModeAutomotive · Italy
+35.1%
Google AI ModeHealth insurance · US
+21.0%
Google AI ModeBanking · US
+20.7%
Google AI ModeHospitality and resorts · US
+9.8%
Microsoft CopilotBest rival engine · Banking
+8.0%
Microsoft CopilotBest rival engine · Insurance
+3.2%
ChatGPTBest rival engine · Hospitality
+2.7%
Pinned Google modelAutomotive · Italy · same vendor, frozen build
−2.4%
Pinned controlHealth insurance · US
−1.5%

Read that chart slowly, because the important part is not the size of the top bar. It is the gap between the blue bars and the grey ones. In every single industry, one engine moved two and a half to eight times more than the best-performing engine next to it, and the model that physically cannot change moved backwards by one and a half percent.

These five industries have nothing in common. They do not share questions, brands, buying cycles, seasonality or news. The only thing they share is the engine answering them. When four unrelated markets move together on one engine and not on the others, the engine is the variable.

The matching event. On the same date, Google shipped a new, faster, cheaper model into Search, positioned for high-throughput answering. A model built for speed and volume producing more enumerative answers that name more brands is a coherent mechanism, and the timing lands inside our detection window.

The Brandlight Take · 02
A vendor shipped a routine model upgrade. Nobody announced a change to brand recommendations. Five industries had their competitive landscape rewritten inside seven days.
Finding Two

So who actually gained from it?

This is where it gets uncomfortable. The engine did not start naming more brands. It started naming the brands that were already winning, more often.

BrandlightResearch

Who the extra mentions went to

All tracked brands in one consumer goods category, split by how visible they already were before the event. Change in visibility, in percentage points.

The four leading brandsAlready the most visible
+11.53
The eight smaller brandsThe long tail of the category
+0.13
Pinned control, leadersSanity check
−0.04
Pinned control, smaller brandsSanity check
+0.02

The category leader gained more than twelve points of visibility in a single week. The brands at the bottom of the same category, answering the same questions, in the same market, on the same engine, gained essentially nothing. The pinned control confirms that neither group had any underlying change to explain it.

Put that in commercial terms. One week of engine behaviour moved the leaders further than most brands move in a quarter of sustained content investment. No campaign ran. No content shipped. No budget moved. The engine simply decided to lean harder on the names it already trusted.

This is the mechanic that should worry anyone who is not already the leader in their category. AI answers are not a neutral surface that rewards effort proportionally. When the underlying model changes, the advantage tends to compound toward whoever already had it, and it happens on a timescale no marketing plan can react to.

The Brandlight Take · 03
If you are the leader, a model upgrade is a windfall you did not earn. If you are the challenger, it is a setback you cannot appeal. Either way, you should know it happened.
Finding Three

Did it stick, or was it just a bad week?

Weekly measurement is noisy, and a one-week spike is usually nothing. This was not a spike. It stepped up, held, and kept climbing for a month.

BrandlightResearch

Four weeks either side of the event

Brands named per answer in the health insurance category, before and after the event date. The pinned control is shown underneath for comparison.

28 21 14 7 MODEL SHIPPED 19.73 23.88 24.41 24.97 pinned control Week before Week after Two weeks Three weeks

The engine jumped 21% in the first week and then never gave it back. It kept edging up for the next fortnight, which is the signature of a staged rollout reaching more of the traffic over time rather than a measurement blip correcting itself.

Underneath it, the pinned control traces an almost perfectly flat line across the same four weeks. Same questions, same category, same period, no movement. That flat line is what makes the blue line believable.

The Brandlight Take · 04
Treat a single week of AI visibility data as noise. Treat a step that holds for a month as a new baseline, and rebuild your expectations around it.
Finding Four

The change your dashboard never showed you

On the same day, a second engine did something completely different, and a visibility score would have reported absolutely nothing.

BrandlightResearch

Change in total sources cited, same week

How many source pages each engine cited across the whole category, comparing the two snapshots either side of the event date.

ChatGPTSources it pointed to
−41.1%
Google AI ModeSources it pointed to
+15.9%
Google AI OverviewSources it pointed to
+1.6%
PerplexitySources it pointed to
−0.1%
Microsoft CopilotSources it pointed to
−1.0%
Version-pinned controlThe model that cannot change
−3.3%

Here is the part that matters. Over that same week, ChatGPT's brand mention rate did not move at all. Visibility went from 82.94% to 82.63%. The number of brands it named per answer went from 18.94 to 19.00. If you were watching a visibility dashboard, the line was flat and your week was uneventful.

Underneath that flat line, the engine stopped pointing at two fifths of the web it used to point at. Same brands named, dramatically fewer sources credited.

The same result in every market we track

Automotive in France gives the clearest single view of it, because that account runs a version-pinned control next to the live engines. Different industry, different market, different language, different questions, different competitive set, same week, same result.

BrandlightResearch

Automotive, France: change in total sources cited

Automotive in France, measured on the same two weekly snapshots. This account runs a version-pinned control alongside the live engines, so the movement can be attributed rather than assumed.

ChatGPTAutomotive · France
−41.2%
GeminiAutomotive · France
−7.0%
PerplexityAutomotive · France
−0.4%
Version-pinned controlThe model that cannot change
−14.9%

Five industries, four markets, two languages, and one number. Every industry in the panel lost between a third and two fifths of its cited sources on this engine, in the same seven days, while the brands it named stayed put.

BrandlightResearch

The same contraction, in every industry

Change in total sources cited by this one engine, in the two weekly snapshots either side of 21 July, alongside what happened to the brand mention rate in the same week.

Industry Market Sources cited Brand mention rate
Hospitality and resorts United States −42.5% held steady
Automotive France −41.2% −0.48pp
Health insurance United States −41.1% −0.31pp
Banking Canada −40.1% held steady
Consumer goods United States −35.3% held steady

The automotive brand's mention rate barely moved either, going from 27.98% to 27.50%. Exactly the same signature: the same brands named, far fewer sources credited. The version-pinned control in that account moved 14.9%, so the live engine fell close to three times as far as the build that was frozen.

Independent corroboration from outside our data

This is not only our measurement. Independent tracking across more than a billion AI citations reported referral traffic from the same engine falling by more than half from that same date. Two different datasets, two different methods, two different organisations, the same direction and the same day.

It also explains something we saw in a different industry entirely: one challenger brand collapsed on this engine from 59% visibility to 12% in that single week, while its pinned control barely moved. A brand whose presence depended on the long tail of citations was always going to be the first casualty of a citation contraction.

−41%
Sources cited by the engine
Across the whole category, in one week
−0.31pt
Change in brand visibility
Statistically indistinguishable from no change
47pts
Lost by one challenger brand
From 59% to 12% on this engine alone
−35 to −43%
The range across all five industries
Four markets, two languages, one week
The Brandlight Take · 05
Being named and being cited are two different assets, and they can move in opposite directions in the same week. If you only measure one of them, you will be blindsided by the other.
The Discipline

How do I know you are not fooling yourselves?

A method that only ever finds things is not a method. Here are two findings that looked real, and did not survive the control.

BrandlightResearch

The drop that looked like a vendor change

Five weeks before the July event, one engine's visibility fell more than ten points in a single week. Convincing, until you look at every engine in the same week.

Engine Change in brands named per answer What that rules out
Google AI Mode−66%The apparent finding
Perplexity−49%Moved too
Google AI Overview−47%Moved too
ChatGPT−37%Moved too
Microsoft Copilot−34%Moved too
Gemini−33%Moved too
Version-pinned controlCannot change−35%Verdict: not a vendor change

Every engine collapsed together, including the model that physically cannot change. A vendor release cannot move a pinned build. So this was a change in the measurement pipeline, not a change in any AI engine. A second industry confirmed it: the same engine there moved four hundredths of a point that week.

Published as a vendor finding, this would have been completely wrong, and it would have looked every bit as convincing as the real one.

The second claim we withdrew

We also expected to report that the engine changed what kind of sources it leaned on. It did shift toward third-party publishers by roughly three points across the July event. But the pinned control shifted by slightly more on the same measure. When the control moves as much as the treatment, there is no finding, so we make no claim about where the engine grounds its answers.

The Brandlight Take · 06
Two of our most promising findings did not survive their own control. That is not a failure of the study. It is the only reason to believe the two that did.
What To Do

So what does this mean for my brand?

Four conclusions that follow directly from the evidence above, and one open question we have not answered.

1. Stop reading a blended AI visibility number

Both events in this study are invisible in a cross-engine average. Averaging the engine that moved 21% with five engines that did not produces a number that looks calm while your position is being rewritten. Per-engine tracking is not a refinement, it is the difference between seeing this and not seeing it.

2. Measure citations separately from mentions

One engine held its brand mentions perfectly steady while cutting cited sources by two fifths. These are different assets with different commercial value, and they moved in opposite directions in the same seven days. A single visibility metric cannot represent both.

3. Expect the gains to go upward, not outward

When a model changes, the evidence here says the benefit concentrates on brands that were already prominent. If you are challenging a category leader, model releases are a headwind you have to plan around rather than a lottery you might win.

4. Judge a change over a month, not a week

The real event held its new level and kept climbing for four weeks. The false one reversed and failed to replicate elsewhere. Duration and replication are what separate a signal from noise, and neither is available from a single snapshot.

The Brandlight Take · 07
AI visibility is not a score you audit once a quarter. It is a position that can be revalued overnight by a decision made at a vendor you have no relationship with, and the only defence is knowing the day it happens.