Ask an AI assistant the same question five times and you can get five different answers. We spent two months measuring that noise, and it settles the two questions every visibility program faces: how often to check, and whether to bunch the checks or spread them out.
Two findings carry everything below. First, where a brand sits decides how hard it is to measure: brands that almost always appear in AI answers disagreed with themselves on roughly one morning in eight, while brands on the bubble disagreed on up to seven mornings in ten. Second, bunching or spreading your repeat asks is a provider setting: on Google and OpenAI a burst of five asks is worth five independent reads, while on Anthropic whole mornings flip together and the same burst is worth about three and a half.
Here's where those numbers come from. In July 2026 we asked three assistants (OpenAI, Anthropic, Google) the same shopping questions five times within a few minutes, every morning, for three weeks: 1,357 scored answers. In August we repeated the run on brands picked because they sometimes appear in AI answers and sometimes don't. Why one AI answer is never enough explains the principle those runs confirmed: an AI answer is a sample. This is the data.
What answer noise actually looks like
One real cell of that data: one question, one brand, OpenAI, five asks per morning for three weeks. The brand and the question held still the whole time.
Each column is one morning. Each dot marks where the brand ranked in one answer; red crosses are answers that left the brand out entirely. The scatter inside a single column often matches the drift across the whole three weeks: same question, same morning, a #1 and a #13 minutes apart.
The movement belongs to the sampler. An assistant generates a fresh answer every time, and brands near the edge of an answer ride that randomness in and out. Any single reading is mostly noise, which is why a reading belongs in a trend, never in a headline.
Always, never, and sometimes are three different problems
The July pass ran on brands that appear in 96 to 98 percent of answers. The August pass deliberately seeded brands picked at a historical appearance rate of 20 to 80 percent. Same instrument, very different behavior:
On always-present brands, presence is nearly settled: 8 to 19 percent of mornings split. On the bubble, disagreement dominates for Google (71 percent) and OpenAI (67 percent). OpenAI's bubble numbers come from 15 mornings before its pinned model retired, so treat them as directional.
That splits measurement into three regimes, and each one gets its own playbook:
Where the brand sits | How often to check | Reads per check | Report as |
|---|---|---|---|
Always shows up (85%+) | Weekly; rank moves slowly and presence is settled | Light; noise is a quarter of contested brands' | A smoothed rank trend |
On the bubble (20 to 80%) | Every 3 to 5 days; a swing lives about 6 days | 3+ on Google and OpenAI; spread across days on Anthropic | A probability smoothed across checks, never a one-read badge |
Never shows up (near zero) | A light probe on the weekly cycle | Minimal; every ask repeats the first | The moment it first flips in, then treat as on the bubble |
The check intervals come from run-length analysis of how long each state lasts; the repeat experiment adds the read counts and the reporting rules.
The bubble row deserves its numbers spelled out. Near a 50 percent appearance rate a single ask is a coin flip, and even at 60 percent, calling the brand in or out with real confidence from one sitting takes about 90 asks. The honest report for a bubble brand is a probability smoothed across checks; a green or red badge from one reading claims more than one reading knows. And a presence swing lives about six days, so a weekly rhythm lets a whole swing come and go between readings.
One more finding cuts across all three regimes: the bubble itself moves. Of 14 brand-and-question pairs we picked because their history put them on the bubble, about six were still there when the August run started, a few weeks after selection. Three had settled at never and two at always; one previously steady pair had fallen almost out of the answers. The bubble is a place brands pass through, and the checking rhythm has to follow them.
Bunch your asks, or spread them out?
Both work. Which one works for a given channel depends on the provider, because each one's noise has a different shape.
Grey is the five asks you paid for; color is what they're worth as independent evidence. Google is cleanly binomial across 56 mornings, and OpenAI points the same way on a shorter window. Anthropic's bursts overlap because whole mornings succeed or fail together.
On Google and OpenAI, bunching pays full value: five asks are worth five asks, and averaging them cuts the noise exactly as fast as arithmetic promises. Anthropic behaves differently. Its answers within a morning agree almost perfectly, but whole mornings occasionally flip together: a brand that appears four mornings in five will sometimes produce a morning of five straight misses, which independent chance would produce about eight times in ten thousand. Against that pattern the fifth same-morning ask buys almost nothing, and the same budget spread across five days keeps close to its face value.
So bunching and spreading are provider settings. Bubble brands always take the sequential half too, because classifying a brand that hovers around 50 percent is a smoothing problem across days rather than a precision problem within one sitting.
Spend your reads where the answer is contested
Noise concentrates exactly where readings carry decisions.
Identical re-asks spread about four times wider on contested brands than on stable ones, for rank and sentiment both.
A stable #1 needs a light touch, and a contested brand deserves every repeat ask in the budget. The same read spent on a settled answer buys confirmation; spent on a contested one, it buys information.
How much of one reading is real
Averaged across the window, how much of the movement in a single daily reading is real change rather than answer noise? The answer splits sharply by provider.
Whiskers show confidence ranges from a cell-level bootstrap. Anthropic readings carry the most signal per read; Google readings need the most averaging before they mean anything.
On Anthropic, roughly half of what a single reading shows is genuine. On Google it's closer to one part in seven. Any confidence attached to a reading should reflect the channel it came from.
The practical rules fall straight out of the numbers. Check stable brands weekly and contested ones every few days. Bunch repeat asks on Google and OpenAI; spread them across days on Anthropic. Report a bubble brand as a probability. Weight every reading by its channel's reliability. This experiment is why Lantern reports read the way they do.
A note on honesty, because measurement earns trust by showing its discards: two early conclusions of this experiment failed our own review and were withdrawn. One claimed movement was mostly noise, so polling could slow down. One claimed day-to-day rank movement had strong momentum. Both collapsed under stricter controls, and every number in this article survived three separate estimation methods. The follow-up instrument keeps collecting, now across seven AI channels including the assistant products people actually use, so these numbers will keep getting sharper.
Common questions
Why does an AI assistant give different answers to the same question?
Each answer is a fresh generation from a probabilistic model, so the assistant samples a slightly different response every time. In our data the same question drew a #1 ranking and a #13 ranking minutes apart, with nothing changed in between.
How often should a brand check its AI visibility?
Weekly covers brands that always or never appear, because their state moves slowly. Brands that sometimes appear need a check every few days: their swings last about six days, and a weekly rhythm can miss one entirely.
How many times should the same question be asked?
On Google and OpenAI, ask several times in one sitting and average the answers; five asks there are worth five independent reads. On Anthropic, spread the same asks across days, because whole mornings succeed or fail together.
Can a single AI answer prove a brand is visible?
For a brand near a 50 percent appearance rate, one answer is a coin flip, and real confidence from one sitting would take about 90 asks. A probability smoothed across several checks tells the truth a single answer can't.