— RESOURCES

Sampling on the bubble: how often to check AI answers

Ask an AI assistant the same question five times and you can get five different answers. We spent two months measuring that noise. It settles how often to check a brand's AI visibility, and whether to bunch the checks or spread them.

In July 2026 we asked three assistants (OpenAI, Anthropic, Google) the same shopping questions five times within a few minutes, every morning, for three weeks: over 1,300 scored answers. In August we repeated the run on brands picked because they sometimes appear in AI answers and sometimes don't. Why one AI answer is never enough explains the principle those runs confirmed: an AI answer is a sample. The data settles two practical questions. How often should you check a brand's AI visibility? And when you check, should the asks bunch into one sitting or spread across days?

What answer noise actually looks like

Here's one real cell of that data: one question, one brand, OpenAI, five asks per morning for three weeks. Every dot marks where the brand ranked in one answer. The brand and the question held still the whole time.

One question asked five times each morning for 19 mornings. Each dot is where the brand ranked in one answer; red crosses are answers that left the brand out entirely.

The scatter inside a single column often matches the drift across the whole three weeks. On some mornings the same question produced a #1 ranking and a #13 ranking minutes apart. The assistant samples a fresh answer every time, and the movement belongs to the sampler.

Always, never, and sometimes are three different problems

Whether a brand appears in the answer at all depends heavily on where it sits. For brands that almost always appear, five identical asks disagreed about presence on roughly one morning in eight. For brands on the bubble (a historical appearance rate between 20 and 80 percent), the five asks disagreed on most mornings for two of the three providers.

How often five identical asks disagree about whether the brand appears at all, for always-present brands versus brands on the bubble.

That splits the problem into three regimes.

Brands that almost always appear

Weekly checks captured everything the data offered: rank moves slowly here, and stable cells carried about a quarter of the answer noise that contested cells did, so they also need fewer repeat asks per check.

Brands on the bubble

This regime decides everything. Near a 50 percent appearance rate a single ask is a coin flip, and even at 60 percent, calling the brand in or out with real confidence from one sitting takes about 90 asks. The honest report for a bubble brand is a probability smoothed across checks; a green or red badge from one reading claims more than one reading knows. A presence swing also lives about six days, so bubble brands need a check every few days. A weekly rhythm lets a whole swing come and go between readings.

Brands that never appear

A brand at zero gives the same answer every time, so extra asks within a morning repeat what the first one said. A light probe on the weekly cycle covers it, and the probe has one job: catch the moment the brand first flips in. From then on it's a bubble brand and gets bubble treatment.

One more finding cuts across all three regimes: the bubble itself moves. Of 14 brand-and-question pairs we picked because their history put them on the bubble, about six were still there when the August run started, a few weeks after selection. Three had settled at never and two at always; one previously steady pair had fallen almost out of the answers. Brands drift between regimes within weeks. The checking rhythm has to adapt with them, and a fixed schedule falls behind.

Bunch your asks, or spread them out?

Both work. Which one works for a given channel depends on the provider, because each one's noise has a different shape.

What a burst of five same-morning asks is actually worth, measured per provider on the bubble brands.

On Google, five asks in one burst behave like five independent coin flips, and OpenAI points the same way on a shorter window. Bunching pays full value there: five asks are worth five asks, and averaging them cuts the noise exactly as fast as arithmetic promises.

Anthropic behaves differently. Its answers within a morning agree almost perfectly, but whole mornings occasionally flip together: a brand that appears four mornings in five will sometimes produce a morning of five straight misses, which independent chance would produce about eight times in ten thousand. Against that pattern the fifth same-morning ask buys almost nothing. A five-ask burst is worth about three and a half independent reads, while the same budget spread across five days keeps close to its face value.

So bunching and spreading are provider settings. Bubble brands always take the sequential half too, because classifying a brand that hovers around 50 percent is a smoothing problem across days rather than a precision problem within one sitting.

How much of one reading is real

How much of the movement in a single daily reading is real change rather than answer noise? The answer splits sharply by provider.

The share of one rank reading's day-to-day variation that is real movement rather than answer noise, by provider.

On Anthropic, roughly half of what a single reading shows is genuine. On Google it's closer to one part in seven, so a Google reading needs the most averaging before it means anything, and any confidence attached to a reading should reflect the channel it came from.

The practical rules fall straight out of the numbers. Check stable brands weekly and contested ones every few days. Bunch repeat asks on Google and OpenAI; spread them across days on Anthropic. Report a bubble brand as a probability. Weight every reading by its channel's reliability. This experiment is why Lantern reports read the way they do.

A note on honesty, because measurement earns trust by showing its discards: two early conclusions of this experiment failed our own review and were withdrawn. One claimed movement was mostly noise, so polling could slow down. One claimed day-to-day rank movement had strong momentum. Both collapsed under stricter controls, and every number in this article survived three separate estimation methods. The follow-up instrument keeps collecting, now across seven AI channels including the assistant products people actually use, so these numbers will keep getting sharper.

Common questions

Why does an AI assistant give different answers to the same question?

Each answer is a fresh generation from a probabilistic model, so the assistant samples a slightly different response every time. In our data the same question drew a #1 ranking and a #13 ranking minutes apart, with nothing changed in between.

How often should a brand check its AI visibility?

Weekly covers brands that always or never appear, because their state moves slowly. Brands that sometimes appear need a check every few days: their swings last about six days, and a weekly rhythm can miss one entirely.

How many times should the same question be asked?

On Google and OpenAI, ask several times in one sitting and average the answers; five asks there are worth five independent reads. On Anthropic, spread the same asks across days, because whole mornings succeed or fail together.

Can a single AI answer prove a brand is visible?

For a brand near a 50 percent appearance rate, one answer is a coin flip, and real confidence from one sitting would take about 90 asks. A probability smoothed across several checks tells the truth a single answer can't.

Loading Lantern…

Use your browser’s Reload button to try again.