# Lantern: full article text > Complete text of every Lantern article. The index with links is at https://lantern.is/llms.txt. --- # Sampling on the bubble: how often to check AI answers Source: https://lantern.is/resources/sampling-on-the-bubble Author: Eric Published: 2026-09-05 Two months of asking AI assistants the same shopping questions five times per morning, and what that noise says about how often to check a brand's AI visibility and whether to bunch the checks or spread them across days. Ask an AI assistant the same question five times and you can get five different answers. We spent two months measuring that noise, and it settles the two questions every visibility program faces: how often to check, and whether to bunch the checks or spread them out. Two findings carry everything below. First, where a brand sits decides how hard it is to measure: brands that almost always appear in AI answers disagreed with themselves on roughly one morning in eight, while brands on the bubble disagreed on up to seven mornings in ten. Second, bunching or spreading your repeat asks is a provider setting: on Google and OpenAI a burst of five asks is worth five independent reads, while on Anthropic whole mornings flip together and the same burst is worth about three and a half. Here's where those numbers come from. In July 2026 we asked three assistants (OpenAI, Anthropic, Google) the same shopping questions five times within a few minutes, every morning, for three weeks: 1,357 scored answers. In August we repeated the run on brands picked because they sometimes appear in AI answers and sometimes don't. [Why one AI answer is never enough](/resources/why-one-ai-answer-is-never-enough) explains the principle those runs confirmed: an AI answer is a sample. This is the data. ## What answer noise actually looks like One real cell of that data: one question, one brand, OpenAI, five asks per morning for three weeks. The brand and the question held still the whole time. Each column is one morning. Each dot marks where the brand ranked in one answer; red crosses are answers that left the brand out entirely. The scatter inside a single column often matches the drift across the whole three weeks: same question, same morning, a #1 and a #13 minutes apart. The movement belongs to the sampler. An assistant generates a fresh answer every time, and brands near the edge of an answer ride that randomness in and out. Any single reading is mostly noise, which is why a reading belongs in a trend, never in a headline. ## Always, never, and sometimes are three different problems The July pass ran on brands that appear in 96 to 98 percent of answers. The August pass deliberately seeded brands picked at a historical appearance rate of 20 to 80 percent. Same instrument, very different behavior: On always-present brands, presence is nearly settled: 8 to 19 percent of mornings split. On the bubble, disagreement dominates for Google (71 percent) and OpenAI (67 percent). OpenAI's bubble numbers come from 15 mornings before its pinned model retired, so treat them as directional. That splits measurement into three regimes, and each one gets its own playbook: | Where the brand sits | How often to check | Reads per check | Report as | | --- | --- | --- | --- | | Always shows up (85%+) | Weekly; rank moves slowly and presence is settled | Light; noise is a quarter of contested brands' | A smoothed rank trend | | On the bubble (20 to 80%) | Every 3 to 5 days; a swing lives about 6 days | 3+ on Google and OpenAI; spread across days on Anthropic | A probability smoothed across checks, never a one-read badge | | Never shows up (near zero) | A light probe on the weekly cycle | Minimal; every ask repeats the first | The moment it first flips in, then treat as on the bubble | The check intervals come from run-length analysis of how long each state lasts; the repeat experiment adds the read counts and the reporting rules. The bubble row deserves its numbers spelled out. Near a 50 percent appearance rate a single ask is a coin flip, and even at 60 percent, calling the brand in or out with real confidence from one sitting takes about 90 asks. The honest report for a bubble brand is a probability smoothed across checks; a green or red badge from one reading claims more than one reading knows. And a presence swing lives about six days, so a weekly rhythm lets a whole swing come and go between readings. One more finding cuts across all three regimes: the bubble itself moves. Of 14 brand-and-question pairs we picked because their history put them on the bubble, about six were still there when the August run started, a few weeks after selection. Three had settled at never and two at always; one previously steady pair had fallen almost out of the answers. The bubble is a place brands pass through, and the checking rhythm has to follow them. ## Bunch your asks, or spread them out? Both work. Which one works for a given channel depends on the provider, because each one's noise has a different shape. Grey is the five asks you paid for; color is what they're worth as independent evidence. Google is cleanly binomial across 56 mornings, and OpenAI points the same way on a shorter window. Anthropic's bursts overlap because whole mornings succeed or fail together. On Google and OpenAI, bunching pays full value: five asks are worth five asks, and averaging them cuts the noise exactly as fast as arithmetic promises. Anthropic behaves differently. Its answers within a morning agree almost perfectly, but whole mornings occasionally flip together: a brand that appears four mornings in five will sometimes produce a morning of five straight misses, which independent chance would produce about eight times in ten thousand. Against that pattern the fifth same-morning ask buys almost nothing, and the same budget spread across five days keeps close to its face value. So bunching and spreading are provider settings. Bubble brands always take the sequential half too, because classifying a brand that hovers around 50 percent is a smoothing problem across days rather than a precision problem within one sitting. ## Spend your reads where the answer is contested Noise concentrates exactly where readings carry decisions. Identical re-asks spread about four times wider on contested brands than on stable ones, for rank and sentiment both. A stable #1 needs a light touch, and a contested brand deserves every repeat ask in the budget. The same read spent on a settled answer buys confirmation; spent on a contested one, it buys information. ## How much of one reading is real Averaged across the window, how much of the movement in a single daily reading is real change rather than answer noise? The answer splits sharply by provider. Whiskers show confidence ranges from a cell-level bootstrap. Anthropic readings carry the most signal per read; Google readings need the most averaging before they mean anything. On Anthropic, roughly half of what a single reading shows is genuine. On Google it's closer to one part in seven. Any confidence attached to a reading should reflect the channel it came from. The practical rules fall straight out of the numbers. Check stable brands weekly and contested ones every few days. Bunch repeat asks on Google and OpenAI; spread them across days on Anthropic. Report a bubble brand as a probability. Weight every reading by its channel's reliability. This experiment is why Lantern reports read the way they do. A note on honesty, because measurement earns trust by showing its discards: two early conclusions of this experiment failed our own review and were withdrawn. One claimed movement was mostly noise, so polling could slow down. One claimed day-to-day rank movement had strong momentum. Both collapsed under stricter controls, and every number in this article survived three separate estimation methods. The follow-up instrument keeps collecting, now across seven AI channels including the assistant products people actually use, so these numbers will keep getting sharper. ## Common questions Why does an AI assistant give different answers to the same question? Each answer is a fresh generation from a probabilistic model, so the assistant samples a slightly different response every time. In our data the same question drew a #1 ranking and a #13 ranking minutes apart, with nothing changed in between. How often should a brand check its AI visibility? Weekly covers brands that always or never appear, because their state moves slowly. Brands that sometimes appear need a check every few days: their swings last about six days, and a weekly rhythm can miss one entirely. How many times should the same question be asked? On Google and OpenAI, ask several times in one sitting and average the answers; five asks there are worth five independent reads. On Anthropic, spread the same asks across days, because whole mornings succeed or fail together. Can a single AI answer prove a brand is visible? For a brand near a 50 percent appearance rate, one answer is a coin flip, and real confidence from one sitting would take about 90 asks. A probability smoothed across several checks tells the truth a single answer can't. --- # AI answers change when nothing else does Source: https://lantern.is/resources/why-one-ai-answer-is-never-enough Author: Eric Published: 2026-09-02 Ask an AI assistant the same question twice and the list of brands moves. What that means for measuring AI visibility. Ask ChatGPT for "the best running shoes for flat feet". Write down the brands it names. Ask again five minutes later. The list moves. Nothing changed in those five minutes. Same model, same question, same morning. These assistants write a fresh answer every time, so a brand sitting near the edge of relevance drops in and out on its own. ## A single reading is a sample A rank tracker reads a search results page that mostly sits still. Check it Tuesday, check it Friday, and the page keeps its shape. An AI assistant writes its answer new each time. We asked the same shopping questions several times inside one morning and compared what came back. Brands that almost always appear still didn't appear every time. Brands sitting on the edge came back closer to a coin flip. Treat one of those readings as the answer and you'll report that a brand doubled its AI visibility on a morning when nothing happened. ## Each assistant is noisy in its own way On some assistants the answers vary inside the moment. Ask five times before lunch and you'll see a spread, so asking repeatedly and averaging settles it. Others behave differently. Their answers inside one morning agree almost perfectly, then the whole day moves together. The assistant is consistent, and tomorrow it is consistently something else. Repeating the question five times in a row teaches you nothing there. You need readings spread across days. Sample every assistant on one schedule and you'll get both cases wrong. ## The contested questions keep moving The sharpest finding came from watching the same questions over weeks. Queries we picked because they were contested had often settled by the time we looked again. The brand was either in the answer for good or out of it for good. Questions that had held steady had drifted into contention. So the thing you measure won't hold still either. A quarterly snapshot of AI visibility describes a set of answers that already moved. ## What to do about it Ask more than once. Report a brand that appears half the time as a brand that appears half the time. Watch the questions that move more closely than the ones that sit still. Measure it naively and you'll report movement that never happened. Lantern watches these answers every day and reports them as tendencies. Ask twice. The second answer is the one that teaches you something. ## Common questions ### Why do AI assistants give different answers to the same question? They write each answer fresh instead of reading a stored ranking. Small differences inside that process change which brands reach the list, and brands near the edge of relevance move the most. ### How many times should I ask before I trust the result? More than once, and the right number depends on the assistant. Some vary inside a single morning, so readings taken close together help. Others hold steady all day and then shift, so readings spread across days matter more. ### My brand appears in half the answers. Did my AI visibility drop? No. A brand near the edge of an answer set moves in and out on its own, so the honest reading is that it appears about half the time. Confirm a real drop across many readings before you act on it. ### How often should I check? Often enough to see the questions move. The contested set changes over weeks, so a quarterly check reports answers that already changed. Watch the questions that decide purchases in your category more closely than the rest. --- # How ChatGPT decides which brands to recommend Source: https://lantern.is/resources/how-chatgpt-decides-which-brands-to-recommend Author: Andrew Lissimore Published: 2026-07-02 An AI names a brand for one of two reasons: it remembered you, or it looked you up. Not long after ChatGPT launched, I asked it which headphones to buy. I had spent ten years building [Headphones.com](https://headphones.com) into one of the most trusted names in audio, and the model never brought us up. It had no web access back then, nothing to look up. The answer came straight out of whatever impression of the headphone market it had absorbed during training, and a decade of winning Google searches had left barely a trace there. On July 1, [Business Insider profiled Lantern](https://www.businessinsider.com/why-ecommerce-startup-lantern-pivoted-geo-agentic-shopping-marketing-pitch-2026-7), the company that moment eventually turned into. Sydney Bradley did it right: hard questions about a crowded market, no softballs, and a story that came out fair. The line of mine that survived is that this is the most important transition since search. I stand by it. What never fits in an interview is the mechanics, and the mechanics are the part another operator can actually use. An AI names a brand for one of two reasons. It remembered you, or it looked you up. Everything I have learned about getting recommended sorts into those two channels, and they reward different work. Audio taught me the first channel before AI existed. The forums decided who was credible in headphones years before any model did: Head-Fi threads, subreddit arguments, the handful of reviewers people actually trusted. Headphones.com earned its name there, thread by thread, return policy by return policy. The models have since read all of it. A model's sense of who is credible in a category forms from exactly that public record: reviews, roundups, forum fights, press. When enough of it treats a brand as a serious answer to a real problem, the association sticks. There is no placement to buy in this channel and no tag to install. In fact, the only way in is to be worth writing about, which is either the most encouraging thing about this whole shift or the most brutal, depending on the week you're having. Retrieval is the second channel, and the more forgiving one. Ask ChatGPT which headphones to buy today and it no longer answers from memory alone: it will often run a search mid-answer, pull a handful of live pages, and ground its recommendation in what it just read. Perplexity does this on every query. The traffic that follows is real. Adobe measured a 1,200% rise in visits to US retail sites from generative AI sources between July 2024 and February 2025. So the practical question becomes what the model finds when it fetches you, and whether it can fetch you at all. Cloudflare ships a one-click setting that blocks AI crawlers wholesale, and plenty of stores run with it on without anyone having decided to. Others render product pages in JavaScript the crawlers never execute. A store in that condition can rank well on Google and still be unreadable to the systems doing the recommending. The cheapest thing to check, and the first thing most teams skip. Then the fetched page has to say something a machine can use. Audio companies name products like perfumes. Clear, Elegia, Atrium, Caldera: gorgeous names that tell a machine nothing. (Somebody fought hard for those names, and I sympathize, but the machine does not.) A listing titled "Focal Elegia: closed-back audiophile headphones, 35 ohm, quiet enough for a shared office" gives the model every attribute it needs to match the person asking what they can wear at a desk without bothering anyone. Multiply that difference across every title, spec sheet, and price in a catalog and you have most of what separates brands that appear in answers from brands that get skipped. Press, it turns out, lands in both channels at once. A national story is retrievable the day it runs, sitting in exactly the set of pages a model pulls when it wants to know who is credible in a category. And it joins the public record the next generation of models absorbs. That, more than any traffic spike, is why coverage matters now. The buying side is being wired up in parallel. OpenAI and Stripe shipped the Agentic Commerce Protocol in September 2025, and Google and Shopify followed with the Universal Commerce Protocol in January 2026, so an agent can increasingly check stock, confirm a price, and finish checkout without a human touching the storefront. Around 60% of searches already end without a click on any result. Where does that leave the storefront? Increasingly, the answer is the storefront. Want to know where you stand? The test costs nothing. Take the five questions that decide purchases in your category, the ones with a budget and a use case in them. In my world that was "best closed-back headphones under $500 for the office." Ask ChatGPT, Gemini, and Claude, write down which brands come back, and do it again next week. Most operators have never looked at this list once. It is the scoreboard now. Lantern is the company that 2022 headphone question turned into. At Headphones.com I could check every Google ranking that mattered any morning I cared to; when the models started answering instead, there was nothing to look at. So we built what I needed that morning: Lantern watches how ChatGPT, Gemini, and Claude answer the buying questions in your category, shows where you stand against the brands they name instead of you, and ranks the changes that actually move the answer. Where all of this settles, nobody knows. I suspect the agents end up doing the buying outright and the storefront becomes plumbing, though anyone who claims to know the timeline is selling something. Possibly including me. What is already true is enough: twenty years of SEO taught brands to win a ranking. The next twenty are about winning a sentence. ## Common questions ### Do AI assistants recommend products from training data or live web data? Both, through two separate channels. Part of the answer comes from what the model learned about brands during training, and part comes from pages it retrieves while answering; ChatGPT, Gemini, and Perplexity all ground shopping answers with live search now. A brand can be strong in one channel and invisible in the other. ### Why doesn't my brand show up in ChatGPT when it ranks well on Google? A Google ranking measures one retrieval system, and AI assistants run their own. If AI crawlers are blocked from your site, if product pages only render through JavaScript, or if product data never states plainly what you sell, the model has nothing to work with when it fetches. ### What is generative engine optimization (GEO)? Generative engine optimization is the practice of making a brand more likely to be cited and recommended in AI-generated answers. Answer engine optimization (AEO) describes the same goal. Both come down to the two channels models actually use: the reputation absorbed during training and the pages retrieved at answer time. ### How long does it take to improve AI visibility? Retrieval-side fixes, meaning crawler access and product data, can show up in answers within weeks. The reputation channel builds the way reputations always have, review by review and story by story, and the models keep reading as it accumulates. Nobody can date that precisely, which is a good reason to watch the answers weekly instead of guessing. --- # Google Now Measures Your AI Search Visibility Source: https://lantern.is/resources/google-search-console-ai-reports Author: Eric Published: 2026-06-26 Google Now measures AI Search Visibility On June 3, Google shipped something brands have wanted since AI Overviews launched. Search Console now has dedicated reports for AI search. You can see how often your pages appear inside Google's generative AI features: the AI Overviews that sit above the old blue links, and AI Mode, the conversational tab beside them. Google is starting with a small set of properties, so the reports may already be waiting in your Search Console. For two years, brands watched AI reshape their search traffic with one hand tied behind their back. A page would shift and nobody could say how much of the change was AI answering questions directly. Google just handed everyone an instrument. Use it. The report covers impressions in Google's AI features across Search and Discover, broken down by page, country, and date, down to the hour, with device splits for Search. In plain terms, it tells you how often your pages show up when Google's AI answers a query. That is a real number, and one brands could only guess at before. Read an impression as the first signal, the moment your page enters an AI answer. The question worth asking next is what happens after that appearance. For an ecommerce brand, the number that moves revenue is whether the agent recommends your product by name and whether a shopper can buy it from there. Google's report starts that conversation. Your job is to follow it all the way to the sale. Most AI shopping happens across more than one company's surfaces. A buyer asks ChatGPT for the best option in their size, shops inside Gemini, asks Claude to compare two products, and checks Google along the way. Complete AI visibility means knowing how you show up in all of those answers, and whether the agent recommends you and can transact. Checking that takes nothing but the questions themselves. Ask ChatGPT, Gemini, and Claude the decision queries your customers ask and count who gets named. Google's new report now keeps watch over its own surfaces daily. Checking the rest by hand works until the query list grows, and keeping that watch running is the work tools like Lantern exist for. Google opening up AI reporting tells you where this is heading. The largest search company in the world now treats AI visibility as a number worth reporting on. The brands that win will measure it everywhere their customers are asking. ## Common questions ### What did Google announce? On June 3, 2026, Google launched Search Generative AI performance reports in Search Console. They show how often your pages appear in Google's AI features, including AI Overviews and AI Mode in Search and generative features in Discover, broken down by page, country, and date, with device data for Search. The rollout starts with a subset of properties. ### What does the report measure? Impressions: how often your pages appear in Google's AI answers, broken down by page, country, and date down to the hour, plus device for Search. Google says more metrics will follow. ### Does it cover ChatGPT, Gemini, and Claude? It covers Google's own AI surfaces. Visibility in the other agents, where a large and growing share of AI shopping happens, has to be measured by querying them directly or using a tracker that does it daily. ### How do I measure my visibility across every AI agent? Run your category's decision queries through each agent and record which brands get named. AI visibility trackers automate that watching across providers, queries, and time.