Asking ChatGPT for your own company or product name to see whether it comes back cannot, on its own, be used as a basis for a decision. An AI answer changes every time, even for the same question. Some days you appear and some days you don't, so what should you look at to judge your own visibility? This article lays out why the check can't be trusted, and what to judge with instead.
Contents
TL;DR#
- Asking ChatGPT whether your name comes back returns a different answer each time, so a single result decides nothing
- What moves is not only whether you appear. The content itself can be at odds with the facts, down to citing a reputation that does not exist
- Research reports that answers still fail to line up under settings meant to fix the output. Accuracy varied by up to 15%, and on the same problems the spread between hits and misses reached up to 70% (both are upper bounds)
- Asking many times and counting how often you came back is an aid for getting a rough feel, no further. The survey effort doesn't pay for itself, and what it measures is the AI's impression, which is not tied to traffic to your own site
- What you judge with is not the answer at the moment you ask, but the record that accumulates on your own site: whether someone arrived via AI, which page they landed on, whether they bought
1. What You Get Wrong When You Judge From One Ask#
Judging from one result gets you wrong in both directions, equally — the reassured direction and the alarmed one. Yesterday your name came back, today it doesn't. Or the reverse: nothing for a long stretch, then suddenly it's there. Anyone who has started working on GEO has usually run into it at least once.
You get it wrong on the days you appear and on the days you don't#
When you do appear, you stop checking there. What actually happened may be that this one run happened to put you among the candidates. Send the same question the following week and you're gone. That is not unusual.
When you don't appear, the problem is larger. Not being picked once is enough to start an investment decision — rewriting articles, adding structured data. The moves themselves may be right, but the observation that gave you a reason to start isn't dependable. And since you check the effect with the same method, if you appear after making the fix, you can't separate whether that was the result of the fix or just the swing.

What swings is not only whether you appear#
There is one more thing that gets missed. What changes is not only whether you appear, but the content of the text that comes back.
An AI can cite a reputation that does not exist as its grounds. You ask about your own site, get judged on the basis of bad word of mouth, and following the source leads to a page that has nothing to do with you. It is an extreme case, and it shows plainly that the text returned by your check is one output assembled on the spot, not a fixed record.
The full picture of how to measure is in measuring AI brand visibility, and where a mid-market store actually stands. The problem that sits before it is this one: why the single check almost everyone starts with can't be used as a basis for a decision.
2. Why the Answer Changes Every Time#
Send the same question twice in a row and the text that comes back is a different thing. That is because an AI answer is assembled from scratch each time, and it is not a defect.
It picks words one at a time by probability#
An AI is not pulling a finished piece of text out of somewhere. It produces candidates for "the word likely to come next," then picks one of them. It repeats that to the end to compose the text. If there is even a little width in how candidates get picked, the opening phrase changes, and everything that unfolds from there changes with it. In a passage recommending online stores, another company's name lands in the position your name was going to occupy. Given the mechanism, that is an ordinary thing to happen. With a service like ChatGPT that searches the web before answering[2], the pages it refers to also differ from occasion to occasion.
Even settings meant to fix the output don't line up#
It is tempting to think that settings which suppress the swing will return the same answer. This has actually been measured in research. In a study that put five models through eight kinds of task ten times each, under settings that were supposed to fix the output, the answers still did not line up[1]. Accuracy varied by up to 15%. And between counting every problem that was answered correctly at least once in ten runs as correct, and counting every problem that was answered incorrectly at least once as incorrect, the gap reached up to 70%.
These two figures are upper bounds, not averages and not medians. The width of the variation differs greatly by task, and most of it lands smaller than this. What matters is less the size of the figures than the fact that even settings that were supposed to fix the output could not fix it completely.

That output swings is well known among people who work with AI daily. What is not known is how far that swing invalidates the everyday task of "checking whether we come up." The conditions that make an article likely to be cited are covered in the two conditions that get you cited by AI.
3. What to Judge Your Visibility By Instead#
Put the ground for the decision on the record that accumulates on your own site, not on the answer at the moment you ask. The former gets rewritten every time; the latter doesn't change after the fact.
The answer of the moment, and the record that accumulates#
An answer returned by asking an AI is one frame of that moment. Since the next answer to the same question is a different thing, a month-scale trend does not come out of one answer.
What stays on your own site is a continuous record. Someone followed a link out of an AI answer and arrived, landed on a particular page, read and bounced, or went through to a purchase. These pile up with dates on them and don't get rewritten later. Last week's visits do not stop having happened. This is what you can make the starting point of a decision.
How far does asking many times get you#
So what about asking many times and counting how often you came back? There is nothing wrong with the method itself. Some practitioners check their footing with it: it was 2 out of 10, now it's 8. But what you get is an aid for grasping a rough tendency, no further, and there are two limits.
One is the effort. Running a meaningful number of times for one search keyword, then doing that for as many keywords as matter, every month. That survey design is what specialist research firms put together to produce trends for a whole market. For a single company that only wants to know where it stands, the cost doesn't balance out (for the exposure level of a whole market, a large-scale survey can measure it).
The other is more fundamental. However many times you run it, what you're measuring is the AI's impression, and it isn't tied to traffic to your own site. Your name can be in the answer while AI-referred visits this month are zero. The reverse happens too: you never see yourself in the answers, yet the traffic arrives. Increasing the number of runs does not connect the two.

What to do about appearing is not covered here#
This article does not step into what to do to get yourself to appear. Google states that there are no special requirements for appearing in AI features[3]. Moves such as structured data and how articles are written are collected in how to show up in ChatGPT and build the base that gets you cited. The only point here is that neither the decision to pick those moves nor the check on whether they worked should be loaded onto one ask. How to read articles that get cited without getting clicks is covered in the numbers to look at before deleting a zero-click page.
RevenueScope solution
Re-aggregating in the same shape every month while keeping AI-by-AI referral detection and bot exclusion intact does not survive as manual work. RevenueScope stores AI-referred traffic as a continuous record. It classifies which of ChatGPT, Claude, Perplexity, Gemini or Copilot a session came through, and returns the page it landed on along with the revenue that stood up on that page (landing revenue = attributed to the entry page). It also returns the date that page last received an AI-referred visit. What you judge with becomes the record left on your own site, not an answer that gets rewritten every time you check. Visits where the AI assistant passes no referrer are not counted. That is a floor common to everyone doing the measuring, and on top of that floor it counts the traffic that actually arrived.
Asking an AI assistant through MCP returns this#
Ask "break out AI-referred traffic by AI assistant" and it comes back in this shape:
| AI assistant | Session share | Landing revenue share | Last referred |
|---|---|---|---|
| ChatGPT | 28.3% | 0% | 2026-07-12 |
| Perplexity | 26.4% | 0% | 2026-07-14 |
| Gemini | 18.9% | 48.2% | 2026-07-10 |
| Claude | 17.0% | 38.6% | 2026-07-14 |
| Copilot | 9.4% | 13.2% | 2026-07-12 |
One example (illustrative, past 30 days)
The first thing to read in this example is the last-referred column on the right. With a date in it, you can separate traffic that came until last month and stopped from traffic that never came at all. One check tells you "we didn't appear today" and no further; with a record, you can go back afterwards and find when it changed. Line it up against the day you fixed an article and the effect of the move can be confirmed here too.
That the AI carrying the most sessions and the AI tied to revenue are not the same one is not unusual in itself. That phenomenon is handled with measured data from the sample store in measuring AI visibility for a mid-market store.
FAQ#
Frequently asked questions#
Q. Is a different answer every time a sign that I'm asking badly?
A. It isn't a problem with how you ask. Send the same wording on the same day and the text that comes back changes. An answer is not pulled from something stored somewhere; it is assembled by picking one word at a time, freshly each time. Making the question specific will line up the direction of the answer to a degree, but it does not go as far as the same result every time.
Q. If I use settings that fix the output, will the same answer come back?
A. Not completely fixed. Even in research that measured it under settings that were supposed to fix the output, the answers did not line up[1]. On top of that, in consumer services such as ChatGPT, users cannot touch those settings. If the answer involves search, the pages it refers to also differ each time.
Q. What should I do if an AI writes something about us that is at odds with the facts?
A. First, separate whether it comes back every time or only came back that once. If it was once, there is a good chance it will be gone on the next generation. If the same error comes back however often you ask, the cause is often on the page being referred to, and stating your own site's description correctly and explicitly is more effective.
Q. We appeared yesterday and not today. Does that mean our ranking dropped?
A. You can't judge that from single results. AI answers have no fixed ranking like search results, and whether you appear is decided on the spot each time. To see whether something dropped, line up how much AI-referred traffic arrived by date and check whether the level itself fell. A single day of appearing or not appearing does not translate into a move up or down in rank.
Summary#
The procedure for judging visibility comes down to three things. First, take the task of asking ChatGPT whether your name comes back out of the judgment. The answer is picked fresh every time, so there is nothing to read from that one run, on the days you appear or the days you don't. What swings is not only whether you appear but the content that comes back. That answers still fail to line up under settings meant to fix the output has been measured in research[1].
Second, settle where counting belongs. How many times out of ten you came back works as a rough gauge. But the survey effort doesn't pay for itself, and what it measures is the AI's impression, which is not tied to traffic to your own site. Keep it as an aid and don't use it for the judgment.
Third, make the judgment on the record of AI-referred traffic. How many sessions arrived from each AI, which page they landed on, whether revenue stood up from there, when the last referral was. With those four lined up with dates, you can confirm afterwards what moved before and after you fixed an article. One check leaves none of it.
See which ads actually drive revenue, at a glance
Free up to 5,000 sessions/month, AI analyst included. No credit card required. Up and running in 5 minutes.
![[Research] Does ChatGPT Name Your Brand? The Answer Changes Every Time](/_next/image?url=%2Fimages%2Fnews%2Fchatgpt-brand-check-single-try.jpg&w=3840&q=75)


![[Research] The Two Conditions That Get You Cited by AI: What a 250K-Trial Study Shows](/_next/image?url=%2Fimages%2Fnews%2Fai-citation-conditions-study.jpg&w=3840&q=75)