How to Track Your Brand’s Visibility in ChatGPT
By Rahul A
By Rahul A

Track your brand in ChatGPT by testing fixed buyer prompts, recording mentions and citations, and comparing results across models and locations.
You can track brand visibility in ChatGPT by running the same buyer prompts regularly, saving the full answers, and measuring whether your brand is mentioned, recommended, cited, or omitted.
Brand visibility in ChatGPT isn't one score. Measure four separate outcomes: whether your brand is mentioned, whether it is recommended, whether the answer gives a reason to choose it, and whether the response links to a page you control.
That distinction matters because a brand can appear in an answer without being a serious option. A model might list your company among several providers, describe it incorrectly, or mention an outdated page. Record each outcome separately instead of marking the test as simply visible or invisible.
Use a small tracking sheet with columns for prompt, platform, model if shown, date, location, brand mention, recommendation, stated reason, linked source, competitor mentions, and factual errors. Add a final column for buyer usefulness. A response that names your company but gives the wrong service area is not a successful result.
Treat the sheet as a measurement record, not a prediction of every customer conversation. ChatGPT responses can vary with wording, conversation history, account settings, location, and whether web search is available. Your result shows what happened under a defined test condition.
For more context, read How to Improve Your AI Chatbot Visibility This Week.
Start with the questions a ready-to-buy customer would actually type, not prompts that contain your brand name. A useful starter set covers category discovery, local selection, problem-specific selection, comparison, and a follow-up asking for evidence.
For example, a freelance bookkeeper could test, “Who can help a small business in Manchester catch up on overdue accounts?”, “What should I look for in a bookkeeper for a growing consultancy?”, and “Compare local bookkeeping services for a business that needs monthly reporting.” Replace the details with your market, service, location, budget, and urgent problem.
Write every prompt down before testing. Do not quietly improve the wording after getting a weak answer, because that turns measurement into persuasion. Keep one version broad enough to reveal which brands the model retrieves naturally and another that reflects a real customer constraint.
Use the same prompt set each time, then add new prompts only when your offer, market, or customer questions change. The important test is not whether ChatGPT can mention you after being led to your name. It is whether a buyer can reach you from a plausible question.
For more context, read What ChatGPT Visibility Means for Your Brand Buyers.
Run a manual prompt panel in a fresh chat, using the same wording, account type, location, and browsing setting wherever possible. Open ChatGPT, paste one saved buyer prompt, and save the complete response with the date and visible model or search indicators.
Use a spreadsheet or a shared document for the record. Give each prompt its own row and each platform its own column, or store one row per prompt-platform-date combination. Copy the answer rather than recording only a yes or no. The surrounding wording reveals whether the mention was prominent, qualified, or wrong.
Repeat each important prompt in a new conversation. A previous conversation can teach the model your preferred answer and make a later result look stronger than an ordinary buyer would receive. Avoid switching between personalisation settings, logged-in accounts, and locations without recording the change.
A repeatable check does not require an API, automation platform, or technical setup. If you later automate collection with a tool such as Zapier, Make, or n8n, keep a manual sample as a control. Automated runs can preserve consistency, but they cannot make a changing model response objectively stable.
Record the answer’s evidence, not just the brand name. Capture the exact prompt, full response, date, model label, whether web search was used, cited URLs, and the position or prominence of your mention.
Then classify the result using consistent labels. “Absent” means no mention. “Listed” means your brand appears without a clear preference. “Recommended” means the answer gives a reason to choose you. “Cited” means the response links to a page that supports the claim. “Incorrect” means the answer includes a material error, such as the wrong location, service, audience, or availability.
Add competitor names and the claims attached to each company. A competitor appearing more often isn't automatically a problem if the prompt fits its offer better. The useful comparison is whether your company is present and accurately represented for the jobs you want.
Save screenshots when the interface displays citations or search context, because those elements can change after the answer is generated. A copied answer without its source links makes later checking difficult. Keep the raw record alongside your interpretation, so you can distinguish what the model said from what you think caused it.
Compare platforms by holding the buyer prompt and scoring rules constant, then treat each platform as a separate environment rather than combining all results into one visibility score. ChatGPT, Claude, and other assistants can use different models, search features, source handling, and account settings.
Paste the same prompt into a fresh conversation on each platform. Record whether browsing or citations are available, because a response generated without current web access is not directly comparable with one that searched the web. Keep the location and language consistent, and note any platform that does not expose a model or search state.
Compare patterns, not isolated answers. If your brand is recommended in ChatGPT but absent in Claude across repeated tests, that is a platform-specific finding. It may reflect different training data, retrieval sources, or ranking behavior rather than a universal weakness in your brand presence.
Official product documentation is the right place to check how search, citations, and model selection currently work. Those features change. Your tracking sheet should therefore include the platform configuration for every test, so a later change in results can be linked to a changed environment instead of blamed immediately on your website.
Separate retrieval failure from representation failure by checking whether the assistant can find your business and whether it describes the business correctly. These are different problems and need different evidence.
If your brand never appears for a relevant prompt, check whether independent pages describe your category, service area, audience, and use cases clearly. If your brand appears but the details are wrong, the tracking result is a representation problem. Record the exact false claim and the page or profile that should correct it.
If ChatGPT cites your website but ignores the relevant service page, the problem may be source selection or page clarity rather than brand absence. If it recommends you only after you name the company, the test shows recognition after prompting, not organic discovery. Label that result separately.
The common failure mode is treating every weak answer as a content task. Sometimes the prompt is outside your actual offer, the location is ambiguous, or the model has no reliable source to consult. Check the prompt fit, search context, and cited evidence before changing a page. A clean diagnosis prevents random edits based on one surprising response.
Trust an AI visibility score only when you can see the prompts, platform conditions, scoring rules, and underlying answers behind it. A single percentage or rank hides the difference between a passing mention and an accurate recommendation.
A useful score is an internal trend indicator. For example, you might assign separate points for being mentioned, recommended, accurately described, and supported by a relevant citation. Keep the rules fixed, document them in the sheet, and report the component results beside the combined score.
Do not compare scores from two tools unless they test the same platforms, prompts, locations, model settings, and definitions. One tool may count any mention, while another may count only an uncued recommendation. Neither score is automatically wrong, but they answer different questions.
Treat results as directional because model outputs are probabilistic and interfaces change. A score that rises after one run may reflect prompt variation or a temporary source choice. Look for a pattern across repeated checks and inspect the raw answers before taking action. The score is a filing label for evidence, not proof that a specific number of buyers will find you.
Choose one corrective action from the most repeatable failure, then rerun the same prompt panel after the change. If the model cannot identify what you do, clarify the category, audience, locations, and services on your core pages. If it gives outdated facts, correct those facts at the source and on relevant business profiles.
If your brand is mentioned but not recommended, compare the reasons given for competing providers. Look for a missing, supportable detail such as a clear service boundary, suitable customer type, process explanation, or proof of a specific capability. Do not add claims merely because a model seems to prefer them.
If citations point to irrelevant pages, improve the page that should answer the buyer’s question and make its title, headings, and supporting details unambiguous. If the answer repeatedly invents information, log the error and avoid treating the model as a dependable source for that claim.
Rerun the unchanged prompts and record the difference. Keep the old answers. A before-and-after comparison is more useful than a new result with no baseline. The weekly habit is simple: test, preserve evidence, diagnose one failure, make one justified change, and test again.
That’s the whole lesson. Try it on a real task while it is fresh, then come back for the next one.
The same corner of the library, one job further on.
Drafted with AI assistance from our own research and Search Console data, and reviewed by Rahul A before publishing. Tools and prices change; check the linked official source before you act.