# How to Check What AI Says About Your Business (Free Test Sheet) By Maria Marnovora You can start this afternoon, with a browser and about two hours for the short version below, or a long afternoon for the full set. The mistake is doing it once, screenshotting whatever came back, and treating that as the answer. The short version: ask the same questions twice in two tools, write down what came back and which pages it cited, and repeat it on a cycle. One answer tells you very little, because the answer moves. Why one AI answer proves nothing One answer from one assistant is a sample of one. In September 2026 we asked two questions, twice in ChatGPT and once in Google's AI Mode: who the best digital marketing agencies in Dubai are for a small business, and which agencies help brands show up in AI search results. Six answers, one account, one evening, from a Dubai connection. The ChatGPT runs were temporary chats with personalisation off and thinking effort set to high; the Google runs were made while signed in. Those are different settings, and worth knowing before you set yours against ours. Far too small to be a study. We logged what the assistants said and checked none of it, and we name no company here. The first question produced fourteen different companies across three answers. Each answer named five. Only one name appeared more than once. The two ChatGPT runs were minutes apart, in the same mode, on the same account, with the question typed identically, and they shared one name out of five. The Google answer named five more, none of which either ChatGPT run had mentioned. The sources moved as much as the names. One ChatGPT run cited two company websites and a business directory. The next cited the directory alone. The Google answer cited five company websites and no directory at all. On the second question, Google's answer opened on a forum thread and a video, neither of which appeared in either ChatGPT run. What was said about each company came from something already published. A package price range in dirhams. A total number of UAE clients served, lifted from the company's own site and passed on as the company's claim. A directory profile's review count, star rating, minimum project size, hourly band and service-mix percentages. A published case study. A professional profile. An office address down to the floor and office number. Not all of it came from the company itself: reviews are written by reviewers, directory profiles are compiled by the directory, and one answer opened on a forum thread. But all of it was already public somewhere. All four ChatGPT runs closed with a buyer's checklist nobody had asked for. One added a caution that a company's published performance figures were self-reported, and that a buyer should ask to see client-level evidence during a pitch. Both Google answers ended instead with a comparison table and an offer to narrow the list. Screenshot a single answer and send it to your team, and you have circulated one draw from a pack nobody has counted. How do you set the check up so it is repeatable? Seven settings decide whether your next check is comparable to this one, and all of them are fixed before you type. Use a temporary chat, with personalisation off. OpenAI says that "before starting a temporary chat, you can choose whether it uses existing memories, custom instructions, and plugins", and that temporary chats "do not create or update memories, even with personalization on" (OpenAI Help Center (https://help.openai.com/en/articles/8590148-memory-faq)). In a tool with no temporary mode, use a private window and stay signed out. Pick one of the two, note it in your log, and use the same one every cycle. Write down the location the tool is using. ChatGPT "may use an approximate location based on your IP address to provide relevant local results" (OpenAI Help Center (https://help.openai.com/en/articles/9237897-chatgpt-search)). You will not always be shown it. On a Google results page it is usually printed at the bottom, though that moves. In ChatGPT, ask in a second temporary chat rather than the one you are testing, because any extra message makes the run something other than the one you meant to test. If you cannot establish it, write down where you were and which connection you used, and use the same desk next cycle. Use at least two tools. On our first question the two tools shared no names at all; on the second, one company appeared in every answer. ChatGPT and Google's AI Mode are the pair we used. Perplexity, Gemini, Copilot and Google's AI Overviews are worth adding once the first cycle is done, and the sheet lists them. Adding a tool next cycle is fine; swapping one out halfway is not. Record the tier and the settings. Note the account tier and whichever model or effort setting the tool shows you, then stay on it next cycle. Ours were ChatGPT temporary chats at high thinking effort and Google AI Mode while signed in. A different setting can produce a different answer, so a check run on one is not comparable with a check run on another. Run every question twice. Not a word different the second time. If the assistant asks you a question back, do not answer it. A question back, about your budget or which emirate you mean, is itself a result. Log that it asked, start a fresh chat, and type the question again unchanged. Answering turns your run into a conversation nobody else can repeat. Open the citations and write down the domains, because the pages behind an answer are the part you can act on. OpenAI's guidance for ChatGPT is that it "ranks search results using multiple factors intended to help users find relevant, reliable information", with "Placement is not guaranteed", and that "Search results and citations can be incomplete, outdated, or incorrect" (OpenAI Help Center (https://help.openai.com/en/articles/9237897-chatgpt-search)). Which questions should you ask? Ask eight questions, in four groups. All eight are written out in the test sheet, with brackets to fill in. Category. "Who are the best [your category] in [your city]?" and "Which [your category] is best for [the customer type you serve]?" The first is asked before anyone knows a name. The second is usually where a smaller company can appear. Your brand. "What is [your company]?", "Is [your company] any good, and what do people say about it?", and "Where is [your company], what does it cost, and how do I contact them?" The first shows what the assistant thinks you are. The second shows who is speaking for you when you are not in the room. The third is the fastest way to find a wrong fact. Comparison. "[Your company] vs [a competitor]: which should I choose?" This shows the criteria you are being judged on, which are rarely the ones you lead with. The buying question. One question your customers actually ask before they buy, and one local phrasing, your category plus a district you serve. Between them they test whether your pages answer anything and whether your listing is right. Two tools, eight questions, two runs each: thirty-two runs, and with the logging that is a long afternoon. If that is more than you have, run the three brand questions and the first category question, in both tools, twice each. That is sixteen runs, and you can add the rest next cycle. What you must not do is run six questions this quarter, eight next, and compare the two numbers. Lunasol, which owns this publication, starts its own audit from the same place, asking "every AI your customers use" the questions "your customers actually ask" (Lunasol (https://lunasol.ae/geo-aeo)). What should you record for each run? Record the run itself: the date, the cycle, the assistant, the tier and settings, the location the tool showed, and the question as typed. Then the result: whether you were mentioned, where in the list, what it said about you, which other companies were named, whether any fact was wrong, and every domain it cited. Record other companies by name only, exactly as the assistant returned them. Leave the rest out: what the assistant judged about them, what it said their prices or addresses were, anything about a named individual. Do not republish any of it either. An assistant's claim about another business is unverified, and repeating it is your risk rather than the tool's. A minute or two a run, most of it spent opening the citations, and it is the difference between a folder of screenshots and something you can compare in three months. Record the failures too. An answer that names nobody is a result, and so is a run that returns nothing at all. All six of our runs came back with an answer, but a run that stalls is still worth a row. How do you read the results? Read the log on four axes: presence, accuracy, sources and variation. Presence. How often you appear at all, out of your own runs. One in ten and six in ten are different situations. It is not market share, and it is not a ranking. Do not let anyone present it as either. If you appear in none of them, treat that as a starting point rather than a verdict: our six runs cannot tell you whether a zero is usual. Read the sources column instead: the pages the assistants did use are the ones you now have to appear on. Accuracy. What is said about you, and whether it is true. A wrong price, an old address or a service you dropped is worth more of your attention than a missing mention, because it reaches a customer who was already interested. Sources. Every answer we logged was assembled out of pages somebody had published. In our runs those were company websites, a business directory, two agency directories, a forum thread, a video and a professional profile. If your answers are built from pages you do not control, that is the work. Variation. In our six answers the movement was much larger than a name or two: the same question, minutes apart in the same tool, came back with four of the five names changed. What counts as normal is something only your own log can tell you. A change in the type of source tells you the answer is not settled. We did not test whether a change on your side moves it. This is a first manual check, not a reporting system. If you want a number every month, you will need something that runs itself. What do you do if AI gets a fact wrong? Sort every wrong fact by where it probably came from, because the fix is different for each source. Your own page is wrong or out of date. The cheapest fix there is. Correct it, date it, and re-run the question that produced the error. Your own page is silent. Nobody invented the fact to spite you. Publish it plainly, in text, on the page where it belongs. A third-party listing about your business is stale. Claim the profile through the platform's verification process and correct it there. Then check the other listings carrying the same fact about you, because they were usually filled in at the same time. Edit listings for your own business only, and never anyone else's. You have been mixed up with another company. Make your own identity unambiguous everywhere: the same legal name, the same address format, the same category, the same founding year. You cannot tell where it came from. Log it and look again next cycle. Some of these settle once the rest is consistent. Re-run the question after each fix, and write down the date you re-ran it. That date is what turns an impression into a record you can show. How often should you repeat it? Quarterly is a sensible cycle for most businesses. Run it sooner after anything that changes your facts: new pricing, a move, a new service, a rebrand, a branch closing. What makes the second check comparable is holding six things constant: the same questions word for word, the same tools, the same tier and account type, the same number of runs per question, the same place and connection, and the same fresh-session setting. Change one and you are comparing two checks rather than two quarters. Lunasol, which owns this publication, works to the same quarterly cadence, re-asking the models "your customers' questions" (Lunasol (https://lunasol.ae/geo-aeo)). The free test sheet The AI Visibility Test Sheet is published with this article as a free Excel download, the same file we would hand to a marketing lead running this for the first time. It has the question bank, a run log carrying the fields above, an error tracker and a source tally. A summary tab counts the runs, the mentions and the wrong facts for any cycle you name, and a second column compares it with the cycle before. The open-error count runs across all cycles. The Run log, Errors and Sources tabs each open with a greyed-out example row. Each is a format guide, not a result, and none of them describes a real company. Start your own rows underneath rather than typing over them, because the summary deliberately skips that row. Keep one copy of the file and add each cycle's runs under the last, with the new cycle name in the Cycle column. On the Summary, type the cycle you want to report in the yellow cell at the top and the one before it in the yellow cell beside it. This article is general guidance on checking what is already public about you, based on six answers logged on one evening in September 2026. AI products change without notice and differ by country, so what we saw may not be what you see, and nothing here can make any tool mention you. Product names are used here and in the test sheet for identification only. ChatGPT is a product of OpenAI; Google Search, AI Mode, AI Overviews and Gemini are products of Google; Perplexity is a product of Perplexity AI; Copilot is a product of Microsoft. All are trademarks of their respective owners. AI Visibility is not affiliated with, endorsed by or sponsored by any of them. Test sheet prepared by Lunasol. Need a more detailed review? Lunasol offers a free AI visibility check, by WhatsApp or email (Lunasol (https://lunasol.ae/geo-aeo)). AI Visibility is owned by Lunasol.