Your AI Visibility Report Is Mostly Made Up

Share

Summary

The good stuff isn't in the report, "rank" is irrelevant when the answer keeps changing anyway and the very data these vendors want to sell you is already yours! Matt Heinz argues that most AI-visibility reports are built on inferred, not observed, query data, that answers from ChatGPT and other tools change too fast for "rank" to mean much, and that the real fix is a corroboration audit of your own sources rather than another vendor dashboard.

Answer-engine tools, and there are a growing number of them, often share the same three slides:

  • Here are the prompts your buyers use to find you.
  • Here’s where you show up in the answer.
  • Here’s the competitor showing up above you.

It’s compelling stuff. But you should be asking where that data comes from and how reliable it is.

Ask that and things get vague quickly.

Nobody actually has the queries

Some vendors are honest about this. Otterly’s own help documentation says it plainly: “AI search engines like ChatGPT and Perplexity do not publish query data. There is no public keyword tool for AI search.”

Their prompts get built from Google Search Console data, SEO keywords and AI-brainstormed ideas. Ahrefs publishes its methodology too – queries pulled from Google’s People Also Ask and its own keyword database, then expanded into related sub-questions by a semantic system they call Fanout. Both companies deserve credit for showing their work.

Others are less forthcoming. One well-funded platform advertises “real prompts submitted to AI platforms by actual users” without explaining anywhere I could find where those prompts come from. Another builds its suggested prompts from your website and your industry context, then scores them against Google search trends. A proxy for a proxy. Clear as mud, right?

This is not to say that the tools are useless. It means the thing being sold as observation is mostly inference. That’s a whole different level of confidence, and could change how you think about making that investment in the first place.

The three companies who do have the data aren’t sharing

Three companies (guess who??) do have the data you want and none of them have a business reason to hand it over.

Google shipped a new Search Console report for site owners recently covering AI Overviews and AI Mode. Impressions, pages, countries, devices, dates. It leaves out the actual queries and the clicks.

Asked about adding them, Google said it’s “continuing to work with website owners to understand what insights will be most helpful.” OpenAI’s advertising product is just as clear – advertisers “only receive aggregate information about how their ads perform such as number of views or clicks,” with no access to chats, history or memories. Anthropic has said it won’t run ads in Claude at all, so there’s no advertiser pipeline there to ask about.

Google built an entirely new reporting surface for marketers and left out the one thing marketers want most. I’d be careful reading that as a roadmap gap.

Even the answer appears to change regularly

Say you had the perfect prompt list anyway. It appears the answers will still change constantly.

A recent crowdsourced study ran a set of prompts through ChatGPT, Claude and Google’s AI Overview nearly 3,000 times using 600 volunteers. It found less than a one-in-a-hundred chance that two responses to the same prompt return the same list of brands and closer to one in a thousand that they return them in the same order.

A separate study looking at 815,000 prompt-page pairs found that after running the same prompt three times in ChatGPT, only about 2 percent of citations survive. Another study tracking more than 82,000 prompts over 17 weeks found ChatGPT swapping roughly three-quarters of its cited sources week over week.

Rank isn’t a thing here. Visibility measured across hundreds of prompts over months is directional and genuinely useful. A screenshot of where you placed on Tuesday is noise and extremely temporary, and a lot of teams are reporting the screenshot to their board.

What it does read mostly isn’t “from” you

When these systems answer a question about your category, they’re largely not reading your website. A recent analysis of more than 100 million citations across ChatGPT, Google AI Mode and Perplexity found the most-cited domains were Reddit, Wikipedia, Medium, Forbes, LinkedIn and YouTube. No vendor site cracked the top five on any platform.

Other research found that on Gemini, the overlap between the brands named in an answer and the domains actually cited can run as low as 30 percent. You get mentioned (great) and something else supplies the evidence (less great!).

Review sites, in this context, are still incredibly important. G2 looked at more than 80,000 of its own product pages recently and found four out of five get cited by AI systems more often than they get viewed by a human. G2 sells vendor visibility and the tracking came from a company selling AI visibility software, so weigh it accordingly. It still points the same direction as everything else.

So, in summary:

  1. Your buyer’s research is being assembled out of sources you don’t own
  2. Weighted by how consistently they agree with each other
  3. And you can’t see the question that started it

Are we having fun yet?

Which makes this a buyer enablement problem (and opportunity)

This all makes buyer enablement more important, not less.

Recent 6sense research (they sell intent data and have a stake in the answer) found 94 percent of buying groups had ranked a preferred vendor before first contact, and 77 percent went on to buy that preliminary favorite. What’s new is that the ranking is now partly assembled by a machine reading forum threads and review pages on the buyer’s behalf.

Gartner recently found that 69 percent of B2B buyers prefer to validate AI-generated insights with a sales rep. By the time anyone gets on a call the opinion already exists and the call is more about validation than discovery.

A few things I’d do about all of this. Some of it is work we do at Heinz Marketing, so weigh it accordingly.

Build your own question set. You don’t need an outside vendor to tell you what your own data already knows! The questions your buyers actually ask are sitting in your sales calls, your support tickets, your win/loss debriefs and your community threads. That’s an invaluable and competitively-protected source nobody can sell you, because you’re the only one with it.

Optimize for corroboration instead of placement. If consistency across independent sources is what these systems reward, the work is making sure your positioning, your review profiles, your customer stories, your analyst coverage and your own site all say the same true thing. Most B2B companies have four versions of who they serve living in four places. That inconsistency used to cost you a confused visitor and now it costs you the citation.

Arm the champion for a validation conversation. Your champion isn’t discovering you in that first call, they’re defending a choice they already made to a buying group that didn’t make it with them. Different content for the CFO than for the security lead. Proof points procurement can forward. When machines are doing the heavy-lifting early in the buying journey, your champion enablement work is about defending a recommendation and helping to drive consensus and velocity towards a group commitment to change. Different job, different call, different content.

Ungate the evaluation. Most of our clients ungated content years ago once they realized the gate cost more than it captured, and every panelist on a recent buyer enablement roundtable I hosted said some version of the same thing about demos. If a buyer can get a straight answer about your product from a machine in nine seconds, making them wait two weeks and sit through a discovery call before they’re allowed to see it is a disqualifying move. They’ll just go ask the machine again.

What a corroboration audit actually looks like

If that inconsistency sounds familiar it should. It’s the same dissonance you get when sales, marketing and product each carry a different definition of who you’re targeting – chaos and tension inside the building, mixed signals outside it. I’ve argued before that your ICP works better as a living wiki than a document gathering dust in a folder. This is that same discipline pointed outward, and the audience for the inconsistency is a lot bigger now.

The audit itself isn’t complicated:

  1. Take ten questions from your own list. The real ones buyers ask, out of the calls and tickets and win/loss debriefs above.
  2. Run each one several times across ChatGPT, Gemini and Perplexity. Save what gets cited, not just what gets said. The repetition is the point – one run tells you nothing.
  3. Open the cited sources side by side. Your site, your review profiles, the Reddit threads, the press coverage, your team’s LinkedIn profiles.
  4. Ask one question of the whole stack. Do these describe the same company, serving the same buyer, solving the same problem? Mark every place they don’t.
  5. Fix what you own first. Site, review profiles, LinkedIn, sales collateral. Then work the ones you only influence – customer stories, analyst briefings, press.

The output is a list of contradictions. Most teams are surprised how long it runs, and a good number of them turn out to predate any of this. They were always there. The machine is just the first reader thorough enough to catch all of them at once.

One parallel caution worth watching

These systems are pattern-matchers across sources, which means the buying question and the employment question are drawing from an overlapping pool. A friend described a CEO years ago who was so upset about his Glassdoor reviews that he paid people overseas to write better ones. There was a time when this worked and you could get away with it.

I’m drawing conclusions ahead of the data here, but the logic feels sound. It’s already failing, and from here it reverses. A cluster of glowing reviews that contradicts every other available signal reads as noise to a system built to reward agreement. And it isn’t a big leap from there to an AI summary telling a buyer that a company builds good products and treats its people badly, then leaving them to draw the obvious conclusion about what kind of vendor that makes.

Oh marketers, we so crave control. And I know what you just read means you’re getting even less of it.

What you can control is whether every source that machine reads tells the same true story about who you help and why it works. And in that way it’s a familiar job: building confidence with buyers we’ve always been responsible for, now graded by a machine that never visits the site and never fills out a form.

New rules, new constraints, new opportunities.

This post originally appeared on Matt Heinz’s Substack.