Skip to content

Uncategorized

The four rules that make an AI-visibility number honest

July 28, 2026 · 5 min read · Ignite Agent

Illustrative confidence-range chart showing a range bar with a centre dot, representing a measured range rather than a single number
Illustrative graphic, for visual presentation only. Generated for Ignite.

Every number on a dashboard is a small promise. This is the fourth chapter in a series about building an AI-visibility product in the open, and it is the one where we make the promise plain: we will not fake a number. We will not round one up, pick the flattering denominator, or draw a trend we cannot defend. We would rather show you a smaller true number than a bigger false one.

That reads like a slogan until it costs you something. It cost us a headline.

Why did Ignite kill a flattering dashboard headline?

We killed a dashboard card because reading its data honestly reversed the story it told. Working on our first pilot, we built a card called “Who AI cites” and framed it to look like a win. When we read the underlying data straight, the picture was the opposite: across the AI answers we sampled, other sources out-cited the pilot brand 270 to 126, roughly 2.1 times as many citations.[2] So we deleted the flattering framing and let the humbling read stand.

That 270-versus-126 figure needs its scope attached. It is our own pilot’s snapshot on one date, single-tenant and point-in-time, not a benchmark and not a law of the category, and we will not name the sources that made up the 270. It is simply our own result, and we left it on the screen because it was true. A product that only shows you good news is not measuring anything. It is decorating.

What four rules make an AI-visibility number honest?

Four rules make an AI-visibility number honest: split branded from non-branded, tag grounded versus parametric, attach a confidence range and never publish off a single sample, and withhold when the denominator is undefined. Killing that headline was not a mood. It was these four rules doing their job.

Split branded from non-branded. “How often does AI mention us” blends two very different questions: does the model know your name when asked, and does it recommend you for the category question a stranger types. We always separate them, because the blended rate flatters you with your own brand recognition.

Tag grounded versus parametric. An answer that cites live sources on the open web is a different kind of evidence from one a model recites from memory. We label which is which, so a confident-sounding sentence never gets counted as a citation it did not earn.

Attach a confidence range, and never publish off a single sample. We do not put a number in front of a customer without more than one sample, a rolling window, and a range around it. One lucky read is an anecdote, not a metric. We also exclude our own measurement traffic from any crawl proof, so the tool never cites itself back to you as evidence.

Withhold when the denominator is undefined. If we cannot honestly say “out of what,” we show nothing rather than invent a floor. A blank is uncomfortable. A fabricated percentage is worse.

What has Ignite actually measured on our first pilot?

On our first pilot we have measured the “how AI describes you” half in full, and shipped no fix yet. A readiness probe, reading the counts accumulated over a rolling 21-day window, showed 1,658 AI answer sweeps and 95 brand claims extracted and checked.[3] Read those as one tenant’s counts, single-tenant and point-in-time, not a proof band and not a result anyone should expect.

The other half, shipping an approved fix to the live site and then measuring the result on recrawl, was not built on that date. Improvement insights for this tenant still stood at zero. That is exactly why we will not tell you the pilot’s visibility improved. We have not shipped a change and measured it. Until we do, “it went up” would be a fabricated number, and this series is a promise not to write one.

What does GEO research actually support?

Published GEO research supports one clear lever: adding statistics, citations, and quotations to visible content raised answer visibility 30 to 40% on the GEO benchmark and 22% Position-Adjusted Word Count (Aggarwal et al., GEO, KDD 2024).[1] GEO stands for Generative Engine Optimization. The scope that must ride with that number: “visibility” there means Position-Adjusted Word Count, how much of your own text the answer quotes back, not traffic, not ranking, not sales. It is a strong reason to write substantive, well-sourced pages. It is not a promise that a given pilot will see that lift, and we will not borrow the paper’s percentage and paste it onto our pilot as if we had earned it.

Why does an honest number beat a flattering one?

An honest number beats a flattering one because the flattering one breaks the moment a customer checks it. Our first pilot is the proof: the blended mention rate moved from 3% to 17.8%, which looks like a clean win, but almost all of that lift was branded recognition. On non-branded prompts, the category questions a stranger actually types, the brand was named 0 times out of 24.[4] The 17.8% is the decorating number. The 0 of 24 is the measuring one.

The temptation in this category is to sell certainty, because everyone is anxious about how AI talks about their brand. We think the durable thing is the opposite. Show the real number, split it, range it, and withhold it when the math is not there. Kill your own flattering headline when it is a lie by framing. The day we would rather look good than be right is the day the product stops being worth trusting. Smaller and true beats bigger and fake, every time.

References

  1. [1] Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024. Adding statistics, citations, and quotations raised answer visibility 30 to 40% on the GEO benchmark and 22% Position-Adjusted Word Count. Scope: “visibility” is Position-Adjusted Word Count, how much of your own text the answer quotes back, not traffic, ranking, or sales. https://arxiv.org/abs/2311.09735
  2. [2] Our first pilot, Share of Voice read: other sources out-cited the pilot brand 270 to 126 across the answers we sampled, roughly 2.1 times as many citations. Single-tenant, point-in-time snapshot on one date. Not a benchmark or a category result.
  3. [3] Our first pilot, readiness-probe counts accumulated over a rolling 21-day window: 1,658 AI answer sweeps and 95 brand claims extracted and checked. Single-tenant, point-in-time; the apply-and-re-measure half was not yet built and improvement insights stood at zero for the pilot on that date.
  4. [4] Our first pilot, mention rate: the blended branded rate moved from 3% to 17.8%, but non-branded answers named the brand 0 times out of 24. Single-tenant, point-in-time, branded-inflated. Not a benchmark or a result to expect.

See what AI says about your business.

Run a free scan and find out whether AI names you when your customers ask.

Run my scan