AI assistants don't return the same answer every time. That makes “are we mentioned in ChatGPT?” a question with a probabilistic answer. Reporting it well takes a little statistics, and the payoff is decisions you can defend.
Visibility is a rate, not a yes or no
If you collect 30 answers to your tracked prompts and your brand appears in 18, your observed visibility is 60%. But that 60% is an estimate. Collect another 30 answers and you might see 15 or 21 mentions just by chance.
Confidence intervals, and why Wilson
A 95% confidence interval gives the range of visibility rates consistent with what you observed. For proportions, the Wilson score interval is a good choice: it stays sensible for small samples and for rates near 0% or 100%, where the simpler textbook formula produces impossible values.
For 18 mentions in 30 answers, the Wilson 95% interval is about 42% to 75%.
| Mentions / answers | Interval | Width |
|---|---|---|
| 6 / 10 | 31.3% – 83.2% | 51.9 pts |
| 18 / 30 | 42.3% – 75.4% | 33.1 pts |
| 60 / 100 | 50.2% – 69.1% | 18.9 pts |
| 180 / 300 | 54.4% – 65.4% | 11.0 pts |
The interval narrows roughly with the square root of the sample size: to halve its width, you need about four times as many answers.
Telling a real change from noise
Suppose visibility drops from 60% to 40%. Whether that's meaningful depends on how many answers sit behind each number. A two-proportion test answers it:
- 18 of 30 → 12 of 30: p ≈ 0.12. Not significant at the usual 0.05 threshold; this drop could easily be noise.
- 60 of 100 → 45 of 100: p ≈ 0.03. Significant: something probably changed.
The same 15-to-20 point drop can be noise or a real signal. Alerting on raw percentage changes produces false alarms; alerting on significant changes doesn't.
How to get enough answers
- Track more prompts on the same topic and report them together.
- Sample each prompt more than once per period.
- Use longer periods (weekly rather than daily) when volume is low.
- Report per engine and never average API answers with consumer-app answers.
Treat failures as gaps, not zeros
Engines sometimes time out or refuse. If a failed collection is counted as “not mentioned”, your visibility silently drops for reasons that have nothing to do with your brand. Exclude failed runs and show them as missing data.
This is how Vedlora reports visibility: every rate has a Wilson interval, alerts fire only on significant changes, and failures are shown as gaps. The methodology covers the details.