If you want to know how to measure AI share of voice for generative engine optimisation, the honest answer is that there are two genuinely different routes, and most businesses end up mixing both rather than picking one outright. Dedicated citation-monitoring platforms automate the process of asking AI engines questions and logging when your brand appears, while manual prompt testing involves someone on your team running the same queries by hand and recording the results in a spreadsheet. Neither approach is universally superior. Paid tools scale better and produce cleaner trend data, but manual testing gives you direct control over prompt wording and context that automated tools sometimes miss. This article sets out exactly what each method can and cannot measure, gives you a decision framework for choosing between them, and shows you how to run a manual tracking process properly if that is where you start.
What AI share of voice actually means in a GEO context
AI share of voice describes how often your brand, product or content is cited, mentioned or recommended when someone asks an AI engine a question relevant to your market. Unlike a search engine results page, there is no fixed list of ten blue links to count. Instead, the AI generates a conversational answer, and your brand either features in that answer or it doesn’t. Measuring this means running representative questions through engines such as ChatGPT, Gemini, Perplexity or Copilot, then recording whether, how and in what context your brand appears relative to competitors.
How it differs from traditional share of voice
Traditional share of voice in search or PPC relies on stable, queryable data: impression share, ranking position, ad auction data. AI share of voice is inherently less stable, because generative engines can produce different answers to the same prompt depending on when you ask, how the prompt is phrased, and which underlying model version is being used. This means any measurement method, paid or manual, is sampling a moving target rather than reading a fixed report. That instability is precisely why the choice of measurement method matters so much, because a method that doesn’t account for variability will give you misleading confidence in either direction.
The three things any tracking method needs to capture
Whichever approach you use, a measurement method only earns its keep if it captures three things reliably: whether your brand is mentioned at all, the context in which it is mentioned (recommended, listed as an alternative, criticised), and how that compares with named competitors across the same set of prompts. A method that only tells you “yes we were mentioned three times this month” without context or competitive framing is not really measuring share of voice, it’s measuring raw mention frequency, which is a different and much less useful metric.
The two ways to track AI citations
Dedicated AI monitoring platforms work by running a bank of prompts against multiple AI engines on a schedule, then using natural language processing to detect brand mentions, categorise sentiment and produce dashboards. Manual prompt testing is simply a person typing questions into ChatGPT, Gemini or Perplexity, reading the answers, and logging what they find in a spreadsheet or document. The mechanics are different, but the underlying goal, understanding how often and how favourably your brand appears in AI-generated answers, is the same.
| Aspect | Dedicated monitoring platform | Manual prompt testing |
|---|---|---|
| Engine coverage | Typically covers several engines simultaneously through one interface | Limited to whichever engines someone actively logs into and queries |
| Ongoing effort required | Low once set up, mostly reviewing dashboards | Ongoing manual time each testing cycle, scales with prompt volume |
| Historical trend data | Automatically stored and comparable over time | Only as good as your own record-keeping discipline |
| Prompt customisation | Often constrained to the platform’s prompt templates or categories | Fully flexible, you write exactly the phrasing your customers use |
| Cost structure | Recurring subscription, usually priced by prompt volume or seats | No direct tool cost, but consumes staff time |
What paid AI monitoring platforms actually measure
Paid platforms in this category generally focus on quantifying mention frequency and competitive positioning at a scale that would be impractical to replicate by hand. They run large batches of prompts on a repeating schedule and use automated text analysis to flag brand mentions, extract surrounding sentiment, and tag which sources the AI engine appears to be drawing from when it names a brand.
Core metrics you’ll see in a vendor dashboard
Most platforms in this space report on a similar set of underlying metrics, even though the naming conventions differ between vendors. Understanding what each metric actually captures matters more than the label attached to it.
| Metric | What it measures | Interpretation |
|---|---|---|
| Citation frequency | How often your brand appears across the tested prompt set within a period | A rising trend suggests growing AI visibility, but check whether the prompt set changed too |
| Prompt coverage | The proportion of tracked prompts where any brand in your category is mentioned at all | Low coverage means the category isn’t strongly represented in AI answers yet, not that you’re failing |
| Source attribution | Which web pages or domains the AI engine appears to reference when citing your brand | Useful for identifying which of your pages are being pulled into AI answers |
| Sentiment or context tag | Whether the mention is neutral, recommending, or comparing unfavourably | Context matters more than raw frequency, a high mention count with negative framing is a warning sign |
| Competitive share | Your mention frequency relative to named competitors across the same prompts | This is the closest equivalent to a genuine share of voice figure |
What these platforms typically cannot do reliably is explain why a particular answer was generated, since the underlying AI models don’t expose their reasoning, and vendors are inferring attribution patterns rather than reading confirmed source logic. Treat source attribution data as a strong indicator rather than a guaranteed causal link.
What manual prompt testing can and cannot measure
Manual testing gives you direct, first-hand visibility into exactly how an AI engine answers a specific question at a specific moment, phrased exactly the way you choose. That precision is valuable when you’re testing high-value prompts tied to particular products, services or comparison queries that a generic vendor prompt bank wouldn’t necessarily include.
Setting up a manual tracking log
A manual process only produces useful data if it is structured consistently from the start. The following workflow describes how to set one up properly rather than starting with an ad hoc spreadsheet that becomes unusable after a few weeks.
- Define a fixed list of 15 to 30 prompts that reflect real customer questions, informational queries and comparison-style questions relevant to your services.
- Decide on a testing cadence, for example fortnightly or monthly, and commit to running the full prompt list on the same schedule every time.
- Run each prompt in a fresh, logged-out session where possible, to reduce the influence of personalisation or prior chat history on the answer.
- Record, for each prompt, whether your brand was mentioned, the exact wording used to describe it, and which competitors appeared alongside it.
- Tag each mention as positive, neutral or unfavourable based on how the AI engine framed your brand relative to alternatives.
- Store results in a single running spreadsheet with a date column, so trends across testing cycles can be compared rather than reviewed in isolation.
- Review the log at the end of each cycle and note any prompts where competitor mentions increased or your brand dropped out entirely.
Where manual testing breaks down at scale
The workflow above works well for a focused set of priority prompts, but it has structural limits once you try to extend it.
- Testing across multiple engines multiplies the time required, since ChatGPT, Gemini and Perplexity all need separate sessions and separate logging.
- Manual logs are vulnerable to inconsistent tagging when different team members interpret “positive” or “neutral” mentions differently over time.
- There is no automated way to detect subtle wording changes in how an engine describes your brand unless someone re-reads previous log entries carefully.
- Historical comparison becomes unreliable if testing cadence slips, which happens easily when the task competes with other marketing priorities.
A decision framework for choosing between paid and manual tracking
The right method depends on how many prompts you need to track, how many engines matter to your business, and how much internal capacity you have to run testing consistently. The table below sets out the signals that tend to point towards one approach or the other, though most organisations will find their situation sits somewhere between the two extremes rather than cleanly on one side.
| Decision criterion | Signal favouring manual tracking | Signal favouring a paid platform |
|---|---|---|
| Number of priority prompts | Under roughly 30 prompts covering your core services | Dozens or hundreds of prompts across multiple product lines |
| Number of engines to monitor | One or two engines where your customers are known to be active | Three or more engines with materially different answer behaviour |
| Internal capacity | A team member can commit a fixed slot each testing cycle | No one has consistent bandwidth to run and log tests reliably |
| Reporting requirements | Internal use only, informal trend tracking is sufficient | Needs to be presented in board or client reporting on a recurring basis |
| Competitive set size | A handful of known, stable competitors | A larger or shifting competitive landscape that’s hard to track by hand |
Cost thresholds worth testing against
There is no universal budget figure that makes a paid platform worthwhile, since pricing and internal staff costs vary considerably between organisations. As an internal example only, if your team already spends more than a day per month running and logging manual tests, it can be useful to compare that time cost against a platform’s subscription fee to see which represents better value for your specific circumstances. That threshold isn’t an industry standard, it’s simply a starting point for building your own comparison.
For organisations running a wider generative engine optimisation programme across content, structured data and citation-building activity, tracking is only one part of a larger picture. If you’re weighing up whether to bring in outside support for the broader strategy rather than just the measurement piece, it’s worth looking at what a specialist offering generative engine optimisation services would cover beyond monitoring, since tracking data is most useful when it feeds directly into content and technical decisions rather than sitting in a report on its own.
Common measurement mistakes and how to avoid them
Both paid and manual approaches can produce misleading conclusions if the underlying process has gaps. The mistakes below turn up regularly regardless of which method a business has chosen.
| Problem | Likely cause | Corrective action |
|---|---|---|
| Share of voice appears to swing wildly week to week | Prompt set is too small or being changed between testing cycles | Fix a stable prompt list and only add new prompts as a separate, clearly labelled batch |
| Brand mentions look strong but leads haven’t changed | Mentions are neutral or comparison-only rather than recommending your brand | Add a sentiment or context tag to every logged mention, not just a yes or no count |
| Competitor appears to have overtaken you suddenly | A single testing session caught an unusual answer variation | Re-run the same prompt across a few sessions before treating one result as a trend |
| Data can’t be presented to stakeholders convincingly | Manual log lacks consistent dating, tagging or a competitor comparison column | Standardise the spreadsheet structure before the next testing cycle begins |
- Testing only branded or product-name prompts, which misses the informational questions where AI engines most often introduce new brands to a potential customer.
- Treating a single AI engine’s results as representative of AI visibility overall, when ChatGPT, Gemini and Perplexity can behave quite differently for the same query.
- Comparing this month’s manual results against last month’s paid platform report, or vice versa, without acknowledging the two methods use different sampling logic.
Building a hybrid measurement approach
Most organisations get more useful data from combining both methods deliberately rather than defaulting to one exclusively. Paid platforms are well suited to broad, ongoing coverage across many prompts and engines, while manual testing is better suited to deep, qualitative checks on your highest-value queries. Used together, each compensates for the other’s blind spots.
Weekly manual checks that complement paid data
Even organisations running a paid monitoring subscription benefit from a lightweight manual layer focused on the small number of prompts that matter most commercially, since these are the queries worth reading in full rather than relying purely on an automated sentiment tag.
- Identify the five to ten prompts most directly tied to purchase decisions, such as “best [service] provider for [use case]”.
- Run these manually once a week, reading the full AI answer rather than just checking whether your brand appears.
- Compare what you read against what the paid platform’s dashboard reported for the same prompts, noting any discrepancies.
- Use discrepancies as a prompt to investigate rather than assuming either source is automatically correct.
- Feed anything notable, such as a competitor being newly recommended, into your content or GEO planning discussions the same week rather than waiting for a monthly report.
What to do differently once you have this data
Collecting AI share of voice data only creates value once it changes what your team does next. The checklist below turns raw tracking results into specific actions rather than leaving them as a passive report.
- If your brand is absent from a prompt where a competitor consistently appears, review whether you have a page that directly answers that exact question, and if not, brief one.
- If mentions are frequent but framed neutrally rather than favourably, look at whether your content includes the kind of specific, comparative detail that gives an AI engine a reason to recommend you rather than just list you.
- If source attribution data points to a particular page driving citations, prioritise keeping that page updated and accurate, since it appears to be functioning as a reference point for AI engines.
- If testing consistently shows gaps across a whole topic area, treat that as a content gap to close rather than purely a measurement issue.
Frequently asked questions
What counts as an AI citation when measuring share of voice
An AI citation, for measurement purposes, is any instance where an AI engine’s generated answer names your brand, links to your website, or references your product or service by name, whether as a direct recommendation, one option among several, or a comparison point. It’s worth distinguishing this from a simple keyword match, since some tracking tools flag partial matches or generic industry terms that don’t actually reference your specific organisation. When reviewing results, check a sample of flagged citations manually to confirm they genuinely name your brand rather than a similarly worded competitor or generic term.
How often should I run manual prompt checks
There’s no fixed interval that suits every business, but a fortnightly or monthly cadence tends to strike a reasonable balance between capturing meaningful change and keeping the workload manageable. Testing too frequently, such as daily, often just captures normal answer variation rather than genuine trend movement, since AI engines can produce different phrasing for the same prompt from one session to the next. Testing too infrequently, such as quarterly, risks missing shifts that would have been useful to catch earlier. Choose a cadence you can sustain consistently, since a broken or irregular schedule undermines the value of the historical comparison more than a slightly longer gap between tests would.
Can I use the same prompts every time I test
Using a stable core set of prompts is actually recommended, since it’s the only way to build a genuine trend line rather than comparing unrelated data points. That said, it’s sensible to periodically add new prompts as customer language or your service offering evolves, provided you log these as a distinct batch rather than mixing them into your historical comparison set. A practical approach is to keep roughly 80 per cent of your prompt list fixed for continuity and dedicate the remainder to testing emerging questions or new competitor entrants.
Do paid AI monitoring tools cover Perplexity and Gemini as well as ChatGPT
Coverage varies by vendor and changes as the market develops, so it’s worth checking a platform’s current engine list directly rather than assuming coverage before subscribing. Some platforms cover ChatGPT extensively but offer more limited or delayed coverage of Gemini and Perplexity, particularly where those engines have different technical access arrangements for automated querying. If monitoring a specific engine is a priority for your business, confirm that coverage explicitly during any vendor evaluation rather than relying on general marketing claims about “AI engine coverage”.
Is AI share of voice the same as traditional search share of voice
They measure related but distinct things. Traditional share of voice is usually based on ranking positions or impression data within a search engine results page, which is queryable and relatively stable. AI share of voice is based on whether and how your brand appears within a generated conversational answer, which can vary between sessions even for an identical prompt. The two metrics often move independently of each other, so a strong traditional SEO position doesn’t guarantee AI citation, and the reverse is also true.
How many prompts do I need for a reliable manual sample
There’s no universally agreed minimum, but a set that’s too small, say under ten prompts, tends to produce results that swing noticeably between testing cycles simply due to normal AI answer variation rather than any real change in visibility. A prompt list in the range of fifteen to thirty focused, high-relevance questions tends to give a more stable picture while remaining manageable for manual testing. Treat this as a practical starting point to refine based on your own results rather than a fixed rule.
Should I track competitors as well as my own brand
Yes, tracking competitor mentions alongside your own is what turns a simple mention count into an actual share of voice metric. Logging only your own brand’s presence tells you whether you appeared, but it doesn’t tell you whether the AI engine consistently favours a competitor for the same questions. Recording which named competitors appear on each prompt, and in what order or framing, gives you the comparative context needed to judge whether your visibility is genuinely competitive or simply present.
Set up your tracking process this month
Rather than treating AI share of voice as something to figure out eventually, pick a starting point now based on where your business actually sits. If you’re tracking fewer than thirty priority prompts across one or two engines with no dedicated reporting requirement, build the manual log described earlier and commit to a fixed testing cadence before adding any tooling cost. If you’re already managing prompts across multiple engines, need recurring stakeholder reporting, or have identified that manual testing is consuming more staff time than it’s worth, use the decision framework table to build a short business case for a paid platform rather than adopting one on the strength of a vendor demo alone. Whichever route you choose, review your method every quarter against the mistakes outlined above, since the biggest risk to this kind of tracking isn’t picking the wrong tool, it’s letting an inconsistent process quietly undermine the data either one produces.