A practical 90 day roadmap for generative engine optimisation works in three phases: a foundation month spent on entity clarity, content structure and measurement setup; a testing month where you expand content and start recording early citation signals; and a refinement month where you scale what is being picked up by AI platforms and rework what is not. This article sets out exactly what to change in each phase, which changes to test in isolation, and how to build a simple tracking system that tells you whether tools such as ChatGPT, Perplexity, Google’s AI Overviews or Copilot are actually referencing your pages. It is written for marketing teams who want a structured, time-boxed way to start GEO testing without guessing which variable caused a change in citation behaviour.
Why generative engine optimisation needs a phased testing approach
Traditional SEO testing usually has a reasonably stable feedback loop: you change something, wait for crawling and indexing, then watch rankings or organic sessions move over a period of weeks. Generative engine optimisation does not behave the same way. AI platforms retrain, adjust their retrieval methods, and change how they select sources far more frequently than a search index refreshes, which means a citation you see today can disappear next week for reasons unrelated to your page.
That instability is exactly why a phased plan matters. If you change ten things on a page at once and it later gets cited by an AI assistant, you have no way of knowing which change caused it. A 30-60-90 day structure forces you to introduce changes in a controlled order, record what you did and when, and check citation behaviour at defined points rather than randomly. It also gives you a realistic timeframe: one month is rarely enough to draw conclusions, but three months of consistent, logged testing usually is.
What “citation” means in this context
Throughout this plan, a citation means any instance where an AI platform references your page as a source, whether that is a visible link, a named brand mention without a link, or your content being paraphrased closely enough that it is clearly the source material. Some platforms show sources explicitly, others do not, so your tracking method needs to account for both visible and inferred citations, which is covered later in this article.
Setting your baseline before day one
Before changing anything, you need a starting point. Skipping this step is the most common reason teams cannot tell whether their 90 day test achieved anything, because they have no record of what the situation looked like beforehand.
- List the pages you intend to include in the test, ideally between five and fifteen pages so the sample is manageable but meaningful.
- For each page, run the queries a potential customer might type into ChatGPT, Perplexity and Google’s AI Overviews, and record whether your brand or page appears, is missing, or appears incorrectly.
- Note the current content structure of each page: heading hierarchy, presence of definitions, use of data, author attribution and publication or update dates.
- Check whether structured data (schema markup) is already implemented, and if so, which types.
- Record organic search position and impressions for the same pages from Google Search Console, so you have a secondary reference point alongside citation data.
Keep this baseline in a simple spreadsheet rather than trying to remember it. You will refer back to it at day 30, 60 and 90, and without a written record the comparison becomes guesswork.
Choosing which pages to include in the test
Not every page is a good candidate. Prioritise pages that answer a specific, well-defined question, since AI platforms tend to cite sources that resolve a query directly rather than pages that are broad or purely promotional. A page explaining “how invoice factoring works” is a better test candidate than a general homepage, because the former has a clear question-answer structure that an AI system can extract and attribute.
Days 1-30: building the entity and content foundation
The first month is about making your organisation and your content easier for an AI system to identify, understand and trust as a source, rather than about volume. Most of the work here is structural rather than about writing new content from scratch.
Entity clarity means an AI platform can confidently connect your content to your organisation, your authors and your area of expertise. This matters because generative engines often draw on multiple signals about who is publishing information, not just the words on the page. If your organisation is described inconsistently across your website, LinkedIn, Companies House listings and industry directories, that inconsistency makes it harder for a system to build a confident association between your brand and the topics you want to be cited for.
Entity and author clarity checklist
- Confirm your organisation name, description and industry classification are consistent across your website, Google Business Profile and any third-party directories you appear in.
- Add or update author bios on key pages, including relevant experience and credentials, rather than leaving content unattributed.
- Ensure your “About” page clearly states what your organisation does, who it serves and how long it has operated, in plain language rather than marketing copy.
- Check that internal linking connects related pages logically, so topical relationships are easy to follow rather than isolated.
- Review whether your organisation is mentioned accurately on any third-party sites, since AI platforms often synthesise information from multiple sources rather than a single page.
Once entity signals are addressed, turn to content structure on the pages included in your test. This is also the stage at which many teams decide to work from an established framework rather than build one from first principles. If you want a structured starting point rather than assembling every phase yourself, adapting a documented generative engine optimization roadmap and mapping it onto your own content calendar and internal resourcing can save several weeks of trial and error during this foundation phase.
Content structure changes to make first
AI platforms generally extract information more reliably from content that answers a question directly near the top of a section, uses clear headings that match how people phrase queries, and separates factual statements from opinion or promotional language. During days 1-30, restructure your test pages so each major section opens with a direct answer, followed by supporting detail, rather than building up to the answer gradually.
Days 31-60: expanding content and testing early signals
By the second month you have made structural changes and given AI platforms time to potentially recrawl and reprocess your content. This is the point to start expanding content depth and to begin your first round of deliberate citation checks, rather than waiting until day 90 to look.
Expansion in this context means adding genuinely useful detail that a person or an AI system could not get from a thin, generic version of the page: specific examples, data you can support, comparisons, and clear explanations of trade-offs. It does not mean padding pages with repetitive keyword variations, which tends to make pages harder for an AI system to summarise cleanly rather than easier.
Structured data and technical additions
During this phase, add or refine structured data where it genuinely reflects the content on the page: FAQPage markup for pages with real question-and-answer sections, Article markup with author and date fields, and Organization markup that matches your entity information from month one. Structured data does not guarantee citation, but it gives AI systems an additional, machine-readable confirmation of facts already stated in your visible content, which can support the trust signals you built in the first month.
At the end of day 45 or so, run your first mid-point citation check using the same queries from your baseline. Record any changes, including negative ones such as a competitor appearing where you previously did not see one. This is a checkpoint, not a verdict; three months of data collection is what makes the pattern reliable, not a single reading at day 45.
Days 61-90: refining based on citation data
The final month is about interpreting what you have observed and making deliberate decisions about what continues, what changes further, and what you abandon. This is where a testing plan earns its value, because by this stage you should have two data points (day 30 baseline and day 45 or 60 check) to compare against a final day 90 review.
Deciding what to scale, pause or rewrite
| Signal observed by day 60 | What it suggests | Recommended action for days 61-90 |
|---|---|---|
| Page is now cited on at least one platform where it previously was not | The structural or entity changes may be contributing, though other factors could also be involved | Keep the page unchanged for the remainder of the test to avoid confounding the result, and apply the same approach to similar pages |
| No change in citation status despite structural changes | The content may still lack sufficient depth, authority signals, or a direct enough answer | Add a genuinely new element such as an original comparison, data point or expert explanation rather than repeating existing points differently |
| Page appears in AI answers but with inaccurate information attributed to you | The AI system is drawing from an outdated version, a third-party source, or misreading ambiguous phrasing | Clarify the specific claim on the page in plain, unambiguous language and check for outdated cached versions or conflicting third-party mentions |
| Competitor now appears where your page previously did, or still does not | Their content may answer the same query more directly or carry stronger entity signals | Review their page structure and public entity information as a diagnostic exercise, then adjust your own content’s directness rather than copying their approach |
Use the final two weeks of the 90 day period to apply the lessons from pages that changed positively to any remaining test pages that have not yet been updated, so the learning compounds rather than staying isolated to one or two pages.
How to track AI citations across platforms
Tracking citations manually is currently the most reliable method available to most UK marketing teams, since automated tools for this purpose are still maturing and vary considerably in accuracy. A consistent manual process, run on a fixed schedule, will give you more trustworthy data than an inconsistent automated one.
Manual prompt testing workflow
- Build a fixed list of 10 to 20 queries that reflect how a real customer would phrase a question related to your test pages, rather than exact-match keywords.
- Run each query on the same set of platforms every time, for example ChatGPT with browsing enabled, Perplexity, and Google’s AI Overviews where they appear for that query.
- Record, for each query and platform, whether your page is cited, whether a competitor is cited instead, and whether the information presented about you is accurate.
- Take a screenshot or copy the exact response text alongside the date, since responses can change between sessions and you will want evidence of what was shown on a given day.
- Repeat this exact process on the same day of the week, for example every Monday, so your data reflects a consistent cadence rather than random spot checks.
- After each round, update your baseline spreadsheet with a new column for that date, so you can see the full history of each query rather than only the most recent result.
Setting up a citation tracking log
A simple spreadsheet with one row per query and one column per check date is sufficient for most teams starting out. Include columns for platform, cited (yes/no), accuracy (accurate/inaccurate/partial), and a free-text note field for anything unusual, such as your page being referenced without a clickable link.
| Metric | What it measures | How to interpret it |
|---|---|---|
| Citation frequency | How often your page appears as a source across your fixed query list over time | An upward trend across several check-ins suggests your changes are having some effect; a single high or low reading is not conclusive |
| Citation accuracy | Whether information attributed to you is correct when it does appear | Frequent inaccuracies point to ambiguous phrasing on the page or outdated cached content rather than a visibility problem |
| Visible link vs brand mention only | Whether the platform provides a clickable source or simply names your organisation | A mention without a link still indicates the AI system associates you with the topic, though it will not generate direct traffic |
| Referral traffic from AI platforms | Sessions in your analytics tool attributed to AI platform referrals, where that data is available | Useful as a supporting indicator, but currently unreliable as a sole measure since not all platforms pass clear referral data |
| Position within the answer | Whether your source is the primary reference or one of several listed | Being the sole or primary source generally indicates stronger perceived authority on that specific query than being listed alongside several competitors |
Common mistakes that undermine a 90 day test
Several recurring problems make it difficult for teams to draw any useful conclusion from a GEO testing period, even when the underlying content work was reasonable.
| Problem | Likely cause | Corrective action |
|---|---|---|
| No clear result after 90 days | Too many pages changed at once, or changes made without a written baseline to compare against | Reduce the test to a smaller page set and ensure a documented baseline exists before the next testing cycle begins |
| Citation appears, then disappears within a few weeks | Platform-side changes to retrieval or ranking of sources, unrelated to your content | Continue tracking rather than reacting immediately; a single disappearance is not evidence the underlying approach failed |
| Team disagrees on whether a “citation” occurred | No shared definition was agreed before testing began | Define citation types clearly in writing at the start, distinguishing visible links, mentions and paraphrased content |
| Content becomes repetitive or keyword-heavy after expansion | Volume was prioritised over genuinely new information during the days 31-60 phase | Review expanded sections and remove any paragraph that repeats an existing point rather than adding a new one |
What to do differently after running this plan
Completing one 90 day cycle should change how you approach content and technical work going forward, not just produce a one-off report. The following adjustments reflect what typically needs to happen once you have real citation data rather than assumptions.
- Treat entity consistency as ongoing maintenance, checking it quarterly rather than only at the start of a test, since directory listings and third-party mentions drift over time.
- Build the manual prompt tracking workflow into a recurring task with a named owner, rather than a one-off exercise tied to this specific test period.
- Use the day 90 results to prioritise your next batch of pages, starting with topics closest to those that already showed positive citation movement.
- Share the day 90 findings with whoever writes or approves content, so structural lessons (direct answers first, clear attribution, supporting data) are applied to new content from the outset rather than retrofitted later.
A useful decision framework at this stage is to ask three questions about each test page: did citation behaviour change, did the change appear connected to a specific action you took, and is the result stable across at least two check-ins rather than one. Pages that answer yes to all three become your template for the next round of pages.
Frequently asked questions
How long does it take for AI platforms to start citing a page after changes are made
There is no fixed timescale, since different platforms recrawl and reprocess content at different rates and some rely partly on cached or delayed data. Some teams see a change within a few weeks, others see nothing measurable within the full 90 day period. This is why the plan builds in checkpoints at day 30, 45 to 60, and 90, rather than expecting a result at any single point.
Which AI platforms should I include when tracking citations
At minimum, include Google’s AI Overviews where they appear for your queries, ChatGPT with browsing or search enabled, and Perplexity, since these represent different retrieval approaches. If a meaningful share of your audience uses Microsoft Copilot or another assistant, add that too, but keep the list consistent across every check so your comparisons remain valid over time.
Do I need structured data for generative engine optimisation to work
Structured data is not confirmed as a requirement for citation, but it gives AI systems a clearer, machine-readable version of facts already present in your visible content, such as author identity, publication dates and question-and-answer pairs. Treat it as a supporting technical signal alongside clear content structure and entity consistency, not as a substitute for either.
How is testing for generative engine optimisation different from normal SEO testing
Traditional SEO testing usually relies on a relatively stable ranking signal and a defined crawl and index cycle. GEO testing deals with retrieval systems that can change frequently and do not always disclose their sources, which means results can be less consistent between check-ins. This is why manual, scheduled tracking with a written baseline matters more here than in typical keyword rank tracking.
What actually counts as a citation from an AI platform
A citation can be a visible clickable link to your page, a named mention of your brand or organisation without a link, or content that is clearly paraphrased from your page even without direct attribution. Recording all three types separately in your tracking log gives a fuller picture than only counting visible links.
Should I test on new pages or existing ones
Existing pages with some history and a clear topic focus are usually better test candidates for a first 90 day cycle, because you can establish a genuine baseline and are not also waiting for a brand new page to be discovered and indexed. Once you have a working process, apply the same structure to new content from the point of publication.
How much should I change a page between the day 30 and day 60 checkpoints
Make one meaningful addition per checkpoint, such as a new comparison, a clarified claim, or additional structured data, rather than rewriting the page substantially each time. Frequent large-scale rewrites make it harder to connect any citation change to a specific action, which defeats the purpose of a phased test.
Next steps: prioritising your first testing cycle
Start by selecting no more than ten pages, write down their current citation status and content structure today, and commit to the manual tracking workflow on a fixed weekly or fortnightly schedule before making any content changes. Once that baseline exists, work through the days 1-30 entity and structure checklist in full before moving to content expansion, since skipping the foundation phase tends to produce unclear results later. Review your findings honestly at day 90: where citation behaviour improved and appears stable across two or more checkpoints, apply the same approach to a wider set of pages; where nothing changed, revisit whether the page answers its target question directly enough before assuming the whole approach has failed.