Quick answer: A 90-day GEO roadmap runs in three phases. Days 1–30 fix the entity, schema and crawlability, pick three intent clusters and publish the first four answer-first pages. Days 31–60 test framing, format and freshness against citation data in Bing Webmaster Tools. Days 61–90 scale the winning framings, broaden the query base and add third-party corroboration. AI Studio measures all of it in citations, not traffic.
On this page
What should you measure before day 1?
Start with a baseline you can compare against at day 90, or the quarter becomes anecdote. It takes one afternoon and no paid software.
- Citation count and cited pages. In Bing Webmaster Tools' AI Performance report (released February 2026, still the only free first-party record of which pages an assistant used), record the last 30 days: total citations, distinct pages cited, top grounding queries. If it is empty, write down the zero.
- Google's view, where you have it. Search Console's generative-AI views (launched June 2026, still rolling out by region) report AI Overviews and AI Mode impressions but no clicks. Record them if available; do not wait on them.
- The 10 buying questions. Write the ten questions a buyer asks an assistant before choosing a vendor like you: "best X in Singapore", "how much does X cost", "X vs Y". Run each through ChatGPT, Copilot, Perplexity and Gemini and note whether you are named, and who is.
- Entity check. Ask each assistant "what is [your brand] and what does it do?" If the answers disagree with each other or with your site, fix the entity first.
Our guide to tracking AI citations across Bing, Search Console and GA4 covers each source in detail.
Days 1–30: what does the foundation phase include?
The first month has one job: make the site something an assistant can crawl, understand and attribute, then give it four pages worth citing.
Weeks 1–2: entity, schema and crawlability
- Write one canonical description of the business (who, what, where, for whom) and use it verbatim on the site, in Organization schema and in every directory profile. Disagreement between these is the most common reason a brand is retrieved but never named.
- Validate structured data on every template: Organization, Article with real dates, BreadcrumbList, FAQPage. Broken JSON-LD fails silently.
- Confirm robots.txt is not blocking the crawlers you want, Bing has a verified property and sitemap, and key pages return clean 200s without JavaScript-only content.
- Publish an
llms.txtfile describing the site and its key pages. Engine support is uneven; treat it as cheap hygiene, not a lever.
Weeks 3–4: three clusters, four pages, one scoreboard
- Pick three intent clusters from your buying questions: typically "what is / how does", "cost / pricing" and "best / vs / how to choose". Three is enough to compare; more dilutes the test.
- Publish the first four answer-first pages, at least one per cluster. Each opens with a direct answer, uses question-phrased H2s, carries one comparison table and a FAQ mirrored in schema. Guides and comparisons, not service pages.
- Set up measurement. A spreadsheet with one row per week: citations, cited pages, top grounding queries, top-query share, plus your manual prompt sample. Fill it every Monday for 13 weeks.
Then stop changing things. Engines need a stable site to re-evaluate; a team that keeps editing during the read window cannot tell what worked.
Days 31–60: how do you test what earns citations?
Month two is about hypotheses, not volume. You have four pages and a few weeks of citation data; the question is which properties of those pages the engines respond to.
Four hypotheses worth testing first
- Format. Does a page built around a comparison table out-cite the same topic as prose? Publish one of each in the same cluster and compare after three weeks.
- Framing. Does "how to choose a [vendor]" out-cite "best [vendor] in Singapore" for the same intent? The grounding-query report shows which phrasing the assistant actually searched for.
- Freshness. Does a visible "Updated [month]" line plus a refreshed dateModified change retrieval on an older page? Refresh two pages, hold two back.
- Hub-and-spoke. Does linking four spokes to one hub guide lift citations for the hub, the spokes or neither? This test most often surprises teams.
What a valid test looks like with noisy data
Citation counts swing week to week for reasons unrelated to you: a model update, a change in report aggregation, a single query spiking. So a valid test changes one variable, runs at least three weeks, and is judged on three numbers: total citations, distinct cited pages, and top-query share. A "win" that only moved the total is suspect. For the discipline of forming and controlling hypotheses, use the AEO experimentation roadmap.
Reading the grounding-query report
The grounding-query view is the closest you get to the assistant's intent. Pages cited for queries matching their intent are winners: reuse the phrasing. Pages cited for queries they were not written for are mis-framed: rewrite the opening and H2s to match. Pages never cited after four weeks indexed are losers: kill or merge them, and write no more in that framing.
Days 61–90: how do you scale what worked?
Month three doubles down on the framings that earned citations and deliberately spreads the risk.
- Double down across clusters. If "how to choose" out-cited "best of", write "how to choose" pages in the other two clusters, not a fifth variant in the first.
- Broaden the query base. If most citations sit in one grounding query, you are one model update from zero. Add adjacent-question pages, open a fourth cluster, and treat a falling top-query share as success.
- Third-party corroboration. Assistants weigh what other sites say about you. Pursue mentions on publicly verifiable third-party pages: directories, partner pages, guest pieces, review platforms. It is slow, which is why it starts now.
- Internal linking. Link every new page from the hub, two siblings and the relevant service page, with descriptive anchors. Orphaned guides get cited late.
- Write the day-90 report against the day-0 baseline: citations, cited pages, query spread, the prompt sample, and which hypotheses won. It is the brief for the next quarter.
Our GEO strategy guide for Singapore brands covers where Copilot, Gemini and AI Overviews behave differently.
Week-by-week checklist
Weeks are indicative; the sequence is not.
| Week | Phase | Do | Measure |
|---|---|---|---|
| 0 | Baseline | Record citations, cited pages, grounding queries; 10 buying questions; entity check | Day-0 snapshot saved |
| 1 | Foundation | Canonical entity description; Organization and Article schema validated | Zero schema errors |
| 2 | Foundation | Robots, sitemap, Bing property, llms.txt; fix broken or JS-only pages | Key pages indexed in Bing |
| 3 | Foundation | Choose three intent clusters; draft four answer-first pages | Drafts match the buying questions |
| 4 | Foundation | Publish the four pages; request indexing; start the weekly scoreboard | First weekly row filled |
| 5 | Test | Hold the site stable; read first citations and grounding queries | Which pages were retrieved, for what |
| 6 | Test | Launch format and framing tests (one variable each) | Test pages indexed |
| 7 | Test | Launch freshness and hub-and-spoke tests | Per-page citations on the scoreboard |
| 8 | Test | Read results at three weeks; kill or merge pages never retrieved | Winners and losers named |
| 9 | Scale | Publish the winning framing in the other two clusters | Cited pages rising |
| 10 | Scale | Add adjacent-question pages and a fourth cluster | Top-query share falling |
| 11 | Scale | Third-party corroboration outreach; internal links from hub, siblings, service page | New referring pages live |
| 12–13 | Report | Re-run the 10 buying questions; write the day-90 report against the baseline | Citations, cited pages, spread, hypotheses won |
Takeaway: four weeks of engineering and writing, four of patience and reading, five of deliberate repetition.
Which KPIs matter, and which ones mislead?
Four numbers tell you whether the roadmap is working; one popular number tells you almost nothing in 90 days.
- Citations (Bing AI Performance): the headline count. Useful for direction, too noisy to judge a week on.
- Cited pages: distinct pages earning citations. The healthiest growth metric: it shows the programme, not one lucky page, is working.
- Share of grounding queries: how many distinct queries you are cited for, and your citation share within the ones that matter. Bing's citation-share metric, previewed in April 2026, makes the second half measurable.
- Concentration: the share of citations in your single biggest query. High and rising is fragility, even when the total looks superb.
The number that misleads is traffic from AI referrers. Several assistants pass no referrer, so GA4 undercounts it, and it lags citations by weeks or months. Report it as a floor. If leadership wants one chart, give them cited pages over 13 weeks, with the honest caveat that citations are visibility, not leads. Our guide to the best AEO tools in 2026 covers the point at which paid trackers add something the free reports do not.
What AI Studio saw in its own data
We ran a version of this roadmap on our own site in mid-2026. Two caveats: these are Bing Copilot citations as reported in Bing Webmaster Tools, not traffic or leads, and they are one site in one market.
- Citations went from roughly 34 a day to roughly 2,066 a day within 24 hours of publishing eight guide pages on 8 July 2026: a step change, not a curve.
- July 2026 closed at about 37.6K citations against 754 in June, roughly fifty-fold, and the quarter totalled 38,573.
- The uncomfortable finding was concentration: a large majority of those citations sat in a single query cluster, which is why "broaden the query base" is a named task in days 61–90. A one-day reporting artefact that looked like a ranking loss is why the roadmap judges on three-week windows.
The full write-up is in what 38,000 AI citations taught us. If you would rather have it run for you, our GEO Singapore and AEO agency Singapore pages describe how we scope it and what we report at day 90.
Frequently Asked Questions
Is 90 days long enough to see AI citations move?
Yes, if the foundation is done in the first month and the site is then left stable long enough to be re-crawled. In our own data the first citations appeared within days of publishing a batch of guide pages. What 90 days will not give you is a settled trend, so judge the quarter on cited pages and query spread, not a single peak.
Do we need paid AI-visibility tools to run this roadmap?
No. The roadmap runs on Bing Webmaster Tools' AI Performance report, Search Console's generative-AI views where available, and a spreadsheet of prompts you test by hand. Paid trackers earn their keep when you need daily multi-engine sampling across many prompts, which is usually a day-61 decision rather than a day-one one.
What if our citations are flat after the first 30 days?
Treat it as diagnostic rather than failure. First confirm the new pages are indexed in Bing and Google. Then check the grounding-query report: if your pages are retrieved for queries that do not match their intent, the framing is wrong. If nothing is retrieved at all, the problem is usually crawlability or entity clarity, and days 31–45 should fix that before any new content goes out.
How is this different from the AEO Experimentation Roadmap?
The AEO Experimentation Roadmap is about experiment design: how marketers form hypotheses, control variables and read conversion outcomes. This roadmap is the calendar for a GEO programme whose unit of measurement is the citation. Use that one to design a clean test; use this one to decide what to do in each of the 13 weeks and what to report at day 90.
Can we run this alongside our existing SEO programme, or does it compete?
It complements it. The entity, schema and crawlability work in days 1–30 is the same work a careful SEO team does, and answer-first guide pages tend to earn organic rankings as well as citations. The one real conflict is cadence: GEO testing needs the site held stable during a read window, so schedule redesigns, migrations and mass rewrites around the test calendar.
Should we measure traffic from ChatGPT and Perplexity instead of citations?
Not as the primary KPI. AI referral traffic is undercounted because several assistants pass no referrer, and it lags citations by weeks or months. It is a useful secondary signal, and a lead that names an assistant as its source is worth recording. But the number a roadmap can actually move within 90 days is citations, so build the scoreboard on that.
Related reading
- GEO Singapore — generative engine optimisation services
- AEO agency Singapore
- GEO strategy guide for Singapore brands in 2026
- The AEO experimentation roadmap: designing clean AI-search tests
- How to track AI citations in Bing, Search Console and GA4
- What 38,000 AI citations taught us about GEO
- The best AEO tools in 2026, and when you need them
Want to be the answer, not just a search result?
AI Studio builds AI search visibility for Singapore brands — entity, schema, and the content that assistants actually cite. Start with a free AI Visibility Audit of your own site.