How Does AEO Work?

When does AEO fail or not make sense?

AEO fails when an engine has no verifiable evidence to retrieve about a company, and it makes little sense when buyers in the category aren't using AI to research the purchase. Technical access is the other common blocker, since a site that refuses AI crawlers or renders its text through scripts can't be quoted, however good that writing may be. An engine's answer is bounded by the proof already published somewhere it can reach.

When your buyers aren't asking AI in the first place

The demand test comes before everything else. In B2B software, the behavior is well documented: 6sense's 2025 Buyer Experience Report found 94% of B2B buyers used LLMs somewhere in the process, and 95% ultimately bought from a vendor that was already on their Day One shortlist. That combination is what makes AEO worth funding, because the shortlist is being assembled inside a chat window before anyone talks to sales.

Plenty of categories don't work that way. If your revenue comes through a referral network or a small set of accounts you already know by name, there's no AI assistant sitting between you and the buyer, and the same budget usually does more in outbound or in sharpening the offer itself.

There's a volume question underneath this too. Semrush's 50,000-site channel study found AI traffic grew 66% in 2025 but still accounted for less than 0.15% of total visits. So a plan that needs a step change in session volume this quarter is unlikely to get it here, because the position AEO buys you inside a recommendation accrues over several quarters, and the first thing you tend to notice is that the visitors who do arrive are further along in their evaluation, which shows up in your funnel long before it shows up in your session count.

When the source material isn't ready yet

An engine can only repeat what the market has already said about you somewhere it can retrieve. Omniscient Digital's branded-query research found owned content supplied only 23% of AI citations, while reviews and other social proof supplied 57%, which means the majority of the evidence pool sits on properties you don't control. If you have a handful of reviews, no published case studies, and no subject matter expert who'll go on record, that majority is empty, and on-site publishing can't fill it for you.

That's a normal place for an early company to be, and it reads as a sequencing question about when to start rather than a judgment on the business. Get four customers to write honest reviews, publish two case studies with real numbers in them, and the same AEO program lands very differently ninety days later.

Positioning has the same dependency. When two people inside a company describe the product in two different ways, the engines will repeat both versions and pick whichever one retrieval happens to surface, which is how we've watched a single brand get described accurately in one engine and mischaracterized in another during the same week. Settling that story internally costs a lot less than trying to correct it later through published content, so we run the messaging exercise before the writing starts.

When engines physically can't read the site

Google states there are no additional eligibility requirements for AI Overviews or AI Mode beyond its normal Search requirements (Google's AI features guidance), which means a page that isn't indexed isn't a candidate for an answer. Meanwhile OpenAI discovers content for ChatGPT search through OAI-SearchBot (OpenAI publisher FAQ) and Perplexity separates its indexing crawler from user-requested fetches, so each engine has its own access path.

We run into this on sites that have been performing well for years, where a WAF rule or a bot-management default configured long before OAI-SearchBot existed is quietly turning the newer agents away, and client-side rendering can produce the same outcome by hiding the text behind a script the crawler never executes. The bot list changed under everyone, so this is almost never a decision anyone got wrong. It's usually a two-week fix, and it has to clear before content work can do anything at all. Your own team can check most of it in an afternoon by pulling server logs to see whether OAI-SearchBot and the Perplexity crawlers are getting through, then viewing the served HTML to confirm the answer text is genuinely present rather than being injected by a script after the page loads. Whatever your bot-mitigation vendor is doing by default deserves a close read, since most of those defaults were written before these agents existed.

What we tell you if you're not ready yet

We sell this. Our AEO Authority System starts at $12k for the on-site foundation and $5k a month for off-site authority work, and we still turn engagements away, because a company whose actual bottleneck is conversion or positioning will spend that money and feel nothing. When a site converts at half the rate it should, our advice is to fix the lead generation path first and then point AEO at a page that closes, because the engine hands the visitor off at the click and everything after that belongs to the page.


How do AI model updates affect AEO results?

Model and index updates change which sources an engine retrieves, so AI citation results move even on pages that haven't been touched. Source-level citations are the most volatile layer, while brand-level visibility tends to be stickier, and each engine shifts on its own schedule. Reliable measurement depends on a fixed prompt panel run repeatedly over time, with model releases annotated so movement isn't misread as a content result.

How much the source layer actually moves

Semrush's AI visibility trend research found the average prompt-coverage change among ChatGPT's top 100 source domains was roughly 120% over three months, which is a much larger swing than most teams plan for. Reddit is the clearest single example: its coverage as a ChatGPT source fell about 82% in that window while its use in Google AI Mode rose about 74%, which is one platform moving in opposite directions on two engines at the same time.

Engine-level behavior diverges just as sharply. Semrush observed ChatGPT expanding its source diversity by around 80% in a single month while Google AI Mode barely moved, and separately found that ChatGPT and AI Mode agreed on brands 67% of the time but on sources only 30% of the time, which is why a single "AI visibility" score averaged across engines hides most of what's actually going on underneath it.

Brands move more slowly than citations

Brand presence is a slower-moving asset. In the same three-month window, 25 new brands entered the top 100, mostly in lower positions, while the top 50 stayed comparatively stable. That gap between volatile sources and stable brands is the practical argument for investing in the entity itself, since a brand the engines recognize survives the constant churn in which individual pages get cited from one month to the next.

It also explains why we weight off-site corroboration so heavily. When a model update reshuffles which domains it trusts, a brand that appears across reviews, expert bylines, comparison articles, and community discussion keeps most of its surfaces intact, while a brand whose entire footprint is its own blog is carrying a single point of failure.

How we read the data when a model ships

Our monitoring runs on a fixed prompt panel, scored the same way every time, across ChatGPT, Claude, Gemini, and Perplexity. Every prompt gets run several times, because these models are non-deterministic and a single run tells you very little. We've watched the same brand appear in two of five runs on the same prompt on the same afternoon.

Known model and index releases get dated markers on the reporting timeline, so a drop that lands on a release date can be read for what it probably is instead of being pinned on whatever we published last. Movement gets judged over 30 and 90-day windows across the panel as a whole, since individual prompt wins and losses at any smaller scale are mostly noise.

The failure mode we see most often is attribution by proximity: something moved, the last thing we did was publish a comparison page, therefore the comparison page did it. That's occasionally correct, and it's at least as likely that the engine quietly changed what it reads.

What this means for expectations

Volatility is why our public timeline is a range. Most clients see measurable citation and recommendation activity within 60 to 90 days of finishing setup, and what follows is accumulation: a few more prompts each month where an engine has a reason to name you, arriving on four different schedules because the engines don't update in step with each other. The programs that hold their position through model updates are the ones treating monitoring as ongoing work, which is why the citation dashboard and competitive positioning are built into the system and reviewed monthly with a strategist. An engine can reframe an entire category between quarters, and the dashboard is where you'd see that happen while there's still time to respond to it.


What are the biggest AEO mistakes to avoid?

The most common AEO mistakes are adding markup and files that engines ignore, publishing generic content at volume, and treating your own site as the whole program when most citations come from third parties. Google states that it ignores llms.txt for Search, that no special AI schema type exists, and that content need not be cut into artificial chunks. Measurement is the fourth, since most AI brand exposure never produces a click for analytics to record.

Believing in special AI markup

Google's AI optimization guidance is unusually blunt about this: it ignores llms.txt for Search, there is no special AI schema type, and content does not need to be pre-cut into artificial "AI chunks" for a model to use it. Google also states plainly that structured data isn't required for generative AI search and that overfocusing on it is a mistake, while its structured data policies note that correct markup never guarantees appearance or ranking.

We still ship valid JSON-LD on every build, because it keeps brand facts, services, authorship, and relationships consistent as a site grows, and that consistency is genuinely useful to an engine trying to work out who you are. It functions as a clarity layer sitting on top of good crawlable answers. Where programs go wrong is treating it as the whole program, running an audit that produces a schema checklist and then declaring the AEO work done.

Publishing at volume without adding evidence

Ahrefs studied 75,000 brands and their AI visibility correlations and found that site page count barely correlated with AI visibility at all, around 0.194. YouTube mentions had the strongest tested correlation at roughly 0.737, and branded web mentions landed between 0.66 and 0.71. Page count sits near the bottom of the tested signals, and it happens to be the one that AI writing tools make easiest to scale.

The GEO research from KDD 2024 points in a more useful direction. Its experiments found that expert quotations, statistics, and cited sources measurably improved visibility inside generated answers (GEO paper). Content that contributes evidence gives a model something specific to pull, while content that restates the existing consensus is interchangeable with the forty other pages saying the same thing, and the model has little reason to reach for yours.

Optimizing only the pages you own

Owned-only programs are the most expensive mistake on this list. Omniscient's research found owned content accounted for only 23% of citations on branded queries, with reviews and social proof accounting for 57% (Omniscient research). That leaves a program working on roughly a quarter of the evidence pool while the reviews, comparisons, and community threads supplying the rest go unattended, which is why off-site authority work sits alongside on-site structure in everything we build.

HubSpot's own program is the clearest counterexample. It ran on-site content, off-site publisher amplification, and community growth together, ending 2025 with partnerships on nearly 1,000 third-party pages and Reddit citations that grew from 178 to 146,000 between May and December, alongside 433% more citations overall (HubSpot AEO case study). One caution before anyone tries to shortcut that: Google explicitly warns against inauthentic mentions, so the durable version asks customers who actually used the product to review it, puts your experts' names only on articles they actually wrote, and shows up in the communities your team would be in anyway. Bought mentions carry a liability that tends to surface later.

Measuring AEO with a traffic report

Ahrefs found AI assistants linked to its site in only 28% of the brand mentions it tracked (citations versus impressions study). The other mentions still happened, they just left no trace in any analytics tool, so a team grading AEO on session counts is grading the program on the small fraction of its output that happened to be clickable, and then concluding not much is going on.

Pew's data on click behavior reinforces it: users clicked a traditional result in 8% of visits where an AI summary appeared, versus 15% where it didn't (Pew Research Center). Fewer clicks is simply the environment now, so we track share of voice, citation share, and how accurately the engines describe a brand on high-intent prompts, then connect that back to pipeline. Referral traffic stays in the citation reporting as one input among several, because on its own it consistently understates what the program is doing.