[FAQ] Answer Engine Optimization

What are the biggest AEO mistakes to avoid?

Written by Kevin Barber | Jul 17, 2026 7:53:28 PM

The most common AEO mistakes are adding markup and files that engines ignore, publishing generic content at volume, and treating your own site as the whole program when most citations come from third parties. Google states that it ignores llms.txt for Search, that no special AI schema type exists, and that content need not be cut into artificial chunks. Measurement is the fourth, since most AI brand exposure never produces a click for analytics to record.

Believing in special AI markup

Google's AI optimization guidance is unusually blunt about this: it ignores llms.txt for Search, there is no special AI schema type, and content does not need to be pre-cut into artificial "AI chunks" for a model to use it. Google also states plainly that structured data isn't required for generative AI search and that overfocusing on it is a mistake, while its structured data policies note that correct markup never guarantees appearance or ranking.

We still ship valid JSON-LD on every build, because it keeps brand facts, services, authorship, and relationships consistent as a site grows, and that consistency is genuinely useful to an engine trying to work out who you are. It functions as a clarity layer sitting on top of good crawlable answers. Where programs go wrong is treating it as the whole program, running an audit that produces a schema checklist and then declaring the AEO work done.

Publishing at volume without adding evidence

Ahrefs studied 75,000 brands and their AI visibility correlations and found that site page count barely correlated with AI visibility at all, around 0.194. YouTube mentions had the strongest tested correlation at roughly 0.737, and branded web mentions landed between 0.66 and 0.71. Page count sits near the bottom of the tested signals, and it happens to be the one that AI writing tools make easiest to scale.

The GEO research from KDD 2024 points in a more useful direction. Its experiments found that expert quotations, statistics, and cited sources measurably improved visibility inside generated answers (GEO paper). Content that contributes evidence gives a model something specific to pull, while content that restates the existing consensus is interchangeable with the forty other pages saying the same thing, and the model has little reason to reach for yours.

Optimizing only the pages you own

Owned-only programs are the most expensive mistake on this list. Omniscient's research found owned content accounted for only 23% of citations on branded queries, with reviews and social proof accounting for 57% (Omniscient research). That leaves a program working on roughly a quarter of the evidence pool while the reviews, comparisons, and community threads supplying the rest go unattended, which is why off-site authority work sits alongside on-site structure in everything we build.

HubSpot's own program is the clearest counterexample. It ran on-site content, off-site publisher amplification, and community growth together, ending 2025 with partnerships on nearly 1,000 third-party pages and Reddit citations that grew from 178 to 146,000 between May and December, alongside 433% more citations overall (HubSpot AEO case study). One caution before anyone tries to shortcut that: Google explicitly warns against inauthentic mentions, so the durable version asks customers who actually used the product to review it, puts your experts' names only on articles they actually wrote, and shows up in the communities your team would be in anyway. Bought mentions carry a liability that tends to surface later.

Measuring AEO with a traffic report

Ahrefs found AI assistants linked to its site in only 28% of the brand mentions it tracked (citations versus impressions study). The other mentions still happened, they just left no trace in any analytics tool, so a team grading AEO on session counts is grading the program on the small fraction of its output that happened to be clickable, and then concluding not much is going on.

Pew's data on click behavior reinforces it: users clicked a traditional result in 8% of visits where an AI summary appeared, versus 15% where it didn't (Pew Research Center). Fewer clicks is simply the environment now, so we track share of voice, citation share, and how accurately the engines describe a brand on high-intent prompts, then connect that back to pipeline. Referral traffic stays in the citation reporting as one input among several, because on its own it consistently understates what the program is doing.