Back to Blog
AEOSEOGo-to-Market

Answer engine optimization: what actually works

Abhishek Singla Jul 31, 2026 12 min read

A COO of a 60-person B2B software company forwarded me a screenshot in June. He had asked ChatGPT to recommend tools in his category. Five vendors came back. He was not one of them. Two of the five were companies he had never lost a deal to, because he had never been in a deal with them.

His question was reasonable: "How much does it cost to fix this?"

My answer was less satisfying. Before you spend a euro, someone has to tell you which of the twenty tactics circulating under the label "answer engine optimization" have any evidence behind them. I went and looked. Most of them do not. Two do, and one of those two is just SEO wearing a new hat.

This is the honest version of that answer.

What answer engine optimization actually is

Answer engine optimization, or AEO, is the practice of getting your company mentioned and cited inside AI-generated answers. Same idea shows up as GEO, generative engine optimization, or LLMO. The names are marketing. The problem is real: a buyer types a question into ChatGPT, Perplexity, Claude, or Google AI Mode, gets a synthesised answer with a handful of vendors named in it, and never sees a list of ten blue links where you were sitting at position four.

The B2B numbers are the part that should get a founder's attention. G2 surveyed 1,076 B2B decision makers across North America, EMEA and APAC in March 2026 for its 2026 AI Search Insight Report. Fifty-one percent said they now start software research with an AI chatbot more often than with Google. That was 29 percent eleven months earlier. Sixty-nine percent said they ended up choosing a different vendor than they originally planned because of what a chatbot told them. Thirty-three percent bought from a company they had never heard of before the chatbot named it.

That last number is the one I keep coming back to. A third of B2B purchases in that sample went to a vendor who was, from the buyer's point of view, invented by the model at the moment of asking.

51%
start research with a chatbot
33%
bought from a vendor they had never heard of
1.08%
of web sessions come from AI referrals

Hold those together, because they are in tension and nobody selling AEO software wants you to notice. Buyer behaviour has moved fast. Traffic has not. Semrush clickstream data covering October 2024 to February 2026 across 13,770 domains puts AI referrals at 1.08 percent of sessions. The same dataset shows those visitors converting at 4.4 times the rate of average organic traffic, which is the reason to care, but 1.08 percent is 1.08 percent. If someone tells you AI search is replacing your pipeline this quarter, they are selling something.

The advice everyone gives, and where it came from

I read seventeen of the top-ranking guides on this topic. Fourteen of them share the same skeleton and most of the same tactics. Answer the question in the first 40 to 60 words. Use question-shaped H2s. Add FAQ schema. Publish an llms.txt file. Get on Reddit. Refresh quarterly.

Not one of the seventeen cited a source for the 40-to-60-word rule. It is folklore that spread by copy and paste.

The other recycled statistic in the category, "adding statistics and quotes lifts AI visibility 30 to 40 percent," traces back to a 2023 paper by Aggarwal et al.. Real research, but it ran on GEO-bench, a synthetic generative engine the authors built. Not ChatGPT. Not Perplexity. The data predates every model your buyers are actually using. It is nearly three years old and nobody repeating it says so.

Then there is the freshness contradiction, which is my favourite. One widely quoted 2026 report claims 83 percent of AI citations come from pages updated in the last twelve months. Ahrefs analysed 17 million citations and found the average cited URL is 1,064 days old. Roughly three years. I have seen both numbers quoted in the same article, two paragraphs apart, by someone who clearly did not read either.

The point

Almost every AEO tactic you have been sold has no control group behind it.

Two do. Search rank, and getting your story published on someone else's domain. Everything else is either unproven, untested, or actively contradicted by the people running the experiments.

The tactics with evidence against them

This is the section that will save you money, so I am putting it before the playbook.

Schema markup does not earn AI citations

Ahrefs ran the cleanest test I have seen on this. May 2026. They screened 6 million URLs, found 1,885 pages that added JSON-LD structured data, matched each to roughly three control pages with similar existing citation levels, and measured 30 days before and after. Result: AI Overview citations went down 4.6 percent. AI Mode and ChatGPT moved 2.4 and 2.2 percent, both inside the noise. The write-up is here.

Honest caveat that Ahrefs states themselves: the study only covered pages already receiving 100-plus AI Overview citations. It does not prove schema fails to help an invisible page. But it does kill the claim that schema is the lever.

A separate 2026 reanalysis by Kurt Fischman found the schema association collapsed to null once control-set construction was corrected. Pages with attribute-rich Product and Review markup cited at 61.7 percent; pages with no schema at all, 59.8 percent; p equals 0.71. Meanwhile Otterly ran a neat experiment: they planted facts that existed only inside FAQ schema and asked the models about them. No platform extracted any of it. The systems read visible HTML at fetch time.

Every study claiming schema works shares one flaw. None of them control for search rank, which is the dominant predictor. As Fischman puts it, any schema study that ignores rank is mostly measuring rank.

Google's own documentation, updated 10 July 2026, is unusually direct about this: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add."

Keep your schema. It still does things for rich results in classic search. Just stop paying an agency to add FAQPage markup as an AI visibility play.

llms.txt is dead

Ahrefs checked 137,210 domains on their Web Analytics product. Twenty-eight percent had published a valid llms.txt file. Ninety-seven percent of those files received zero requests. Across every file in the study there were about 22,000 total fetches, 96 percent from bots, of which SEO audit tools accounted for 21.7 percent and actual AI retrieval bots 1.1 percent.

The single largest fetcher of llms.txt files was Claude Code. A coding agent, reading them as developer documentation. That is what the file is for now.

Zero AI bots ever probed for a missing file. Google's John Mueller has compared it to the keywords meta tag; Gary Illyes said in July 2025 that Google does not support it and has no plans to. Cyrus Shepard's meta-analysis of 54 experiments, patents and official statements (published May 2026) scored llms.txt 2.0 out of 10, dead last of 23 factors.

It takes ten minutes to publish. Fine, publish one. Just do not put it on a slide.

Planting Reddit threads backfires

Reddit removes roughly 25,000 spammy posts and comments a day and revokes about 2 million inauthentic votes daily. There is now documented reporting on AI-generated posts planted specifically to manipulate LLM citations.

The deeper problem is that it does not work even when it survives. Models cite threads with detailed lived experience, and positive and negative sentiment get cited at close to equal rates. A scrubbed, uniformly positive thread is the wrong shape for the thing you are trying to game.

The two tactics that hold up

Search rank, still

In Shepard's meta-analysis, search rank scored 9.4 out of 10 and query fan-out rank 9.3. URL accessibility scored 9.5, the top factor. Structured data 5.6. Domain authority 5.0.

Roughly 38 percent of AI citations come from Google's top 10 results. Which also means 62 percent do not, so ranking is necessary and not sufficient. Fan-out rank matters as much as the head term: the models decompose your buyer's question into sub-questions and go looking for answers to each one. Ranking for "how do you calculate CAC payback" matters as much as ranking for your category name.

Google's own July 2026 guidance says the quiet part plainly: "optimizing for generative AI search is optimizing for the search experience." Their words, not mine. It is still SEO.

If your content strategy is working, most of your AEO is already done.

Third-party distribution, which nobody wants to hear

Stacker ran the one properly controlled experiment in this whole category. March 2026. Eighty-seven stories across 30 brands, 2,600-plus prompts, 8 platforms.

Brand-domain-only citation rate: 7.6 percent. With third-party publisher distribution: roughly 27 percent median. Median lift 239 percent, p below 0.006. And here is the number that should reorganise your marketing budget: 64 percent of all citations landed on the publisher's version of the story, not the brand's own site.

Muck Rack's analysis of more than 25 million links across ChatGPT, Claude and Gemini points the same direction, with earned media making up the large majority of citations and paid or advertorial content at 0.3 percent. Their "earned media" definition is loose and I would not quote the headline percentage, but the direction is not in dispute.

Ahrefs found branded web mentions correlate with AI Overview appearance at 0.664, against 0.527 for branded anchor text and 0.392 for branded search volume. Mentions beat links.

Which means the highest-return AEO work is not on your website. It is PR, analyst relations, podcast appearances, guest research, and getting into other people's listicles.

Where most AEO budget goes
FAQ schema on every page
An llms.txt file
Rewriting intros to 50 words
A $500 a month prompt tracker
Planted Reddit and Quora threads
Where the evidence points
Ranking for the sub-questions, not just the category
Getting your data published on someone else's domain
G2 and Capterra reviews, refreshed
Comparison and alternatives pages you own
Checking your WAF is not blocking GPTBot

The B2B specifics nobody covers

Everything above applies to any company. Four things are different in B2B and I have not seen them written up properly anywhere.

Review sites carry disproportionate weight. In the G2 study, 45 percent of buyers named review-site citations as the single most confidence-inspiring signal inside an AI answer, rising to 50 percent among heavy chatbot users. Review sites ranked second in shortlist influence at 43 percent, behind chatbots themselves at 54 percent and ahead of analyst firms at 36 percent. Peec AI's analysis of 30 million citation sources found G2 shows up meaningfully in Perplexity specifically. So the boring quarterly job of asking twelve happy customers for a G2 review is now an AI visibility tactic. I did not expect that either.

Competitor-shaped prompts are a third of the opening move. G2's breakdown of first prompts: category-based 33 percent, competitor-based 31 percent, requirements-based 22 percent. Nearly a third of your buyers open with a competitor's name. If you have no page that honestly compares you to them, you are absent from that entire branch of the conversation. Ahrefs' page-type data supports this from the other side: "best" pages pull 7.06 percent of AI traffic and "vs" comparison pages 4.88 percent.

Claude matters more than its market share suggests. In one B2B referral dataset, Claude went from 1.4 percent to 18.5 percent of AI referrals in about eight months while ChatGPT fell from 89 to 63 percent. G2 found Claude usage highest among engineering and R&D roles. If you sell to technical buyers, tracking ChatGPT only is measuring the wrong room.

Your buying committee has six to eight people and they prompt differently. The champion asks for features. The economic buyer asks about pricing and risk. Security asks about SOC 2. Each of those is a separate prompt hitting a separate part of your content. This is the same buying committee mapping problem as the rest of RevOps, applied to a new surface.

Measurement, and why most dashboards are noise

Here is the part that made me stop recommending most tools in this category.

SparkToro ran 2,961 query runs with about 600 volunteers across ChatGPT, Claude and Google AI. Fewer than 1 in 100 runs returned the same brand list. Fewer than 1 in 1,000 returned it in the same order. Semrush's own AI Visibility Index, tracking 2,500 prompts, shows 40 to 60 percent of cited sources changing month to month.

So when your tool reports you moved from 4th to 2nd, that is very probably nothing.

The constructive half of the SparkToro finding is the one vendors under-quote: top brands still appeared in 55 to 77 percent of responses regardless of phrasing. Rank order is noise. A stable consideration set is real. Measure whether you appear at all, across many runs, and ignore position entirely.

There is one genuinely useful development. Google Search Console launched generative AI performance reports on 3 June 2026, expanded 23 June, covering AI Overviews, AI Mode and Discover AI features. It gives you impressions, pages, countries and devices. It gives you no clicks, no CTR and no query data, and it does not backfill before 18 May 2026. It is free, it is first-party, and there is nothing equivalent from OpenAI, Anthropic or Perplexity.

Also worth knowing before you build a longitudinal chart: when GPT-5 shipped in September 2025, platform-wide citation volume dropped and reported brand visibility fell across every tool. Not because anything about those brands changed. Every trend line in this category has that discontinuity in it and no vendor annotates it.

If you want a tool anyway, the honest tiers are roughly: free (Google Search Console plus HubSpot's AEO Grader), cheap (Otterly at $29 a month, Rankscale from $20, HubSpot AEO at $50 for 25 prompts), and real budget (Profound, Peec AI from just under €100, Semrush's add-on at $99 a domain). Full disclosure, I work at Peec AI, so weigh that. I would still start with Search Console and a spreadsheet for the first quarter.

Step 01
Check access
Confirm GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot are not blocked by your WAF or robots.txt. This is the top-scoring factor and it is broken more often than you would think.
Step 02
Baseline honestly
Turn on the GSC generative AI report. Write 20 prompts your buyers would really type and run each 10 times. Record presence, not position.
Step 03
Fix the obvious gaps
Build the comparison and alternatives pages you are missing. Refresh G2 and Capterra. Answer the sub-questions, not just the category term.
Step 04
Go earn media
Publish original data. Get it onto third-party domains. This is the only tactic with a controlled experiment behind it, and it is the slowest one.

What I would actually do with a €5k budget

Nothing on this list is exciting, which is roughly the point.

Spend zero on tooling for the first 90 days. Search Console plus a spreadsheet of 20 prompts run manually gets you a baseline that is more honest than most paid dashboards, because you will see the variance yourself instead of having it averaged away.

Spend the first €500 on an access audit. Somebody in your organisation turned on bot protection and did not read the list. I have found blocked AI crawlers at three companies this year, twice inside Cloudflare rules nobody remembered setting.

Spend €1,500 on comparison content. One page per serious competitor, written honestly enough that a buyer who reads it trusts you. Thirty-one percent of opening prompts are competitor-shaped and most companies have nothing to say into them.

Spend €500 on reviews. Ask twelve customers, get six. Repeat quarterly. Boring, cheap, and it is what 45 percent of your buyers say they trust most inside an AI answer.

Spend the remaining €2,500 on one piece of original research with a real number in it, and on getting that number published somewhere that is not your blog. That is the Stacker finding applied. Sixty-four percent of citations went to the publisher's version, so pitch the data, do not just post it.

And skip the schema rewrite, the llms.txt file, and the agency retainer to reformat your intros to 50 words. If someone quotes you a price for those three things as an AEO package, you now know what the evidence says.

Not sure whether AI search is costing you deals?

We run a free 30-minute audit: where you show up in AI answers today, which of your competitors show up instead, and the three fixes worth doing first.

Book an audit →

FAQ

Is AEO different from SEO?

Mostly no, and the people telling you it is a separate discipline are usually selling a separate product. Search rank scored 9.4 out of 10 as a citation factor in the largest meta-analysis available; structured data scored 5.6 and llms.txt 2.0. Google's own July 2026 guidance says optimising for generative search is still SEO. The genuinely new part is that third-party publication now outperforms your own domain, which pushes budget toward PR and away from on-page work.

How much traffic should I expect from AI search?

Roughly 1 percent of sessions today, based on Semrush clickstream data across 13,770 domains. The case for caring is conversion rate, not volume: those visitors convert at about 4.4 times the average organic rate in the same dataset. Also expect 15 to 35 percent of AI-driven traffic to land in GA4 as "direct" because several platforms strip the referrer, so your real number is higher than your analytics says. Read more on that problem in our guide to B2B marketing attribution.

Do I need to add FAQ schema to get cited?

No. Ahrefs' controlled test on 1,885 pages found AI Overview citations went down 4.6 percent after adding JSON-LD, with other platforms flat. Otterly showed models do not extract facts that exist only inside schema. Google states directly that no special markup is needed. Keep schema for classic rich results, drop it as an AI tactic.

Which AI platform should a B2B company track first?

ChatGPT still carries roughly 62 percent of B2B research use, so start there. But Claude went from 1.4 to 18.5 percent of B2B AI referrals in about eight months and skews heavily toward engineering and technical buyers, so add it early if that is your ICP. Perplexity matters if review sites are part of your story, since G2 citations show up there specifically.

How do I know whether my AEO work is doing anything?

Measure presence across many runs, never position. SparkToro found fewer than 1 in 1,000 repeated queries return brands in the same order, so rank movement in your dashboard is almost always noise. Run each prompt at least ten times, track the percentage of runs where you appear, and give it a full quarter. Annotate your chart when a major model version ships, because that alone moves everyone's numbers.


We build the CRM and RevOps systems that let you tell whether any of this is working, and the automation that keeps the data flowing without a human copying rows between tools. If AI search is showing your competitors and not you, tell us what you sell and we will tell you what we would fix first.

Second opinion

Wrestling with something like this in your own stack?

Describe the whole problem to us, in total privacy. Within 7 days you get our second opinion in writing: what is actually going on, how we would tackle it, and what we would avoid. We take on a limited number of questions each month.

Ask privately