ProBackend
a b testing seo
5 days ago6 min read

Proving AI Search Impact: Split Testing, Control Groups, and Real Client Results

A practical guide to proving AI search optimization works through split testing, control groups, and real client test results from seoClarity's enterprise clients. Covers Google's new Search Console AI reports, golden prompt sets, FAQ causation tests, and key Q&A insights.

Stop Guessing Whether AI Search Works. Test It.

Visibility scores tell you if you showed up. They don't tell you whether what you actually changed made a difference. That's the gap a lot of AI search teams are sitting in right now—and it's the gap that split testing closes.

A recent SEJ webinar with seoClarity's Mark Traphagen, Mihir Naik, and Suraj Lalchandani laid out a practical testing framework that's actually being run by enterprise clients. The core argument is simple: page-level performance and split testing tell you if what you did mattered. Visibility scores don't.

Here's how to build that proof.

Google's New Search Console AI Reports

On June 3, 2026, Google launched dedicated Search Console reports for AI Overviews and AI Mode. For the first time, sites can see page-by-page how often each URL appears inside Google's AI search features.

Lalchandani called it the biggest measurement upgrade AI search testing has received. "This has been the hardest thing to measure in AI search. Everyone was sampling. Everyone was inferring. But now Google is just giving it to you."

First-party data from Google carries a trust level third-party tools simply can't match. For more on how these reports work under the hood, see our guide to how Google counts impressions in Search Console's AI report. But the seoClarity team was blunt about the limitations: these reports cover only part of what an AI search testing program needs. ChatGPT, Claude, and Perplexity still require structured third-party tracking.

Action step: Check Search Console for the new AI reports. Map where first-party data fits your testing program before you build around it.

Building a Golden Set of Prompts

Not every prompt deserves testing. The seoClarity team starts by building a golden set of prompts spanning the full AI search funnel—awareness through retention. Every prompt gets tagged by stage, then sorted into tiers based on where the brand currently stands in the AI's response.

Tier 1 prompts are the easy wins. As Lalchandani put it: "You're relevant, but AI just hasn't been given a URL worth linking to."

Tier 2 is the heavier lift. One bucket of prompts actually gets dropped from testing entirely—a move that surprised many attendees. The sequencing is deliberate: early wins buy the political capital to run harder tests later.

Each prompt pairs with the exact page you want cited. That tracking unit is the foundation for everything that follows.

How to Split Test on an LLM

You can't split live traffic 50-50 on an LLM. So you build a control group instead: a set of correlated pages that acts as your noise filter against model updates and algorithmic shifts.

"Lalchandani put it this way: "Without a control group, every result would be guesswork. With one, you can tell a real win from the background noise."

Timing is the discipline most teams skip. The methodology sets a specific baseline period before any change goes live and a minimum test window after. AI search doesn't respond overnight the way traditional SEO sometimes does. Cut the window short and, in Lalchandani's words, "you could be reading noise."

For a broader look at how enterprise SEO teams structure these tests across multiple platforms, check our enterprise AI search testing methodology guide. Every test lands in one of three outcomes. Each one tells you something about your hypothesis.

The FAQ Test That Proved Causation

seoClarity ran the same methodology for three clients and got three very different outcomes. Which is exactly the point.

The FAQ test was the clear win. With roughly 1,000 prompts under measurement, adding FAQ sections to test pages pushed citations up versus control—and they stayed elevated as long as the change was live. Then the team reverted the change.

"The citations fell back down. That's the second half of proof," Lalchandani said. "Not that citations just went up when we added FAQs, but that they went back down when we took them away. That's causation, not correlation."

That reversion is the difference between correlation and causation. And almost no team measuring AI search today can produce it.

The other two tests—one on meta descriptions and one on listicle formatting—ended differently. The reasons why hold lessons for anyone about to invest in either tactic.

Naik's framing: every result is a win, because you have evidence instead of guesses. That's more than most teams in AI search have today.

The Q&A You Actually Need to Hear

Measuring AI Authority Without a Clean Metric

"AI authority is basically how much the model trusts you as a source for this topic," Lalchandani said. "I don't think there's a clean number for it, but there are a couple of signals you can stack."

He named four stackable signals, starting with citation share on your top prompts and cross-engine consistency. "Consistency across engines just means that you become the authoritative source in your category for specific kinds of questions."

Can AI Bots Read Collapsible FAQs?

"Collapsible can mean many different things. It's how you're having it collapsible."

Some setups keep collapsed FAQs fully readable to AI search engines and Google. Others make the content invisible to both—"even Google will not click around on your site," Lalchandani noted. His standing advice: "If you're unsure of something, just test it out. It takes effort, but it'll give you a sure answer." For practical steps on auditing what AI search engines can and cannot see on your pages, see our guide to auditing and correcting AI search answers about your business.

What's the ROI of a Citation That Doesn't Drive Traffic?

"You want to be cited because you are controlling the answer that is actually going to be showing up," Naik explained.

Even without a click, your cited page shapes the narrative inside the answer. This matters especially in comparison queries where citations do the heavy work of positioning both brands. The question shifts from traffic to representation: are your USPs highlighted correctly, is the comparison set right, are inaccuracies surfacing?

Lalchandani shared a cautionary example from a real restaurant client showing exactly what happens when AI cannot reach your content.

Is Traditional SEO Still a Factor?

"Absolutely. It is foundational. It is the foundation," Traphagen said.

seoClarity's longest-standing clients—the ones with well-optimized content and technically healthy sites—are also performing best in AI search. AI optimization is the extra layer on top. Lalchandani added: "When we run tests with our clients, we've rarely, if ever, found a situation where something works for SEO and does not work for AI search."

What's Next

The full webinar recording contains everything the recap holds back: the golden prompt set build, the tier definitions, the control group construction with exact baseline and test windows, the platform-by-platform crawler reference, the meta description and listicle results, and the schema and markdown test blueprints.

Register to watch the full session on demand at Search Engine Journal.


Sources: Search Engine Journal webinar recap "AI Search is Working: How to Prove It With Real Tests" by Heather Campbell, July 23, 2026. Speakers: Mark Traphagen (VP Product Marketing & Training, seoClarity), Mihir Naik (Senior Product Manager, AI, seoClarity), Suraj Lalchandani (Sr. IT Project Manager, seoClarity).

Stop Guessing Whether AI Search Works. Test It

More blogs