Every SEO agency on the market now calls itself an AI SEO agency. Some rebuilt their methodology around how large language models select and cite sources. Most changed their homepage headline and nothing else. If you are evaluating an AI SEO agency right now, your real problem is not finding candidates. It is telling the two groups apart before you sign a contract, because the difference only becomes visible months into an engagement, after the budget is spent.
This guide gives you the evaluation criteria that expose the difference early: what a genuine AI-search specialist must be able to demonstrate, the questions that collapse weak pitches, and the engagement structures that align the agency's incentives with yours. We will not give you a ranked list of vendors. Rankings in this category are mostly pay-to-play, and the right agency for a B2B SaaS company is the wrong one for a marketplace. Criteria travel. Rankings don't.
Why choosing an AI SEO agency is genuinely hard
Traditional SEO had a brutal but useful property: results were publicly verifiable. You could check any agency's claims by looking at rankings. AI search removes that transparency. Answers from ChatGPT, Perplexity, and Google's AI Overviews are assembled per query, per user, per session. Two people asking the same question can get different answers citing different sources. There is no public leaderboard to audit.
That opacity is exactly what makes the category attractive to opportunists. An agency can promise "AI visibility," run a handful of prompts, screenshot a favorable answer, and declare victory. You have no easy way to know whether that screenshot represents a repeatable pattern or a lucky draw.
The second difficulty is that AI search optimization is not a separate discipline bolted onto SEO. It is mostly SEO done to a higher standard, plus a layer of work specific to how language models retrieve and synthesize information. LLM-based engines lean heavily on retrieval: they fetch web content, often through search indexes, and select sources that are crawlable, and tend to favor those that are clearly structured, factually consistent, and corroborated elsewhere. An agency that cannot do classic SEO well cannot do AI SEO at all. That means your evaluation has to test both layers, and most buyers only think to test the shiny one. Our guide to what AI SEO services actually include breaks down that stack in detail if you want the full scope before vetting anyone.
What an AI SEO agency must demonstrate
Four capabilities separate specialists from rebrands. Ask for evidence of each one, in this order, because agencies fail them in this order.
A measurement methodology you can interrogate
This is the first filter and the most reliable one. AI search visibility is probabilistic, so any honest measurement approach has to deal with variance. A serious agency will describe something like this without being prompted:
A defined prompt set built from your actual buying journey, not generic category questions, covering comparison, alternative, "best," and problem-framing queries.
Repeated sampling across multiple runs and multiple engines, because a single run of a single prompt proves nothing either way.
Tracked dimensions beyond "mentioned or not": whether your brand is cited as a source, how it is described, which competitors appear alongside it, and which pages the engines actually cite.
A baseline before work begins, so that change can be attributed rather than asserted.
The tell is how they talk about uncertainty. Practitioners who measure this daily are quick to say what they cannot prove. Vendors who have never measured it speak in certainties. If a pitch includes a precise-sounding visibility score with no explanation of sampling, ask how the score behaves when the same prompts are re-run a day later. The silence is informative. For a deeper look at what visibility tracking should cover, see our piece on measuring AI visibility.
Entity and brand-knowledge work
Language models do not rank pages. They resolve entities: they build an internal representation of what your company is, what it does, who it serves, and what it is associated with. If that representation is thin, wrong, or entangled with a similarly named company, no amount of content will fix your AI-search presence.
A capable AI SEO agency should be able to show you, on a past client, how they:
Audited how the major engines currently describe the brand, and where those descriptions are wrong or outdated.
Established consistency across the sources models draw from: the client's own site, structured data, directory and profile pages, third-party coverage, and review platforms.
Built corroboration, because claims corroborated across independent sources are far more likely to survive into answers than claims that exist only on the brand's own domain.
This is slow, unglamorous work. Agencies that skip it are easy to spot: their case studies are all content, no entity layer.
A content system, not a content promise
AI engines synthesize answers from content that directly and clearly addresses the underlying question. Winning citations at any meaningful scale requires producing a large volume of genuinely useful, tightly structured content, and keeping it current, because stale pages lose citations to fresher corroborated sources.
The word that matters here is *system*. Anyone can write ten good pages. Ask instead:
How do they decide what to produce? You want a defensible prioritization model tied to your revenue lines, not a keyword dump.
Who writes, who reviews for factual accuracy, and how is subject-matter expertise injected? Expertise a model can verify elsewhere lowers its risk in citing you, and fact errors are disqualifying in categories where trust matters.
How does content get maintained after publication? A publish-and-forget operation loses citations to fresher, corroborated sources.
Can they show throughput on a real account: volume, cadence, and the commercial result it drove?
This is where published case studies earn their keep. We maintain more than 200 of them at Junto precisely because content-system claims are cheap and production evidence is not. One example of what a functioning system looks like: for Empruntis, a French credit brokerage, the content and SEO program produced more than 2,400 pieces of content, multiplied organic traffic by four, and cut paid search spend by 20% while Google Discover grew into a channel delivering a quarter of monthly leads. That kind of output discipline, not any single clever page, is what earns durable presence in machine-assembled answers.
Classic SEO foundations underneath everything
Retrieval-augmented engines discover content the same way search engines always have: crawling, indexing, parsing structure. Google is explicit in its Search Central documentation that AI features in Search build on the same crawling and indexing infrastructure as classic results. A site with rendering problems, weak internal linking, duplicated or contradictory pages, or broken structured data is handicapped in AI search before any "AI work" begins.
So test the boring layer. Ask the agency to walk you through a recent technical audit they delivered: what they found, what they prioritized, what moved. An SEO agency with a real technical bench will have specifics. A rebranded content shop will pivot the conversation back to AI as fast as possible. That pivot is your answer.
How to test an AI SEO agency's claims in a pitch
Criteria are only useful if you can verify them in a sales process, where everyone is on best behavior. These tests are cheap to run and hard to fake.
Run their own medicine on them. Before the pitch, ask ChatGPT and Perplexity who the strong agencies are in their market and how each candidate is described. An agency selling AI visibility that has none of its own is not automatically disqualified, but it had better have a sharp answer for why.
Ask for the methodology document, not the deck. Decks are marketing. A specialist has an internal playbook: prompt-set construction, sampling cadence, citation tracking, reporting format. If no such document exists, the methodology doesn't either.
Give them a live query. Pick a commercial question in your category, run it in an engine together, and ask them to explain the answer: why these sources were cited, what those pages have in common, what your brand would need to change to appear. Practitioners narrate this fluently because they do it every day. Pretenders generalize.
Demand a named team. AI search moves fast enough that expertise is individual, not institutional. Ask who will actually work your account and interview that person, not the sales lead.
Ask what they refuse to promise. The honest answer includes: guaranteed placement in any specific AI answer, exact traffic forecasts from AI surfaces, and overnight results. An agency that guarantees a specific ChatGPT recommendation is describing something it cannot control. Our breakdown of how ChatGPT selects and cites sources explains why that guarantee is structurally impossible.
Check the case studies against the claims. If the pitch is about AI search but every case study is generic paid media, the capability is aspirational.
One pitch meeting run this way tells you more than three rounds of RFP scoring.
Engagement structures that actually work
How the work is scoped predicts whether it will succeed. Three structures are common; each fits a different situation.
Audit-first. A fixed-scope diagnostic covering current AI visibility, entity health, content gaps, and technical foundations, ending in a prioritized roadmap. This is the right entry point when you don't yet know the size of your problem, and it doubles as a low-risk trial of the agency's actual thinking. The deliverable should be specific enough that you could, in principle, execute it with someone else. If it reads like a proposal for more work rather than a diagnosis, you learned something too.
Ongoing retainer. The default for sustained work, because entity building and content systems compound over quarters, not weeks. What matters is what the retainer commits to: insist on defined monthly deliverables and a visibility reporting cadence against the baseline, not a bucket of hours. Cost scales with content volume, the number of markets and languages, the competitiveness of your category, and how much technical remediation your site needs. Those are the drivers worth interrogating in any quote; a number without them attached is a guess.
Sprint or project. Useful for a bounded objective: fixing entity confusion after a rebrand, building out one content cluster, preparing a product launch. Weak as a primary strategy, because AI-search presence erodes without maintenance.
Whatever the structure, contract for measurement independence: you should own the prompt sets, the baseline data, and the reporting, so that switching agencies later doesn't mean starting measurement from zero.
Red flags that should end the conversation
Guaranteed placements in AI answers or "we'll get you recommended by ChatGPT" promises.
Visibility scores from a proprietary tool the agency cannot explain mechanically.
Precise market-wide statistics about AI search adoption deployed as urgency levers. The honest version of that pitch is qualitative, because the reliable numbers mostly don't exist yet.
No interest in your technical SEO or analytics access. Anyone proposing AI SEO without looking at your crawl health is selling content by another name.
A pitch that never mentions what happens in classic Google results. AI surfaces and blue links draw on the same underlying authority; an agency treating them as separate worlds misunderstands both.
The team you meet in the pitch is not the team on the account.
Where Junto fits
We built our AI search practice on top of fifteen-plus years of SEO, paid media, and data engineering, and it operates the way this article says a specialist should: baseline measurement across engines before any recommendation, entity and corroboration work alongside content, and content systems that have shipped at the scale the Empruntis program shows. Every claim we make in a pitch maps to one of our published case studies, and we will happily run the live-query test in the first meeting. If that is the standard you hold every candidate to, including us, you will end up with a good agency.
Talk it through before you commit
If you are comparing agencies right now, a short conversation will sharpen your shortlist faster than another spreadsheet: talk to our team about your market, and we will tell you what we would test first, including whether your situation needs an AI-search specialist at all.
Frequently asked questions
Is an AI SEO agency different from a regular SEO agency?
Increasingly, no good SEO agency ignores AI search, and no good AI SEO agency skips classic foundations. The label matters less than the capability set: measurement methodology for AI surfaces, entity work, a content production system, and a real technical bench. Vet for those four regardless of what the agency calls itself.
How long does AI SEO take to show results?
Longer than paid media, and on a similar horizon to classic SEO. Entity consistency and corroboration build over months, and engines re-crawl and re-synthesize on their own schedule. Any agency quoting results in weeks is either planning to cherry-pick favorable prompts or hoping you won't check.
Can an agency guarantee my brand appears in ChatGPT answers?
No. Answers are generated per query and per session, influenced by phrasing, context, and model updates the agency does not control. What a good agency can do is systematically raise the probability of being retrieved and cited, and prove that probability is rising through repeated sampling against a baseline.
Should I hire separate agencies for SEO and AI search?
Usually not. The work overlaps heavily: the same crawlable structure, authority signals, and content quality drive both. Splitting them creates coordination overhead and lets each vendor blame the other's layer. Split only if your current SEO partner has demonstrably failed the AI-specific tests in this article.

Founder and CEO of Junto
Founder & CEO of Junto, Étienne has been an entrepreneur and digital marketing consultant for over 15 years. An expert in Paid Media, SEO, Data, Automation, AI, Growth and Performance, he helps ambitious companies build high-impact growth strategies — generating lasting results and helping brands move forward in a constantly evolving digital environment.





