← All posts

How to Choose an AI Visibility Agency, According to the Category's Own Data

TODD PIECHOWSKI · AUG 06, 2026 · 6 MIN READ

62.8% of what ChatGPT reads to answer who should I hire sits on vendor sites
62.8% of what ChatGPT reads to answer who should I hire sits on vendor sites

Every agency in this category will tell you they’re the best one. Including us. That’s not much help if you’re the person who has to write the check.

So instead of another pitch, here’s what we measured.

We ran 16 buyer questions through ChatGPT™ five times each — the questions a brand actually types when their Google traffic is sliding and they don’t know who to call. 86 answers came back with 1,258 citations behind them. Then we sorted every one of those citations by who owns the page.

62.8% of them sit on the marketing sites of agencies and tools competing for the job. 295 different vendor domains. Directories — Clutch and the like — were 4.1%.

So when you ask an AI which agency to hire, most of what it reads was written by agencies describing themselves. The “Top 11” list you’re reading was very likely published by one of the eleven. We checked one at random: the publisher ranked itself first.

That changes what a useful evaluation looks like. Here’s what we’d ask.

Ask for their own number

Any firm selling AI visibility can run the measurement on themselves in an afternoon. Ask what it says.

We did it on ourselves and published it: on those same 16 buyer questions, Vektor10 appeared in zero of 86 answers. It’s an embarrassing number and it’s a real one, and you now know more about our measurement than you do about a firm that only shows you charts of other people’s brands.

You’re not listening for a good number. You want a specific one, with a denominator, that they didn’t have to think about before giving you.

Ask how many times they asked

This is the one that separates measurement from a screenshot.

Ask ChatGPT™ the same question five times and you get five different answers. Not slightly different — different companies named. A vendor showing you a single result is showing you one roll of the dice. If they can’t tell you their repeat count, there’s nothing underneath the chart.

Ours is five per question, minimum. Below that the noise is bigger than the finding.

Ask which questions, and who wrote them

Visibility is meaningless without knowing what was asked. A firm that reports “you’re at 40% visibility” and can’t hand you the question list is reporting on a set they chose, which they can quietly tune until the number looks good.

Also ask how the questions were generated. Automated category generators tend to produce shopping queries — fine for a product, useless for a service. Ours are hand-written per account, and we’ll show you the file.

Ask what happens on brand-name questions versus discovery

This one gets used on you, so it’s worth knowing.

Ask ChatGPT™ about a company by name and the tooling will almost always report a hit. In our own audit we scored 41 out of 41 on brand-name questions, and when we read the answers, only 12 of the 41 described our actual company — 18 described a textile consultancy with a similar name. A brand-name score can be a perfect 100% and still mean the model has no idea who you are.

The questions that matter are the ones where nobody’s name appears — “who should I hire,” “what are the best options for a brand like mine.” That’s where we scored zero, and that’s where you should demand a vendor’s number too. If they show you a strong result and every question in the set contains their client’s brand name, you’re looking at a report that can’t fail.

Ask who they’d be competing against, by name

We found 133 firms getting named on those 16 questions. The most-named one appears in 24.4% of answers. Nobody is close to owning this.

That matters for your expectations. In a category with no incumbent, an agency promising you the top spot is promising something that doesn’t exist yet. The realistic goal is being in the set of names that comes back, and that’s a different, more achievable piece of work.

If a firm can’t tell you who else surfaces on your category’s questions, they haven’t run your category.

The conflict nobody in this category likes to mention

Most firms here do the optimization work and also report on whether it worked. That includes us.

It’s a real tension and I’d rather say it out loud than pretend we’ve solved it. What we do about it: the question set is fixed and handed to the client at the start, so it can’t be tuned mid-engagement; the raw answers and citations go to the client, not just the summary; and we report the questions where we lost. If you’re already working with an execution vendor, an independent read from someone who isn’t grading their own work is a fair thing to want — sometimes that’s us, sometimes it should be somebody else.

Ask any firm you’re evaluating how they handle it. “We’re objective” isn’t an answer.

Red flags, short list

Rankings. ChatGPT™ doesn’t have positions the way a search results page does. A firm reporting your “rank” is porting an SEO frame onto a surface that doesn’t work that way.

One engine described as all of AI. We read ChatGPT™ and we say so on every deliverable. A vendor claiming full coverage of ChatGPT, Claude, Gemini, Copilot and Perplexity is either running a much bigger operation than the price suggests, or estimating.

Blog volume as the plan. Getting a product recommended is mostly a catalog problem — feed accuracy, price, structured data, merchant enrollment. Publishing helps, and it isn’t the lever for a product brand.

No denominator anywhere. “Visibility up 30%” from what, out of how many, asked how many times.

A guarantee. Nobody controls this surface. Anyone promising a placement is describing a system they don’t have access to.

Is it worth paying for at all

Sometimes no. If you’re under a few million in revenue and your category gets few AI shopping queries, the honest answer is to run the free measurement, watch the number, and spend the money on something with a shorter payback.

The case for doing it now is that the category is empty. 133 firms splitting a set of answers, none above a quarter, and most of the pages the models read were written in the last twelve months. Markets don’t stay that open for long.

Which is exactly why we published our own zero rather than waiting until it looked better.


Method: 16 hand-written discovery questions, five repeats each, ChatGPT™ only, probed 5–6 August 2026. 86 answers, 1,258 citations, 364 domains. Page ownership classified per domain. Firm shares are the percentage of the 86 answers naming each firm. One engine, one point in time.