Ben Wills started in SEO in 2001 at KeywordRanking. Then ontolo with Garrett French. Then he left and did engineering for about 13 years – large scale crawlers and scrapers written in pure C by hand, all the way through to the embedded firmware for an ESP32 sitting inside high end audio equipment.
Now he’s back, building OppAlerts. And he brought the one thing I’ve been waiting for somebody in this industry to say out loud.
He ran one prompt and changed one word
The prompt was flying to Los Angeles to pick up a new car, give me hotel recommendations. Then he changed the car. That’s the whole experiment.
“From a Honda Civic to a BMW to like a Ferrari or something like that. And the hotel recommendations were very different.”
Ferrari got expensive hotels. Civic got cheap ones.
“It was across all three of the ChatGPT models and across all levels of thinking.”
And here’s the part I keep chewing on:
“And the more thinking there was, the more variance it was”
More reasoning, less stability. That is not what most people assume when they turn thinking up.
His takeaway, and I think it’s the sentence of the episode:
“One of the most important things for people to understand about the prompts is how one word can completely change the results, the response that you get back.”
He describes himself as “thirty-three percent is marketer, thirty-three percent is engineer, and like thirty-four percent is scientist.” That checks out.
The real unlock is that you can test it
This is why I wanted him on. Ben:
“That’s the kind of testing you can’t do with SEO. You gotta wait for the index to change and figure out how to get the personalization out of the way.”
“But with the APIs you get direct LLM access in a way that you’ve never really been able to do on the SEO side.”
Twenty years of waiting on an index refresh and guessing at how much of the result was personalization. Now there’s an API and an answer in four seconds.
I pushed back on the idea that LLMs are new at all. They aren’t. BERT was an LLM based implementation. Hummingbird had LLM based heuristics in it, multiple layers of it. This is new and magical is the dumbest world view that I’ve ever heard of. What changed isn’t that the technology showed up. What changed is that it’ll answer you directly.
Mark Williams-Cook posted about twiddlers on LinkedIn the day we recorded, and I said it on the show: Google intentionally has patents, has deployments, has programs to mess with you as an SEO. Literally one hundred percent intentionally designed. I feel honored and insulted.
Ben was more gracious about it than I was. He called it a weird cat and mouse game and pointed out that if somebody was gaming his product, he’d probably be spending money to make it hard for them too.
Ann Smarty was right, and Ben brought her up first
I didn’t prompt this. He did:
“And Ann Smarty said – I think this was last week with you – she said the fastest way to get into the LLM responses is via SEO.”
Then he used it to launch straight into fan-out queries. Even when you scrape ChatGPT results from a logged out session, that response is still hitting an index behind the scenes. Which is how he gets here:
“The search principles are gonna come right back. You gotta have good quality content. There’s gotta be a lot of it distributed across the web. You gotta have backlinks.”
That’s the show in three sentences. If backlinks and distributed quality content are still the engine, then the agencies already doing that work at scale are the ones best placed for this – it is why I keep pointing people at Matt Brooks and the crew who treat an LLM as an off-site customer support team you have to educate rather than a channel to game. Ann’s episode is here if you want the setup.
The correlation study, with his own asterisks already attached
Ben ran two analyses, one in May and one in July. The July one was 100 industries, 10 personas per industry, 1,100 personas across the board. He also analyzed his own Reddit archive and went through Wikidata and Wikipedia looking for correlations between links, keywords and mentions.
“there was a pretty decent correlation between your backlink strength, your search ranking strength, and your likelihood to show up in the LLM results.”
I asked for his caveats. He’d already written them himself, before I asked:
“There was a full page before I got to the data saying: this is correlation, not causality, this is what it means.”
“Don’t look at it as, like, this is baked into any algorithm. We’re just looking at patterns here.”
And then the part that makes him worth reading:
“So I basically treat it as: it’s not perfect, but it’s at least data. It’s at least decent enough to make a decision on.”
That’s the right posture. A strong correlation tells you where to go look in your own industry. If PageRank is the piece for budget travelers in airlines, chase that. If Reddit’s the stronger signal, chase Reddit.
AI slop was the best four minutes
I asked him about slop expecting a hot take. He handed me a fork in the road instead.
“I think the answer to that depends on where you place the accountability.”
Put it on the producer and you’re telling everybody they have to adopt your standard. Put it on the consumer and the answer collapses to stop accepting it, which doesn’t work:
“And the difficulty there is, as long as people are willing to accept it, it’s gonna keep getting produced.”
His example was a friend’s kids on iPads, and you can hear what they’re watching, and it’s all slop. Then he turned it on himself, which I did not see coming:
“are the reports that I put together, is that considered AI slop? Because I didn’t sit there, I didn’t write the code.”
He doesn’t know the answer. He says so:
“I don’t know if it’s a cultural thing that we have to shift, where we insist on higher quality. How do we get everyone to agree on that?”
“I think it’s a really difficult and fascinating conversation, and I’m really curious to see what it’s like in five years.”
For what it is worth, my friend Michael McDougald of Right Thing SEO has gone at the same question from the measurement end and argues that the forensic case that slop is detectable. Which does not settle Ben’s accountability question, but it does mean the line is at least detectable.
More people should answer questions that way.
What he’s actually building
news.oppalerts.com pulls a few million RSS and podcast feeds plus over a hundred thousand news sources into a marketing dashboard and re-ranks every fifteen minutes. The main product adds rank tracking, a historical PageRank database built on Common Crawl’s own PageRank data, contact info scraped off 10 to 20 million sites, campaign setup driven by three written personas, news specific link gap analysis, and a semantic classifier that checks your copy against over 20,000 taxonomy categories pulled from ten taxonomies.
That last one got me. We built an IBM Watson content classification feature into Raven Tools in 2012. Same idea, about a decade of maturity apart.
He also spent a week in a Slack channel with Mike King and Russ Jones tearing apart the Yandex source leak. What stuck with him was the category vectors:
“So CatBoost, I think CatBoost and BERT were both strong categorical taxonomies. And I think they were showing up in lots of different ways. So lots of different vectors per search query and per document.”
Homework
“Just open up Claude Code, tell it to build a Docker image using Manticore, tell it to have the title and the body.”
Scrape ten thousand pages. Put the title and the body in. Then figure out how to design searches that get the most relevant documents back.
“You’ll learn more about search in that hour and a half than anything else.”
I’m turning that into an SOP. Go listen to the whole episode.
Where to find Ben Wills
Ben is building OppAlerts – AI search visibility, historical PageRank off Common Crawl, news-specific link gap analysis and a classifier that checks your copy against 20,000+ taxonomy categories. His marketing news aggregator is free and re-ranks every fifteen minutes at news.oppalerts.com. He is on LinkedIn, and his own site is benwills.com.
Full episode and transcript are on the site. If you take one thing out of it, take the homework: spin up Manticore in Docker tonight and make relevance work.

