Unscripted SEO Podcast, hosted by Jeremy Rivera. Listen to the full episode and read the complete transcript.
This episode is one of the more technical conversations we have had on the show, and one of the more directly useful. I sat down with Malte Landwehr, who runs product and research at Peec AI, a platform that measures and improves brand visibility inside LLM-based answer engines. We covered whether PageRank still runs under Google, why bot crawling breaks the reasonable surfer model, how consensus across the web overrides your own pricing page, a step-by-step method for working query fan-outs, and four measurable ways to catch AI slop before it goes live. Full audio and transcript are linked at the bottom.
The guest: 20 years of SEO and a thesis on PageRank
Malte has been doing SEO for more than 20 years, starting as a teenager when he noticed people randomly landing on his website. He studied computer science, started a PhD he never finished, and worked in social media analytics, web scraping and social network analysis. His bachelor thesis applied PageRank to scientists, ranking successful ones against unsuccessful ones. He co-founded an SEO agency, led the product team at Searchmetrics, then spent five years doing in-house SEO for the largest price comparison website in Europe. On why anyone should trust him, his own answer was that he has always focused on things that actually work and does not like to be bullshitted by fluffy statistics.
Is the old link graph still running?
Jeremy opened by asking whether the algorithm Malte wrote his thesis on is still in play. Nobody outside Google can confirm the exact random surfer algorithm from the original paper is in use, he said, but some variation of it definitely is. He is fond of it because it generalises so well, and gave the example of mapping a graph of animals and which animal eats which, then using PageRank to calculate which species would go extinct first.
His view is that the core principle survives: a graph of the whole internet, links as the edges, a calculation of which nodes are most prominent. What has probably changed is which links get counted. Some websites may not count at all anymore, and white-on-white footer links likely carry far less than they used to.
Why bot traffic is not the reasonable surfer
Jeremy raised a finding from his own analytics: comparing Microsoft Clarity against GA showed that much of what looked like direct traffic to SEO Arcade was bots crawling and scraping pages. Does the reasonable surfer idea apply to robots?
Largely no, Malte said. A crawler that checks a competitor’s top ten product prices every five minutes refreshes the same URLs endlessly and follows no links at all. Many crawlers now log into websites, which the original calculation never covered. There is probably some correlation, in that high-PageRank pages tend to get crawled more, but he described bot crawling as having very different characteristics from human crawling.
Markdown for crawlers, HTML for people
With publishers experimenting with serving Markdown, Jeremy asked whether technical SEO now needs a parallel .md file for every page. Malte said no. That puts the same content on two URLs, wastes crawl resources, and strands any human who lands on the Markdown version with no links to click.
"If humans land on the .md version, there are no links to click. There’s nothing for them to do. It’s a horrible experience."
What can work is a server-side decision on the same URL: HTML for a human, Markdown for an LLM crawler. He was candid that this is a form of cloaking based on user agent, said he would be very careful and would not do it for Google, but sees the case for a ChatGPT crawler, especially on a JavaScript-heavy site with client-side rendering.
He drew a firm line at the Time magazine version of the story, where ads are injected specifically for the LLM. If they are doing what he thinks they are doing, he called it a huge risk, because the human never sees that ad. He expects a product manager at OpenAI or Anthropic to eventually object, or to insist the ads be wrapped in markers the model can ignore. For now, he noted, it is as far as he knows the only way anyone is monetizing bot traffic.
Fast movers, broken things
Jeremy asked whether the LLM companies are being reckless with protocols. Malte pushed back on the word and reframed it as the Silicon Valley mantra of move fast and break things. He pointed to AI Overviews recommending glue on pizza, and the funnier one about horses having five legs. Because these are not predictable algorithms, unusual outputs are inevitable. His analogy was that if you train 1,000 humans to do a job, one will do it wrong and write the rude support email.
He gave two live examples of how that gets exploited. First, ChatGPT recently moved all users, including non-paying ones, to a model doing more fan-out queries, including site: searches. Some of those contain hallucinated domains. Across prompts he tracks, 2 to 3 percent contain a hallucinated domain, currently just parked. Someone could register one and publish false pricing or negative content about the brand it appears to belong to.
Second, advertorials. They are marked, so humans skip them, but LLMs use them as grounding sources and sometimes cite them directly. In a set of insurance prompts he monitors, roughly 2 percent of the sources the models used were advertorials. As he put it, you can buy that influence.
Why the internet’s opinion beats your pricing page
Jeremy asked for action items, referencing an earlier guest’s advice to treat these platforms like uneducated customer support representatives who need training material. Malte liked the metaphor and said he would steal it, but narrowed where it applies. Publishing as much as humanly possible eventually means publishing low quality and duplicates. The real mechanism is consensus.
"So if you only talk about your pricing on your pricing website, and then you change your pricing, and then there are five Reddit threads and two reviews on random blogs that still talk about your old pricing, ChatGPT will answer with your old pricing if a user asks about it."
His recommendations follow from that. Check whether you have a G2 profile, or a Yelp profile if that suits your business, and update them when facts change. Put your key message in your site footer, on your social profiles, and in the footer of your press releases. When you ship a feature, cover it in the blog post, the product page, the help centre and the product docs, each from a different angle, so the models find agreement fast.
A five-step read on any query fan-out
For non-branded prompts, Malte pointed out the LLM does not start with vendor content. It runs fan-outs like "10 best podcast recording software 2026" or the same query narrowed to a solo freelancer, a traveller, or an enterprise. His method for working that:
- Run your target prompts multiple times, not once.
- Look at which brands are currently winning, for inspiration.
- Study the cited sources at the URL level and ask to be added to pages that already mention several competitors. Skip anything that is an interview with a competitor’s CEO.
- Study the same sources at the domain level and ask what new content you could create there, through digital PR, press releases, their commercial content team, a paid article, or affiliate work. On social platforms, join the community or do an AMA.
- Read the fan-out queries themselves for recurring concepts, especially words the model added that were not in your prompt.
On that last step, if half the fan-outs contain the word review, reviews become your topic. He has watched ChatGPT append Reddit for a stretch, and more recently the word official, which is why he currently suggests putting official in your site footer. Yearly numbers are very common, so adding "in 2026" or a bracketed 2026 to titles can work. And because these systems favour freshness, refreshing content with a machine-readable last-updated date that genuinely changes can increase the chance a page gets retrieved.
Dates, cannibalisation and other loosened rules
On publish date versus updated date, Malte described two people on his shoulder. One says be transparent and show both. The other says Google will still surface the old publication date, so delete it. In practice he would display both but push the original into JavaScript, or break the date with a line break and remove it with CSS, so the updated date is what crawlers read. He summarised it as the maximum short-term SEO impact being different from the long-term trust and brand impact.
On cannibalisation, he used to be firmly against it, having worked sites with a million-plus URLs where ten thousand pages chase one topic. Now he thinks multiple pages on a topic help LLMs find consensus, so long as the intent differs: best health insurance providers, then award winners, then by state, then by income band. Overlapping, but distinct at the title level. He is still not a fan of three pages with the same title and nearly identical content.
Spotting AI slop with four measurements
Asked about guardrails for publishing at scale, Malte gave three answers. If you only care about short-term AI visibility and are willing to lose Google rankings, publish a thousand AI articles a day. He said it works, especially on an older established domain, and that he hates that it works. If you are very cautious, publish none. Most people are in the middle, and he would still not use AI for editorial text.
"Basically, if you can create it with a prompt, why would ChatGPT or OpenAI or Google crawl it, index it, and rank it? They could just use that prompt on their own."
Where he does think AI is defensible is when you supply something unique. His examples were generating articles from structured data on college basketball games nobody has written up, and summarising real user reviews on product pages. And if you are publishing at volume, he named four measures to run against a human-written baseline:
- Perplexity, how predictable each next word is. If your AI text scores a lot higher than the baseline, adapt the prompt and humanise it.
- Compression rate, borrowed from email spam detection, roughly how many words you can remove without losing information. Generic prompts with no data supplied score badly.
- Jaccard similarity, which catches the same sentence repeated across pages with a couple of words swapped.
- Cosine similarity, which catches text where every word is unique but the information is identical.
He noted you can tell Claude to write the Python script that checks all four, that it is not rocket science, and that you do not need to understand what is happening underneath to use it in-house.
He also described a briefing chain rather than a single prompt: write a briefing, use a second prompt to check the briefing makes sense, write a briefing per paragraph, generate one paragraph per prompt, add a fact-checking prompt, then check for empty paragraphs and repetitive concepts. That runs about five to seven and a half euros in token costs, and produces content he called really good. His warning was blunt: publish a lot of AI content without quality measures and you often lose your Google rankings, and then regularly your ChatGPT citations too.
Jeremy made the case that the human belongs at the brief, not the edit, and that an editor moving four sentences around is rubber stamping. Malte agreed, and said expert quotes from inside the business plus a full transcript of that person talking about their topic is the best input for a brief and then an article.
MCP and what happens to software you log into
Malte has added an MCP to Peec AI and already has customers who use the product primarily that way. He uses other tools the same way, saying he is more likely to ask Claude for a Linear ticket status than to open Linear, and more likely to have Claude search Notion than open Notion.
"I think all systems that have this character of being a system of record, like a CRM, task management, knowledge management, I think these are becoming basically databases for an MCP. Because why would I log in? There’s nothing there for me that I need."
His caveat is the uneducated user, by which he means anyone logging into 30 tools a day who is an expert in none of them. Given only a chat box, they often do not know what to ask because they cannot see what is available, which is why navigation, an interface and good charts still matter. He also described experimenting with an agent chat inside the product itself, where a social media manager could simply ask for a monthly reporting dashboard and have it built and explained on the spot. A year ago he would have said it was 100 percent user interface. Today he puts it at roughly 90 percent UI, 9 percent MCP and 1 percent API, while noting a handful of API power users skew that last number.
Where to find Malte
Malte said LinkedIn is the best platform to follow him, and you can also reach him on X or by email at his first name at peec.ai. Peec AI measures and improves visibility in LLM-based search and answer engines. It started with prompt tracking and now also covers brand perception, log file analysis and ingesting your web analytics data, along with recommended next steps to increase visibility. His writing on ChatGPT query fan-out patterns, the real risk of AI-generated content and using MCP for SEO is worth reading alongside this episode.
Listen to the full episode
The complete interview and full transcript are on the Unscripted SEO Podcast, where you will find dozens more conversations with SEO practitioners. Subscribe wherever you get your podcasts so you do not miss the next one.
Connect with Malte Landwehr on LinkedIn | Peec AI: peec.ai | Host: Jeremy Rivera, Unscripted SEO Podcast

