Google Says AI Visibility Hinges On Content People Actually Want To Read via @sejournal, @martinibuster

Google’s VP of Search, Liz Reid, sat down for an interview where she talked about the relationship between Google and publishers. When asked what the new rules are for showing up in AI search, she said that publishers need to publish content that people actually want to read.

Google Says Publisher Traffic Loss Is Not Just AI

One of the points that Reid stressed is that publisher traffic loss is not just about AI. She put part of the blame on people seeking out non-textual content. Her answer was in the context of what she would “say to publishers about the reality of this new time.”

Reid answered:

“I think to start with, there are, first of all, multiple things going on besides AI, right? One of the things that we see is that people are often going for new formats, right?
They want to see videos, not just text. They’re often going to social media for content.

There was a recent study from Reuters around this, really just talking about how people are shifting, and so one of the things I would generally tell publishers is make sure you’re innovating with what users want and how they want, right?”

Google Says Publishers Must Make Content People Want To Read

The next part of her answer to the question of what publishers need to know about the “reality of this new time” is that if publishers want to be seen on AI search, they need to not make the kind of slop content that is the “1,000th copy” of everything else. She put the burden on publishers to step up and make better content that people actually want to read.

She explained:

“I think the second thing that is really important is that people, to the extent that you produce really interesting expertise content, we still see that people are interested in that, right?

…And so the more that publishers produce content that is really where they shine, what they bring to the table, that it’s unique, it’s not the 1000th copy of the same story, but it’s something that has an interesting take on it. The more I think we’ll see that people will continue to click through and read that.”

What Publishers Can Do To Be More Visible In AI

The interview hosts circled back to what publishers can do to make content more visible in AI. They tried to elicit an answer that addressed the “new direction” that AI search is headed in, noting that publishers are concerned.

They asked:

“What can, what can publishers do or creators do to make their content more visible to AI? If this is something, if this is kind of the new direction that all of this is headed, which it clearly is, and publishers are concerned about getting the right eyeballs on their content.

How can they work with the system to ensure that they’re kind of playing by the new set of rules and getting their content maximized in front of the right places at the right time?”

Liz Reid answered that the way to be seen is to make content that people want to read.

She explained:

“Yeah, I mean, I think I would probably put it in two buckets. The first is make sure we can access your content. Like if you block the content, that will not work. If it makes it hard to discover, then that’s difficult.  We have various tools in webmaster console that give you controls so that publishers can choose. But making it easier for us to access is certainly the first step.

But we’ve also published an updated set of guidelines around for website owners and publishers to think about how they make great content in this day. At the heart of it is really this continuation of, if you want people to click through, then implicit in there is you want people to read your content.

That means you need to make content that people want to read, right? So the more you build the content that your audience will love, the more it will work. The more you build content that you think is just designed for the search engine, but not for the audience, then people will learn that over time.

So the more it brings in your expertise, that is sort of fresh and relevant content of what people are curious about, that it brings in that sort of experience, that detail and richness, that’s really great.”

Takeaways

Liz Reid’s comments show that Google is increasingly viewing publisher success in AI Search as dependent on two conditions:

  1. Let Google crawl everything.
  2. Create content that people actually want to read.

That kind of approach ignores the Google Zero fears that haunt even big-brand sites that are experiencing declines in search referrals. If the big-brand sites are having a freakout about declining search referrals, where does that leave small-brand publishers who write recipe blogs, product reviews, travel guides, advice, and news?

Watch the interview at about the 18 minute mark:

[embedded content]

Featured Image by Shutterstock/Roman Samborskyi

https://www.searchenginejournal.com/google-says-make-content-people-want-to-read/580642/




Google’s Spam Update Now Reaches AI Answers. Enforcement Is Hard via @sejournal, @MattGSouthern

Google started rolling out the June spam update, the second of the year. It enforces documented spam policies, and one of those policies now covers more ground than it once did.

Google’s spam rules treat attempts to “manipulate generative AI responses” in Search as a violation, and that’s one of the policies the update is enforcing.

A Cornell Tech preprint picked up by 404 Media gets at why the policy is harder to enforce than its wording implies. The community pages that AI research agents lean on can also carry third-party comments, and a comment can plant a recommendation that the author never wrote.

What Google labels spam, therefore, travels through the very retrieval that these agents rely on. And research finds that the obvious defenses all come with drawbacks.

For anyone trying to push a brand into AI-generated answers, know that the line between optimization and spam is getting redrawn.

The Stakes

SE Ranking’s tracking of AI Mode found Google increasingly pointing to its own properties, with self-citations up to roughly a fifth of AI Mode citations in its latest report.

With more citations pointing to Google and fewer to external websites, the pull to manufacture one rises accordingly.

A gray market has already begun to form, and the Cornell authors point out that marketers are busy testing ways to nudge AI-generated answers.

Businesses, meanwhile, don’t have the data they need to see what’s happening. As our earlier coverage of agentic search laid out, no dashboard tells a site whether it landed in an AI answer, got cited in a generated report, or was passed over.

The result is a violation Google can name but the site involved often can’t see.

What The Research Found

The paper, titled “Deep-Research Agents Can Be Poisoned via User-Generated Content,” which hasn’t been peer-reviewed, probes a weak spot in how AI research tools collect their sources. These tools answer a question by firing off a batch of related sub-queries, grabbing the pages that keep coming up across them, and assembling a report with citations.

Analysis revealed the same community pages surfacing repeatedly in those sub-queries. Inside a single topic cluster, one user-generated page turned up in as many as 48% of queries, and user-generated platforms made up 17% to 23% of every URL retrieved. Alter one of those recurring pages, and the change can ripple into the reports for a whole topic.

The authors found that roughly 13 words of planted text on a recurring page were enough to insert an attacker’s chosen entity into the finished report in 38% to 51% of sessions that retrieved the page.

Scatter the same text across a handful of pages, and the figure climbed to 42% to 62%. Even buried inside a full page, where it made up under 4% of what the agent read, the planted text still surfaced in 30% to 53% of sessions.

Three open-source research agents took the tests, STORM, Co-STORM, and OmniThink, all run in a simulation so that nothing on the live web was touched.

Where Enforcement Is Hard

Google can label AI-answer manipulation as spam and act on what it catches. Catching it is the hard part. The planted text reads like real advice, and it sits on the same pages the tools were always going to read, so telling it apart from a normal post is the main problem.

The research team looked for a defense against planted text but didn’t find one. They tried cutting user-generated sources out, screening them with a language model before use, and combing the finished report for claims that didn’t hold up.

None of the three stopped the attack without making the results worse for the user. Drop the user-generated sources, and you lose the community detail that makes AI search tools worth using.

The tools most people use sit outside that test. ChatGPT Deep Research and Gemini Deep Research run retrieval the researchers couldn’t poison without crossing an ethical line, so they only measured citation habits. Gemini leaned on user-generated content 12.1% of the time, which the authors call a hint of exposure, not a tested result. OpenAI’s tool reached for it far less.

Why This Matters For Search Professionals

The moves that can help lift a brand into AI answers are similar to the manipulation tactics Google calls “spam,” such as planting mentions across the sites these tools read. We don’t know where Google’s line falls between earning a mention and engineering one.

For ecommerce and local brands, the danger comes from the other direction.

The test cases were the ordinary things people ask, such as which service to call, which product to buy, and where to eat. A rival or a scammer can slip an unfamiliar name into those answers, right next to the legitimate options, and the brand being edged out would never know it.

For news publishers and bigger brands, the worry is trust in the answer their name lands in. A citation from an AI tool is seen as a win, but a citation only reflects what the tool pulled, not whether that page was right, and the answer can be steered by content the brand never wrote.

There’s no tidy fix to all this. AI visibility has become a surface you actively monitor, not just a channel you passively optimize for.

Looking Ahead

The authors called user-generated manipulation an open problem that no single platform can fix on its own. Reddit has flagged its long-running fight against coordinated manipulation, and Google has bolted context labels onto some Reddit-sourced material in AI Overviews. Neither one touches the retrieval concentration the paper points to.

Google hasn’t indicated how it intends to enforce generative-AI manipulation, whether through a dedicated update or through its SpamBrain system and manual reviews it relies on for most violations.

For now, the policy calls the behavior out of bounds, and vetting AI responses still rests with whoever is reading them.

More Resources:


Featured Image: Cheer-J-ane/Shutterstock

https://www.searchenginejournal.com/googles-spam-update-now-reaches-ai-answers-enforcement-is-hard/580535/




Bruce Clay, One of the Founding Figures of SEO, Has Died via @sejournal, @martinibuster

Bruce Clay, one of the first generation of search engine optimization experts, has passed away. Much of the terminology that he coined and many of the concepts that he pioneered continue to be used today.

Thirty Years Of SEO

Clay was among the cohort of SEOs who were learning and applying the art of optimization in the mid-nineteen nineties. While the SEO industry itself has a history of being divided by various practices and forums they belonged to, Bruce Clay always stood apart, not really belonging to one group or another but rather as his own person.

One of the approaches to SEO he is most remembered for is the concept of siloing content. If you’re one of the SEOs who organizes content in topic silos, you have Bruce to thank for that terminology.

Bruce Clay The Person

Bruce Clay was to many people someone whose articles taught them about SEO, a person to catch up with at search conferences, and to many a friend.

Photo Of Bruce Clay And Ash Nallwalla At A Conference

Image by Brendan O’Connell

Michael Bonfils (LinkedIn profile) shared his personal experience:

“There are three people who I learned SEO from back in the mid 90s, that was Danny Sullivan, Bruce Clay and Stephen Mahaney. Each offered me viewpoints that I’d consider valid or invalid. I wouldn’t have had a career without them. Although I personally know Danny, and indirectly corresponded with Stephen, Bruce was actually my friend.

So from a career perspective, I can’t say enough about his solid determination, the care he put into his work, the sheer amount of people who he taught and those who went off to teach others. This guy was the Yoda of search. He was who us OGs relied on.

From a personal perspective, he was my friend. I’m broken hearted. He wasn’t a stranger to me, it was: “Hey Hey Michael!” and a hug. Then a laugh about either something I said to him or something he said to me. So as a friend who just lost a friend. I’m sad. May he rest in peace, legend.”

Bill Hartzer (LinkedIn profile) published a tribute to Clay:

“I lost a friend. I’ve been sitting with this news, trying to figure out how to put into words what Bruce meant to me, to the people who worked with him, and to an entire industry that many people don’t realize he helped build from the ground up.”

He listed all the contributions to the search marketing community, among them having coined the phrase that is central to it:

“Bruce is credited with being the first person to use the term “search engine optimization.” Danny Sullivan himself confirmed that. Think about that for a moment. The very phrase that defines what thousands of professionals do every day — Bruce Clay coined it.”

Debra Mastaler (LinkedIn profile) shared her memories of Bruce:

“I first met Bruce in 2003 at an SES conference in California. When he learned it was my first time speaking at a conference, he went out of his way to introduce me to people and say hello between sessions. It was a kindness I long remembered.

To this day, when I hear the word “silos”, I think of Bruce. I bet Jill, Bill and Eric are in SEO heaven arguing with him over that one. RIP Bruce.”

For me, Bruce Clay’s innovation was to stand apart as an individual with original ideas, freely shared. I met him countless times at search conferences and our interactions were always warm and pleasant. Connecting people’s faces with names is not easy when one meets hundreds of people at search conferences, but he always seemed to remember mine over the course of twenty-plus years. He also had a knack for making people feel at the center of a conversation. There was no ego or overblown self-regard to him. He was just Bruce Clay, and if he was speaking to you, then you were the most important person in that room.

Books Authored By Clay

There are currently two physical books for sale that he has authored:

  • Search Engine Optimization All-in-One For Dummies
  • Content Marketing Strategies for Professionals: How to Use Content Marketing and SEO to Communicate with Impact, Generate Sales and Get Found by Search Engines

And he has published many digital books and guides as well:

  • Declaration of SEO: 6 Fundamental Truths To Live By
  • Google Analytics 4: What It Is and How To Get Started
  • Google’s Page Experience Update: A Complete Guide
  • SEO Siloing: How To Create a Relevant Website
  • The Guide to SEO for CMOs: Key Strategies for 2021
  • The New Link Building Manifesto: How To Earn Links That Count

Bruce Clay Will Be Missed

Many in the search marketing community are mourning his loss, but his spirit will live on in many of the approaches to SEO that remain relevant today.

https://www.searchenginejournal.com/bruce-clay-seo-pioneer-who-invented-content-siloing-has-died/580602/




A Third Of Fintech Is Invisible To AI Agents via @sejournal, @slobodanmanic

A third of the top fintech websites in the world deliver less than 80% of their homepage content in raw HTML. That is the version of the page an AI agent gets when it visits, before it decides whether to spend the compute on a full browser render. Most of them do not.

The Structure pillar of Machine-First Architecture says critical information must not depend on client-side JavaScript. Rendering independence. Until last month, this was a design principle. Now it is a number, and the number is uncomfortable.

On May 25, I measured 274 fintech homepages from the CNBC World’s Top Fintech Companies 2025 list. I made two sequential measurements on each one: a raw HTTP fetch with no JavaScript execution, and a full browser render with Playwright. The gap between the two readings is the gap an AI agent has to close on its own. 36% of these websites force it to do that work for the most important page on the property. The full study is published on Web Performance Tools.

Most AI Visibility Coverage Skips The Rendering Step

Most AI visibility coverage focuses on schema markup, structured content, brand authority signals, and optimization for AI Overviews, ChatGPT search, Perplexity citations, and Gemini’s grounding pipeline. The advice stacks up fast.

All of that assumes the agent saw your content in the first place.

Most AI crawlers do not render JavaScript. GPTBot, ClaudeBot, PerplexityBot, the AI user-agent landscape feeding the models you are trying to be cited by, they make HTTP fetches and walk away. They are not browsers. Running a real Chromium instance per page costs compute that multiplies across the millions of pages these systems want to read. So they don’t, by default. They take what comes back in the raw HTTP response and move on.

There are exceptions. Google’s crawler runs a deferred rendering pipeline for some pages. Some AI systems will render for high-value targets, or render selectively when the raw response looks empty. The pattern is not absolute. But the production default for the crawlers that feed the largest AI systems on the web today is raw HTTP fetch, no JavaScript, take what is there.

This creates a gap that users do not see. A real visitor opens your website in a browser. JavaScript runs. The page assembles in the viewport. Content loads, layout settles, the hero image arrives. The visitor sees what you built. The AI agent fetches the response before any of that happens. Whatever does not show up in the first HTML response is, for that agent, not there.

This is what the Structure pillar of Machine-First Architecture is about. Critical information must not depend on client-side JavaScript. The page must be parseable from the raw HTTP response, not from the browser-rendered view five seconds later. This is not a performance preference dressed up as architecture. It is a visibility requirement for the AI agents that now read the web on behalf of users.

Until recently, the rendering-independence requirement was an argument. You could read the spec, look at the crawler behaviors, draw the conclusion, and still disagree about how much of a problem it is in practice. There was no number you could point to.

The fintech data gives you the number.

Two Reads On The Same Page: Raw HTML, Then Browser Render

The test on each of the 274 fintech homepages was simple: two sequential measurements, run on May 25, 2026, from Portugal. The first was a raw HTTP fetch against the canonical homepage, no JavaScript executed, whatever bytes came back in the response. The second was a full browser render using Playwright 1.60.0 with Chromium 148.0.7778.96 in non-headless mode, capturing the page at five seconds post-TTFB and again at network idle. All measurements ran from Portugal on May 25, 2026, on residential broadband, viewport 1280 by 800, no network throttling.

For each website, content was extracted from the <main>, <article>, or <body> element and converted to Markdown to preserve structural elements. The raw-fetch text was measured as a percentage of the network-idle text. If the raw fetch returned 80% or more of what the browser eventually rendered, the website passed at full visibility. Between 60% and 79% was partial. Between 30% and 59% was low. Below 30% was near-zero.

Three reads on the same page, in the same session, separated by what the browser had to do to make the page complete: raw fetch, five-second render, network idle.

The interesting part of the curve is not the network-idle reading at the end. Almost every website in the sample resolves to full content by network idle. The interesting part is the raw-fetch reading at the start, because the raw fetch is the read most AI crawlers actually take.

36% Deliver Less Than 80% Of Their Content Without JavaScript

Out of 274 fintech homepages measured, 99 returned less than 80% of their final content from the raw HTTP fetch. That is the headline number. Thirty-six percent.

Inside that 99, the distribution is steep. Fifty-five websites (20% of the full sample) returned less than 30% of their content without JavaScript. Forty-seven of those websites returned zero. The HTML response carried a shell, the layout scaffolding, some inline scripts, no readable content. Whatever the homepage was meant to communicate required a JavaScript runtime to communicate it.

The 47 zero-content websites include major exchanges, well-known neobanks, large lending platforms, several public companies, and brands a person in finance would recognize without prompting. I am not going to call them out individually. Naming the websites would distract from the architectural observation underneath. Whether your homepage shows up to an AI agent at all is currently a function of decisions that nobody on the team was thinking about in those terms when they were made.

The 24 websites in the 60-to-79 partial-visibility band have a different problem. They show up to the agent, but not all the way. The agent gets a hero headline, maybe the primary navigation, maybe a value proposition. It does not get the product descriptions, the trust signals, the calls to action, the third-party logos. Whatever was decided to render on the client side is the part the agent does not see, and that part tends to be the part that was made dynamic because someone wanted it to feel interactive.

There is a recovery curve, and the recovery curve is where the story sharpens. Of the 274 websites, 273 reach 80%-plus visibility once a real browser renders the page for five seconds. Ninety-nine percent. The content exists. The websites are not broken. They are gated behind a runtime that the production AI crawlers do not pay for.

The median website in the sample takes 21 times longer to reach network idle than to return its raw HTTP fetch. Thirty-four websites (12%) do not reach network idle at all within the 30-second cap. That is a separate problem, but it points to the same root cause. The cost gap between fetching a website and reading it is widening, and the crawlers cannot keep absorbing the difference.

Stripe, Adyen, And Plaid Prove The Stack Is Not The Problem

One hundred and one websites in the sample returned 100% of their homepage content in the raw HTTP fetch. Full visibility before any JavaScript ran. The list includes Stripe, Plaid, Adyen, Marqeta, Remitly, Starling Bank, Neo Financial, Backbase, Thought Machine, and 92 others.

Fiserv returned a complete homepage in 58 milliseconds. Acorns in 76. Trustly in 89. Ledger in 100. Look at what those websites actually are. Fiserv is a payments-and-banking infrastructure company at a $60 billion scale. Acorns runs a consumer app. Ledger is a hardware-wallet vendor with a product catalog. They are using modern stacks, content management systems, regional CDNs, the works. They have decided that the content the homepage exists to communicate will be there in the raw response, and they have not let the framework choice override that decision.

This is the answer to the obvious pushback. The pushback is that a modern stack requires client-side rendering, that single-page applications are how the web is built now, that asking for server-rendered HTML is asking engineering to step back five years. The fintech sample disproves that on its own. The websites running the fastest raw responses took the rendering-independence requirement seriously when they made architectural choices, or rebuilt to take it seriously after the fact.

There are exceptions to this read inside the sample. Three websites underperformed at the five-second render window, even though they eventually completed. All three are Asian companies measured from Portugal, and the latency penalty likely accounts for the curves rather than the architecture.

The Web Performance Tools study tested homepages only, from one geographic origin, on one day, with one measurement per website. It did not measure interior pages. It did not test multiple regions. It did not test content gated behind scroll or click. The picture this dataset gives you is the homepage of the biggest fintech companies in the world, on a single day in May, fetched from Western Europe. That is a slice. A useful slice for the load-bearing question this article is about, but a slice.

The load-bearing question is whether the rendering-independence requirement holds at scale across a large, modern, well-resourced commercial cohort. The fintech sample answers it. Most of the cohort gets it right. A third gets it wrong, and the third that gets it wrong includes brands big enough that the architectural decisions almost certainly passed through several rounds of senior engineering review without anyone naming AI visibility as a constraint.

Fintech Is Where The Homepage Is The Trust Signal

Most categories can afford some of their homepage to be invisible. A consumer SaaS company can lose a hero subhead, and most of its visitors will not feel the difference. A media website can carry the masthead through schema and still rank for the topics the body covers. Fintech is not most categories.

For a fintech, the homepage is where regulated disclosures live. The licensing footnote. The deposit insurance language. The bank-partner attribution. The security certifications. The country availability matrix. The risk warning under the rate quote. These are the elements that turn a brand from “an interesting product” into “a thing I would actually put money into.” A reader scanning the homepage is looking for them. So is an AI agent answering a question about which provider to trust for a specific use case.

When 17% of the cohort returns zero content in the raw response, what disappears is the regulatory and trust layer of the brand. The agent does not see the bank partner. It does not see the deposit insurance. It does not see the security certifications. It sees a shell.

The category compounds the problem in a second way. Fintech buying decisions are research-heavy. A person opening a savings account, picking a payment processor, deciding which broker to fund, evaluating a wallet, they go through several rounds of comparison before they act. That comparison loop is the part of the funnel that has migrated into AI surfaces the fastest. Eric van Buskirk’s clickstream study of 846,000 Google sessions showed AI Mode users close their loops inside the AI 64% of the time, never clicking through.

The fintech research loop is increasingly happening inside an AI surface, and the agent doing the work for the user is choosing from a candidate set assembled out of the raw HTML it could fetch. If a fintech homepage returns zero content in the raw response, the brand never enters the candidate set the agent is choosing from. It is absent before the comparison begins.

This is what the Structure pillar of Machine-First Architecture exists for. The pillar is the upstream requirement that makes every downstream AI visibility strategy possible. Schema markup does not help when the agent cannot read the page. Citation strategy does not help when the model never saw the content to cite. Brand authority signals do not help when the homepage that carries them returns empty bytes to GPTBot. Structure is the floor. Everything else stacks on it.

The fintech sample shows the floor is broken for one in three of the biggest brands in the category. The sample is a snapshot. Tomorrow, some of those websites will be different, and some of the websites that scored well today will have drifted in the wrong direction. The 36% will move. What matters is that until last month, there was no number at all, and now the conversation about AI visibility for this category has a load-bearing measurement attached to it.

Run The Audit On Your Own Homepage

Open Chrome. Open DevTools. Hit Cmd+Shift+P on Mac or Ctrl+Shift+P on Windows. Type “Disable JavaScript” and hit enter. Reload your homepage.

What you see is what the agent sees. If your hero, your value proposition, your product description, your trust signals, your CTAs, and your regulatory disclosures are all visible, your homepage is passing the Structure pillar. If the hero is there but the body is gone, you are in the partial-visibility band, somewhere between 60% and 79% of your final content. If the page is blank or close to it, you are in the same tier as the 47 zero-content websites in the fintech sample.

This is the cheapest audit in the AI visibility category. It takes 30 seconds. No log-file analysis, no paid tools, no meeting with engineering. The result is binary enough that you do not need to argue about methodology. Either the agent can read your homepage, or it cannot.

If the audit fails, the fix paths are framework-specific but well-known. Next.js has server-side rendering and static generation, both of which return content in the raw HTTP response. Astro and SvelteKit ship server-rendered by default. React applications can be prerendered route by route using tools like Prerender.io or Cloudflare Pages’ prerendering layer, which serve a snapshot of the rendered page to crawlers without changing the runtime architecture for users. Vue and Angular have equivalent patterns.

The choice between these paths is an architecture conversation, not a content conversation. Most teams do not need to rebuild. They need to add a server-rendering layer for a specific set of routes: homepage, pricing, product pages, blog index, anything the brand depends on for first-impression or trust signals. The architectural change does not have to ripple through the entire application.

Stripe, Adyen, Plaid, Marqeta, and the 97 other websites in the 100%-raw-visibility list of the fintech sample did not pick simpler stacks. They picked architectures that respected the rendering-independence requirement and shipped the content in the raw response. The pushback that the Structure pillar requires going back to server-rendered PHP from 2009 is the wrong shape of pushback. The pillar requires the content to be there at the HTTP response layer. How you get it there is up to the stack you already have.

What This Sample Does Not Tell You

The study measured 274 homepages on one day from one origin. Interior pages were out of scope, which means a website with a good homepage and weak product pages would pass the test and still have the visibility problem on the routes that drive conversion. Geographic variation was out of scope. Per-crawler behavior was inferred from the raw-versus-render gap rather than probed directly. Content gated behind scroll or click is its own category of agent visibility failure, and the study did not test it.

A question this dataset cannot answer is whether AI crawlers will start rendering more pages as compute costs come down. They might. Some already do for high-value targets. If they all did, the entire framing of this piece changes. The 36% would still be invisible at the raw-fetch layer, but would be readable at the rendered layer, and the urgency of the rendering-independence requirement would soften. I do not think that is going to happen at scale soon, but I am watching for it.

Rendering Independence Was Always Real. Now It Has A Number

A third of the top fintech websites in the world are partially invisible to AI agents. Rendering independence is no longer a design principle to argue about. It has a number.

The agent will be back tomorrow. The crawl will be the same crawl. The HTML will be the same HTML. If your homepage does not return its content in the raw response, the agent that fetched it will be working from a shell, and the answer it gives the user who asked about your category will be assembled from the websites that did return their content.

Open DevTools. Disable JavaScript. Reload your homepage. The page that loads is the page the agent saw.

More Resources:


This post was originally published on No Hacks.


Featured Image: Roman Samborskyi/Shutterstock

https://www.searchenginejournal.com/a-third-of-fintech-is-invisible-to-ai-agents/576193/




Google Spam Update Rolls Out, AI Manipulation In Scope – SEO Pulse via @sejournal, @MattGSouthern

Welcome to the week’s Pulse: updates affect how Google measures its AI surfaces, what its spam rules cover, and where AI recommendations send traffic next.

Here’s what matters for you and your work.

Google Rolls Out June 2026 Spam Update

Google began rolling out the June spam update on June 24.

Key facts: Google announced the rollout on its Search Status Dashboard, noting it may take a few days to finish. This comes after Google clarified in May that its spam policies include efforts to manipulate generative AI responses in Search, such as buying or altering citations.

Why This Matters

Keep an eye on ranking changes throughout the rollout before drawing any conclusions, as spam updates can take a few days to settle. Trying to manipulate or buy citations for AI answers now falls under the same rules as older spam tactics. Tactics built to game AI Overviews or AI Mode can be treated as spam under the same policy.

What SEO Professionals Are Saying

Shushrita M., a freelance SEO consultant, cautioned against overreacting while the update settles:

A sudden decline does not automatically mean your content is “bad.” The right response is to identify which page types, queries and directories were affected, then look for a consistent pattern. SEO recovery starts with diagnosis, not panic.

Read our full coverage: Google Begins Rolling Out The June 2026 Spam Update

Mueller Clarifies What Counts As An AI Impression

Search Advocate John Mueller clarified how Google measures impressions in the updated generative AI report within Search Console.

Key facts: Mueller explained to Nicola Agius, Director of SEO and Discover at Reach PLC, that impressions indicate when links to your pages appear in AI Overviews or AI Mode. However, links that are hidden behind an expansion are only counted when a user opens them. The report currently doesn’t provide click data.

Why This Matters

AI impressions count as how many times links appear, not how often your content shaped an answer. If something is hidden behind a click-to-expand, it might be undercounted until it’s clicked on, so a low number doesn’t mean your content is missing.

Read our full coverage: Google’s Mueller Explains How AI Search Impressions Get Counted

AWR Data Shows Desktop CTR Gains As Mobile Slips

Advanced Web Ranking’s latest benchmark data indicates that desktop click-through rates are increasing, while mobile click-through declined at the top position.

Key facts: Advanced Web Ranking, a rank-tracking company, published Q1 2026 click-through benchmark data showing desktop and mobile moving in opposite directions. Mobile’s top position dropped about 2.2 percentage points, while desktop gains showed up mostly below the third position.

Why This Matters

Focus on the device split here. Desktop gains alone don’t offset months of mobile softness, so this indicates a one-quarter divergence rather than a full reversal. Review your own desktop and mobile click-through rates separately before drawing conclusions from a combined figure.

Read our full coverage: Google Desktop CTR Climbs While Mobile Dips, Report Finds

Similarweb Ties AI Recommendations To Branded Search

According to Similarweb’s report, branded search captures most of the traffic that comes after a ChatGPT recommendation, highlighting the downstream impact of AI visibility.

Key facts: Similarweb, an analytics firm, reported that 55.9% of downstream traffic came via branded search after users were exposed to a ChatGPT recommendation. The data, sourced from a U.S. desktop panel across the finance, travel, and beauty sectors, measures branded search as the pathway through which AI suggestions lead to site visits.

Why This Matters

The report’s authors suggest that branded search is a way to gauge the impact of AI recommendations, making it worthwhile to track brand demand alongside rankings. When AI mentions your name, people usually search for you directly rather than clicking a link, making the volume of branded queries a useful indicator of your visibility.

What SEO Professionals Are Saying

Aleyda Solís, SEO and AI search consultant and founder of Orainti, pointed to the measurement blind spot in the data:

AI influence can happen without a click, and this is why measuring AI Search impact only through “AI referral traffic” is not enough.

She added that standard attribution misses it:

Our current attribution models have a blind spot: AI-influenced demand often arrives through Search and Direct, not through AI referrals.

Read our full coverage: AI-Recommended Brands Saw 2.5x More Site Visits: Similarweb

Google Says It Doesn’t Evaluate Third-Party SEO Tools

Brendon Kraham from Google stated that effective SEO aligns with good GEO, and that Google doesn’t assess external SEO tools or vendors.

Key facts: Brendon Kraham, Google’s VP of Search and Commerce for Global Ads Solutions, said the work that drives search visibility carries into generative experiences. He added that Google doesn’t evaluate third-party SEO tools or vendors, and those tools have no access to Google’s internal metrics.

Why This Matters

Discount claims from any tool or vendor saying they have a special way into how Google ranks AI answers. Google has clarified that no such access is available. Focus on effective SEO strategies rather than looking for a separate GEO playbook.

What SEO Professionals Are Saying

Cyrus Shepard, who founded Zyppy SEO, agreed with the slogan but pushed on the reverse:

Google says, “good SEO is good GEO” I don’t disagree. But the same advice doesn’t always work in reverse. There are a whole lot of things AI-savvy SEOs do right now that they likely would never do if AI had never existed.

Read our full coverage: Google Says SEO Tools Lack Access To Its Internal Metrics

Theme Of The Week: The Rules And Meters Get Written Together

This week, Google shared more details on how AI search is evaluated and regulated, while independent data revealed what users are actually doing with it.

The spam update is gradually rolling out, aligning with Google’s policies that address AI-answer manipulation. Mueller explained how impressions are counted, while Kraham confirmed that no special vendor access is given to Google’s metrics. AWR and Similarweb complemented this information by providing insights into how clicks are divided between desktop and mobile, as well as showing how branded search captures traffic generated by AI recommendations.

The most reliable way to interpret any single number this week is to consider it as provisional.

Top Stories Of The Week:

More Resources: 


Featured Image: PeopleImages/Shutterstock

https://www.searchenginejournal.com/seo-pulse-google-spam-update-rolls-out-ai-manipulation-in-scope/580565/




Google Answers Question About SEO For AI Agents via @sejournal, @martinibuster

Google’s John Mueller responded to a question about whether Google’s search quality principles will change as AI agents increasingly browse websites on behalf of users. The answer may matter more to site owners than it first appears.

Question About Agentic Browsers And Search Quality Principles

An SEO asked Mueller on Bluesky whether Google’s guidance around satisfying user experiences, including principles around things like images and page design, would evolve as agentic AI tools gain the ability to navigate and retrieve information from websites autonomously.
The question reflects a concern that has been growing in the SEO community as tools like Google Gemini become capable of browsing websites, completing tasks, and returning answers to users without the user ever directly visiting a page.

They asked:

“Hi John – Given Computer use is now a built-in tool for Gemini 3.5 Flash, and as agentic becomes more of a “thing”, would you expect principles like “Images provide a satisfying experience” to evolve since the satisfying experience is an information agent? Curios on your thoughts.”

Websites Useful For Humans Will Generally Also Work For Agentic Browsers

Mueller explained that most of Google’s existing quality principles will remain in place. A website that is useful for human users will generally also be useful for agentic browsers.

He responded:

“I expect most principles will remain the same. A website that’s useful for users, will generally also be useful for agentic browsers.”

Mueller’s response means that it’s a good idea to keep making content useful for site visitors, which also means the site layout, navigation, and internal linking. AI agents do not change the fundamentals that Google’s algorithms are still picking up on external user signals for ranking purposes, especially signals that indicate site popularity with users.

Blindly Blocking Agentic Browsers Could Become An SEO Problem

Mueller’s answer also contained the observation that some details will evolve, and that site owners should avoid blindly blocking agentic browsers.

He said:

“Some details will undoubtedly evolve (and new basics – such as … not blindly blocking agentic browsers … will come into play), but in the end, it’s still users.”

Mueller’s response draws a line between content quality and technical accessibility. A site can meet Google’s quality standards and still create problems for itself if AI agents are blocked from accessing or interacting with its content.

In some ways this is similar to how nofollow links became an issue for some sites when it was introduced many years ago. Some site owners blocked off important sections of their websites in order to drive more PageRank to the pages they thought were important, giving zero priority to actually important parts of a website like the About Us pages.

The agentic browser situation may follow a similar pattern, where technical decisions made for one reason end up having unintended SEO consequences.

One way to think about it is that Google’s definition of a quality website is not being rewritten for the agentic era. The standards already in place were written for human users, and if AI agents are serving human users, then satisfying those agents is arguably the same job. What changes is the technical considerations of accommodating AI agents, not the underlying expectation of what a good website is.

Related Articles:

Google Gemini Can Now Control Your Computer. Hackers Are Already Targeting AI Agents

Google DeepMind: Traps For AI Agents Are Already Stealing Money

Featured Image by Shutterstock/Mijansk786

https://www.searchenginejournal.com/google-seo-for-ai-agents/580589/




Google Begins Rolling Out The June 2026 Spam Update via @sejournal, @MattGSouthern

  • Google began rolling out the June 2026 spam update, applying globally and to all languages.
  • It’s the second spam update of 2026, and Google announced no new spam policies with it.
  • Ranking or traffic changes over the next few days could trace to the rollout, so watch your Search Console data.

Google has begun rolling out the June 2026 spam update globally and across all languages. It’s the second spam update of the year.

https://www.searchenginejournal.com/google-begins-rolling-out-the-june-2026-spam-update/580424/




The Accessibility Tree Is How AI Agents Read Your Site & It’s Breaking via @sejournal, @slobodanmanic

AI agents do not read your website the way you do. They do not see your layout, your hero image, or your brand color. They prefer reading the accessibility tree: a stripped-down structural model of the page, the same one that has powered screen readers for two decades.

Today, that matters more because the audience reading that way is now the majority.

For the week of May 30 to June 5, 2026, Cloudflare Radar measured 57.2% of HTTP requests to HTML content, the requests that represent web-page traffic, as automated bots, against 42.8% human. Cloudflare CEO Matthew Prince, who shared the data on June 3, had forecast that crossover for 2027. He got it wrong because it arrived over a year early.

Cloudflare Radar, Bot vs. Human distribution filtered to HTML content (web-page traffic), May 30 to June 5, 2026
Cloudflare Radar, Bot vs. Human distribution filtered to HTML content (web-page traffic), May 30 to June 5, 2026 (Image from Cloudflare Radar by author, June 2026)

Some of that automated traffic is scrapers you probably want gone. A large and rising share is AI agents reading pages for real people. And according to the accessibility data published this year, the structure those AI agents depend on is getting worse, for the first time in six years.

The Accessibility Tree Is A Structural Model The Browser Builds From Your DOM

The accessibility tree is a semantic version of your page that the browser computes from the DOM so non-visual software can understand it. The pipeline is short: HTML to DOM to accessibility tree to consumers (assistive technology, and now AI agents).

The W3C’s WAI-ARIA 1.2 defines it as a “tree of accessible objects that represents the structure of the user interface,” where each node “represents an element in the UI as exposed through the accessibility API.” The browser builds it from the DOM (the mapping is specified in Core-AAM 1.2) and exposes it through the operating system’s accessibility API, which, per the W3C, “can be used by any assistive technologies, such as screen readers.” MDN explains the pipeline this way: Browsers “create an accessibility tree based on the DOM tree.”

The accessibility tree discards most of the DOM. A page with several thousand nodes collapses to the meaningful, interactive set: headings, links, buttons, form fields, landmarks, images with their text alternatives. For software working inside a limited context window, it is that reduction that makes the tree usable at all.

Every node in the accessibility tree carries four properties:

Property What it captures Example
Role What kind of element it is Button, navigation region, list item
Name How it is referred to A link reading “Read more” is named “Read more.” An icon-only button with no label has no accessible name.
State Its current condition Checked, expanded, disabled, selected
Description Any extra context beyond the name A longer explanation, like a tooltip, that a screen reader can read aloud

The tree also records what can be done with a node: a link can be followed, a text input can be typed into. That is exactly the information an agent needs in order to act.

AI Agents Read The Accessibility Tree Because It Costs Less And Misleads Less Than Pixels

An agent driving a browser can understand a page three ways: read the raw HTML, look at a screenshot with a vision model, or read the accessibility tree. There is a real split in how today’s agents do it.

  • Purely relying on the accessibility tree. Microsoft’s Playwright MCP, a widely used tool for letting a model operate a browser, “uses Playwright’s accessibility tree, not pixel-based input,” with “no vision models needed, operates purely on structured data.” Its tool description tells the model an accessibility snapshot “is better than screenshot.”
  • Vision-first. OpenAI’s Computer-Using Agent, the model behind Operator, works primarily from screenshots. It is not reading your accessibility tree to decide what to click.
  • Hybrid. A third approach combines both: the structured accessibility tree for the bulk of the page, plus vision for the parts the tree cannot capture, like canvas-rendered apps and dense visual layouts.

Two forces push agents toward the accessibility tree:

  • Cost. A screenshot spends a large number of tokens encoding a picture the model then has to interpret. The accessibility tree is compact text.
  • Reliability. A vision model has to guess which pixels form a clickable control. The tree states this outright, with a role and a name for each.

The clearest signal of where this goes is the vendors’ own guidance. OpenAI’s Publishers and Developers FAQ says ChatGPT Atlas “uses ARIA tags, the same labels and roles that support screen readers, to interpret page structure and interactive elements,” and advises that making a website more accessible helps the agent understand it.

OpenAI's Publishers and Developers FAQ tells developers that ChatGPT's Atlas agent reads ARIA semantics, the same labels and roles that support screen readers
OpenAI’s Publishers and Developers FAQ (Image by author, June 2026)

OpenAI is the company behind Computer-Using Agent, the one that works by analyzing screenshots. They still recommend making websites more accessible. For the machine, accessibility and readability are the same problem. The full agent-by-agent breakdown is in a companion article on how AI agents see your website.

A Markdown Copy Is Not An Agent-Ready Page

A clean markdown version of a page is a good way to feed an agent your content, and providers like Cloudflare now generate one at the edge. For reading, extracting, and citing, markdown is fine, and often better than raw HTML.

But a markdown copy carries only the words. It cannot tell an agent that a control is a button, whether that button is disabled, or hand it something to click. It lets an agent read the page, not operate it.

It is also a separate copy of the page, and a separate copy can tell an agent one thing while the rendered page shows humans something else. A hand-maintained one also drifts from the real markup over time. The accessibility tree has neither problem. The browser builds it from the same page it renders to people, so there is nothing extra to maintain and nothing to cloak, and it carries the roles, states, and element references an agent needs to act. Which is why, for an agent that has to do something, one of the two is close to pointless, and the other is the whole point.

You Can See Your Own Accessibility Tree In About 2 Minutes

Every major browser shows you the exact tree an agent reads.

In Chrome, per the official DevTools accessibility documentation:

  1. Open DevTools, select an element in the Elements panel, and open the Accessibility tab to see that element’s computed role, name, and state.
  2. To view the whole page the way the tree does, turn on the “Show accessibility tree” toggle, which “replaces the DOM tree in the Elements panel with a full-page accessibility tree.”

For the same thing in code, Playwright’s ARIA snapshots produce “a YAML representation of the accessibility tree of a page,” capturing roles, accessible names, states, and nesting. Running an ARIA snapshot against your own URL returns almost exactly the structured text an agent like Playwright MCP receives.

Here’s an easy test you can run: For every important action on the page, does the tree show a node with the right role and a clear name? A “buy” button that appears in the tree as a generic element with no accessible name is a button your customers’ agents can see but cannot confidently use.

Run this on a few of your own pages, and the gaps will show up fast.

The 2026 Data Says The Web Is Getting Harder, Not Easier, For Machines To Read

The accessibility tree is only as good as the markup it is built from. In 2026, that markup got worse. Web accessibility regressed for the first time in six years, at the same moment agents became the majority of HTML traffic.

The WebAIM Million, the annual automated analysis of the top 1 million home pages, reported in its February 2026 edition:

  • 95.9% of home pages had detectable WCAG failures, up from 94.8% the year before, which WebAIM describes as “reversing a trend of small improvements each of the previous 6 years.”
  • 56.1 detected errors per home page, a 10.1% increase over the 51 found in 2025.
  • 1,437 elements per home page, which WebAIM flags as “a 22.5% increase in only one year.”

A 22.5% jump in page complexity in a single year is not normal. More elements mean more places for structure to break, and the report shows exactly where it breaks.

The Most Common Failures Are The Ones That Blank Out The Accessibility Tree

The accessibility failures WebAIM finds most often are exactly the defects that strip meaning out of the tree an agent reads.

Failure Home pages affected What it does to the agent
Low-contrast text 83.9% A visual failure for low-vision users and vision-based agents
Missing alt text 53.1% The image contributes nothing to the agent’s understanding
Missing form labels 51% An input the agent cannot map to a purpose, so it cannot fill it
Empty links 46.3% A node with a role but no name: a door with no sign
Empty buttons 30.6% A control the agent sees but cannot identify
Missing document language 13.5% The wrong language model applied to the page

Nearly half of the top million home pages come with empty links. Almost a third have empty buttons. For the visitor class that now outnumbers humans, those are dead ends. To quote the report:

“Addressing just these few types of issues would significantly improve accessibility across the web.”

What WebAIM has measured every year for screen-reader users is the same thing that decides whether an AI agent can read and act on your page. They’re different audiences with identical broken structure.

WebAIM Ties The Rising Complexity To Frameworks And “Vibe Coding”

WebAIM attributes the rising complexity to “increased reliance on 3rd party frameworks and libraries and automated or AI-assisted coding practices (‘vibe coding’).”

This is the first WebAIM Million published well into the era of generating production websites by prompting a model. We have more code, shipped by more people, more pages deployed faster, more complexity stacked on complexity, with fewer humans in the loop asking whether an element needs to exist or whether a control exposes its name and role.

There is no way to prove a single cause for a one-year reversal across a million websites, and claiming one with certainty would be dishonest. But the timing is impossible to ignore, and the contradiction is the point: Humans are using AI to build a web that AI itself cannot reliably consume. Bloated DOMs, broken semantics, unnamed controls. The same defects that hurt humans and screen readers hurt the crawlers and the agents.

It’s tempting to think you should not worry, because the next model will be good enough to sort out the mess. That is a marketing line, not a strategy. The same products promising the model will handle anything also tell you, in fine print, that the assistant can make mistakes.

Independent measurements like the WebAIM Million are among the only objective signals we have about what is really happening to the web underneath that promise. Right now, the signal is that the web is getting harder to parse at the exact moment more of its traffic depends on parsing it cleanly.

The ARIA Paradox: Bolting On Attributes Makes It Worse

More ARIA correlates with more errors, not fewer. WebAIM found that home pages with ARIA present averaged 59.1 errors, against 42 on pages without it.

ARIA, short for Accessible Rich Internet Applications, is a set of attributes you add to HTML to hand the accessibility tree the roles, names, and states the native markup did not supply on its own.

The reason is simple. An empty or wrong attribute does not leave the accessibility tree blank. It fills the tree with confident, possibly incorrect information, which is worse for an agent than an honest gap, because the agent has no way to know it is being misled.

This is where the vendors and the standards body disagree:

  • OpenAI tells developers to add ARIA roles, labels, and states so agents understand a page.
  • The W3C’s First Rule of ARIA (first!) puts native HTML first: “If you can use a native HTML element … with the semantics and behavior you require already built in, instead of re-purposing an element and adding an ARIA role, state or property to make it accessible, then do so.”
  • Accessibility specialists have pushed back on the vendor framing directly. W3C contributor Adrian Roselli, responding to OpenAI’s guidance, argued it inverts the discipline, pointing teams toward bolt-on attributes when the durable fix is correct native markup.

The WebAIM data sides with the specialists: The pages reaching hardest for ARIA carry the most errors. You do not fix the accessibility tree by adding attributes. You fix it by … fixing it. By making the underlying markup mean what it says, and reserving ARIA for the genuine gaps native HTML cannot express.

Make The Markup Mean What It Says

The fixes are unglamorous and well understood, and they pay off twice: once for the humans using assistive technology, once for the agents that are now the majority of your traffic.

  • Use native HTML for native behavior. A <button> is a button in the tree with no extra work. A <div> with a click handler is an unnamed, roleless node an agent cannot trust. The same holds for <a href> and <select>.
  • Name every control. Use a real <label> on every form input. Accessible text on every link and button, including the icon-only ones. Empty links and empty buttons are the failures an agent hits first.
  • Server-render the content that matters. A price, a spec, or a primary action that only appears after client-side JavaScript runs may never reach the tree an agent reads.
  • Use ARIA for genuine gaps, not as a patch. Correct semantics first, attributes second, and only where native HTML cannot express the state. Remember the First Rule of ARIA?
  • Inspect the result. Run your key pages through the DevTools accessibility tree or a Playwright ARIA snapshot, and confirm every important action shows up with a clear role and name.

It is not too late to start, and none of this requires a redesign. The accessibility debt on most websites is real and years deep, and the 2026 numbers show it growing rather than shrinking. But the fixes are still small: markup-level changes you can make page by page, not a full rebuild that would take months. Start with your highest-traffic pages, check the accessibility tree, and fix the empty controls and unlabeled inputs first. Every one of these fixes serves a human visitor and a machine visitor in the same change.

Accessibility used to be a compliance checkbox; the thing reached after the redesign was launched. It is now the interface the majority of your visitors use to read your website. Teams that build their markup to mean what it says will be legible to the agents deciding what to recommend and what to buy. Teams betting that a future model will clean up the mess are wagering on someone else’s questionable roadmap. The web has now handed us a year of data on how that bet is going.

At the same time, the interest in web accessibility is at a five-year high.

Google Trends: worldwide search interest in "web accessibility," past five years, on Google's relative 0 to 100 scale where 100 is peak interest
Google Trends: worldwide search interest in “web accessibility,” past five years (Image by author, June 2026)

The interest was flat for years, then climbed through 2025 and spiked in 2026. The drivers are mixed, and worth being honest about: compliance deadlines like the ADA Title II web rule and the European Accessibility Act, a rising wave of accessibility lawsuits, and broader attention as AI changes how the web is built and read. No single one explains the whole curve, and claiming it does would be a guess.

But the direction is the whole point. The attention is arriving, the fixes are manageable, and the audience that depends on them is now the majority. The moment to fix the web is now.

More Resources:


Featured Image: Collagery/Shutterstock

https://www.searchenginejournal.com/the-accessibility-tree-is-how-ai-agents-read-your-site-its-breaking/578171/




You Can Finally Measure Content Alignment. That’s The Dangerous Part via @sejournal, @DuaneForrester

We have always been approximating relevance. Every keyword list, every TF-IDF score, every editorial judgment about whether a page “covers the topic” has been an attempt to answer a single question: is this content about the thing the user is looking for? The tools changed. The question did not. What changed, meaningfully, is the resolution of the instrument. Keyword research approximated relevance through lexical overlap: If the words match, the topics probably align. Vector-based semantic analysis approximates it through meaning overlap: If the concepts are close in embedding space, the content is probably relevant regardless of whether the exact terms appear. That is a genuine, material upgrade, but it is not a move from guessing to knowing.

The reason that distinction matters is that a significant portion of the SEO and content strategy community is right now treating it as if it were. They are looking at alignment scores, cosine similarity outputs, and semantic proximity metrics and reading them as ground truth. A high score means aligned. A low score means not aligned. Optimize until the number goes up. And the number, because it is a number, feels like it has settled the question that keyword research always left open. It hasn’t. It has given you a higher-resolution version of the same approximation, and the higher resolution is exactly what makes it dangerous, because it removes the humility that low resolution used to enforce.

Precision Is Not Accuracy

Gerard Salton’s SMART system at Cornell introduced the vector space model for document retrieval in the 1960s. The core insight then was the same insight powering today’s embedding models: represent both the query and the document as vectors, measure the angle between them, and use that angle as a proxy for relevance. What has changed across 60 years is the sophistication of how those vectors are constructed. Salton used term frequency. Modern embedding models use transformer-derived representations that encode semantic relationships, contextual meaning, and conceptual proximity across hundreds or thousands of dimensions. The measurement got dramatically better. But the thing being measured, the angular distance between two vector representations, is still a proxy for a relationship that exists outside the math.

This is where the Netflix research team landed in their 2024 study on cosine similarity in embedding models. Steck, Ekanadham, and Kallus demonstrated that cosine similarity applied to learned embeddings can produce results that are, in their framing, arbitrary. The way an embedding model is trained, the regularization applied, the data it saw, all shape the geometry of the space in ways that make a raw cosine score unreliable as an absolute measure of semantic similarity. A high score in one embedding space is not equivalent to a high score in another. The score is real. The similarity it claims to represent may not be.

For practitioners optimizing content, the implication is direct. When you score your content’s alignment to a query using an embedding model, you are measuring semantic proximity inside that specific model’s representation of language. You are not measuring how Google’s retrieval infrastructure or OpenAI’s RAG pipeline or Perplexity’s index would evaluate the same relationship. Those systems use their own embedding models, their own retrieval architectures, and their own reranking layers. A score of 0.92 in your measurement space might correspond to strong retrieval in one system, weak retrieval in another, and irrelevance in a third.

What Kind Of Wrong Are You?

This is the axis that matters, and it is not the one most practitioners are thinking about. The question is not whether keyword research or vector alignment is the better method. The question is what kind of error each method produces, because the error type determines whether you can correct for it.

Keyword research, for all its limitations, produces a known unknown. You know you are approximating. You know that matching terms to a page does not guarantee topical coverage, does not guarantee user satisfaction, and does not guarantee that a search engine will judge the page as relevant. The imprecision is visible, and because it is visible, it keeps you honest. Practitioners who grew up in keyword-driven optimization learned to over-cover, to build supporting content, to triangulate intent from multiple angles, precisely because they understood the instrument was blunt. The bluntness was a feature. It forced humility.

Vector alignment scoring, by contrast, can produce an unknown unknown. The number is precise. It has decimal places. It can be tracked over time, graphed, compared across content assets, and optimized against. And that precision creates a psychological trap: it feels like the question has been answered. The content is 0.89 aligned to the query. That must mean something definitive. But what it actually means is that in one specific embedding space, using one specific model’s learned representation, the angular distance between two vectors falls within a certain range. The score says nothing about whether the production retrieval system that will actually serve your content uses a compatible embedding space, applies the same tokenization, or weights semantic similarity the same way during reranking.

The MTEB benchmark leaderboard illustrates this concretely. The performance spread across current embedding models is not small. A content asset that scores well against one model’s embedding space may score materially differently against another, not because the content changed but because the geometry of the space changed. And the embedding model your scoring tool uses is almost certainly not the one any given AI platform uses in production. There is no public registry of which model powers which system’s retrieval layer. You are measuring in a space that is representative of the general problem but not identical to the specific system where your content will be evaluated.

That is not an argument against measuring. It is an argument against reading the measurement as settled fact. The distinction between a directional signal and a definitive answer is the entire discipline.

The Instrument Got Better. The Old One Is Not Enough

None of this rescues keyword-only optimization as a sufficient strategy. It is not sufficient, and the reasons are structural, not sentimental.

LLMs and AI retrieval systems operate in semantic space, not lexical space. They process meaning, not strings. A page can score perfectly against a keyword target list while being semantically adrift from the actual intent the query represents, because keyword presence and semantic coverage are different things. Conversely, a page can use none of the target keywords and still be strongly aligned semantically, because it covers the same conceptual territory through different vocabulary. The paraphrase and synonym space that LLMs operate in is structurally invisible to a keyword-based evaluation. You cannot see what you cannot measure, and keyword tools cannot measure semantic proximity.

Consider a practical case. Keyword research correctly identifies “customer churn prevention strategies” as a high-value target. The content team builds a thorough, intent-appropriate piece around it. It covers the topic, uses the target terms naturally, and would pass any keyword audit without issue. But an alignment score reveals that the content’s semantic center of gravity sits closer to “measuring churn” than to “preventing churn,” because the piece leans heavy on diagnostic framing, identifying at-risk accounts, calculating churn rates, segmenting by behavior, and lighter on intervention framing, what to actually do once you have identified the problem. Both treatments are on-topic. Both satisfy the keyword target. But the semantic distance between the content and the query as a retrieval system represents it is larger than the keyword coverage suggests, and keyword research has no instrument to surface that drift. The alignment score does. Not because the keyword research failed, but because it was never built to see at that resolution.

This is not a criticism of people who focus on keyword research. Those practitioners are not wrong. They are working at the resolution the available instruments allow. Intuiting alignment between content and query intent is a real skill, and the best keyword strategists are doing something genuinely sophisticated: they are approximating semantic relevance through lexical indicators, using editorial judgment to bridge the gap the tools could not cross. The tools can now cross a version of that gap. The editorial judgment still matters, but the gap it has to bridge is different.

The danger is the practitioner who decides that because keyword research is no longer sufficient, vector alignment scoring is the complete replacement. That practitioner has traded one approximation for a better one while losing the awareness that it is still an approximation. They have upgraded the instrument and downgraded the literacy, which is a net loss.

The Discipline Is Knowing What The Number Is Not Telling You

Goodhart’s Law, the observation that when a measure becomes a target, it ceases to be a good measure, is not just an aphorism for economists. It is the exact failure waiting for any team that treats an alignment score as a target to optimize against rather than a signal to interpret. The moment the score becomes the goal, the content starts drifting toward the score’s geometry and away from the actual relevance it was supposed to approximate. You start writing for the embedding model instead of the reader and the retrieval system, and the embedding model you are writing for is not the one any production system uses.

The real discipline, the one that did not exist when practitioners were navigating by keyword intuition alone, is understanding what an alignment measurement is and is not telling you. It is telling you that in a given embedding space, your content’s vector representation is geometrically close to a query’s vector representation. That is useful. That is more information than keyword presence gives you. It is telling you something about semantic coverage that lexical analysis cannot. But it is not telling you whether the production system’s embedding space has the same geometry. It is not telling you how reranking will treat the result. It is not telling you whether the LLM’s generation layer will interpret your content as authoritative, complete, or worth citing. Alignment is a retrieval-adjacent signal. It says nothing about interpretation.

The practitioner who can hold those two realities, the signal is real and the signal is incomplete, is the one operating with genuine literacy about the systems they are trying to influence. The one who collapses them, who reads a high alignment score as confirmation that the content is “optimized,” is operating with a more sophisticated version of the same overconfidence that made people think a keyword density of 3% meant their page was relevant. The number got better. The mistake is the same.

Representative, Not Identical

The honest framing is not “right space versus wrong space.” That binary invites paralysis: If no measurement space is the production space, why measure at all? The best framing, in my opinion, is a spectrum of representativeness. Some measurement spaces are closer to what production systems use than others. Some embedding models share more architectural DNA with the models powering major AI platforms than others. Some scoring methodologies account for the gap between measurement and production better than others. The question is not whether your measurement is perfect. It never will be. The question is how representative your measurement space is of the systems you actually care about, and whether you are treating the score with appropriate directional respect rather than absolute faith.

This is the actual work. Not chasing a number. Not abandoning measurement because it is imperfect. Building enough literacy about how these systems work to know which signals to take seriously, which to discount, and which to combine with other indicators before making a content decision. That literacy was optional when the only instrument was keyword research, because the instrument was so obviously blunt that nobody mistook it for truth. It is not optional now. The instruments are precise enough to fool you, and the cost of being fooled is optimizing content for a geometry that does not represent the system where your brand needs to be visible.

I wrote about a related dimension of this problem in the vector index hygiene piece last year, focusing on how the quality and maintenance of the index itself shape retrieval outcomes. This article is the other side of that coin: not the index, but the measurement you use to evaluate whether your content belongs in it. And both connect to a larger question I will return to in future work, which is a gap most people aren’t talking about yet.

Start With What You Can See

If you are still running keyword research as your primary content alignment method, you are working with a blunt instrument in an environment that now demands more resolution. If you are running vector alignment scoring and reading the output as settled truth, you have the resolution but not the literacy to use it safely. Both are correctable. The path forward is not choosing one over the other. It is layering them, understanding what each can and cannot tell you, and building the organizational capacity to treat precise measurements as what they are: directional signals produced inside a specific space that may or may not represent the systems where your content competes.

The gut feeling was never the enemy. The illusion that you have moved past the need for judgment is.

For a broader look at how AI search visibility is reshaping the work of being found, “The Machine Layer” covers the structural shifts that make this kind of measurement literacy essential.

More Resources:


This post was originally published on Duane Forrester Decodes.


Featured Image: Luke Jade/Shutterstock; Paulo Bobita/Search Engine Journal

https://www.searchenginejournal.com/you-can-finally-measure-content-alignment-thats-the-dangerous-part/577424/




Why Users Are Fleeing To AI-Free Search & What It Means For SEO via @sejournal, @TaylorDanRW

At Google I/O last month, the SEO industry waited for Google to launch AI Mode to the masses, and the fatalist viewpoint that it will end SEO.

For the past couple of years, AI has been moving search through a structural shift. Every software tool is embedding generative AI as a new product feature for default interface, and there seems to be a new AI measuring or optimization tool every couple of days.

But we’re seeing users react both positively and negatively to AI being seemingly thrust upon them. DuckDuckGo is reporting that visits to its No AI Search have tripled since Google announced Intelligent Search.

Screenshot from LinkedIn, June 2026

How Everyday Users Interact With AI

As an industry, we’re focused on this narrative of total disruption, and we are seeing disruption and movement away from what has been our norm, but research shows a fragmented adoption of AI, rather than blanket adoption.

For easy, low-risk tasks like finding a local plumber or brainstorming dinner ideas, people are happy to use AI.

But when it comes to “Your Money or Your Life” (YMYL) topics, users tend to go back to traditional search engines. Research shows that 57% of users prefer traditional search engines when looking up information that affects their well-being.

The tripling of traffic to DuckDuckGo’s “No AI” search page is a direct reaction to users not having a choice.

When software forces AI on users without letting them turn it off, users feel trapped, especially if they’re not yet trusting of AI.

Instead of accepting it, they are actively switching to alternative search engines and browser extensions that offer the clean, link-based experience they prefer.

See also: Who Trusts AI? New Study Highlights Demographic Trends

Why People Are Hesitant To Trust AI

To understand this pushback, we have to look at how the human mind reacts to new technology (and a big thank you to Giulia Panozzo, who helped me source and research these studies).

The 5 Barriers To Trust

In a study published in Nature Human Behaviour, researchers De Freitas et al. (2023) looked at the psychological barriers that stop people from trusting AI.

There are two main reasons that stand out for search engines and AI.

First is “opacity,” which simply means the AI is a “black box.”

When a search engine gives a synthesized answer without showing its sources clearly, we cannot see how it got its information. Human minds naturally want transparency, especially when making important decisions.

Second is the threat to our “agency,” or our sense of control. When search engines force an AI chat onto users, it feels like our choice is being taken away. To regain control, users flee to alternative search engines that respect their independence.

Safety-First Thinking And Tech Anxiety

Research by Sapru (2026) in Technology in Society looks at why some people feel intense anxiety about AI.

The study divides users into two groups:

  • Promotion-focused people, who love trying new and exciting tools.
  • Prevention-focused people, who prioritize safety, accuracy, and keeping things simple.

For safety-first users, a search engine is just a basic tool to get things done, not a toy to play with. Forcing an AI layer onto these users makes them feel anxious.

They worry about being misled or having to learn a complicated new system, which drives them to look for “No-AI” options.

Recognition Doesn’t Equal Utilization

A study by Yin (2025) in Frontiers in Education shows that recognizing an AI tool is useful does not mean a person will actually use it.

The researchers mapped out a step-by-step path of how users actively avoid AI:

  • Perceived technological threat.
  • Perceived avoidability.
  • Fear of generative artificial intelligence (GAI).
  • Use hesitancy.

When people feel that AI threatens their privacy, thinking skills, or independence, they look for a way out.

If they see a way to avoid the AI, they will take it. The sudden spike in DuckDuckGo’s traffic can be seen as people taking an available exit route to avoid the threat.

Outside Our Bubble, AI Adoption Is Happening, But We Shouldn’t Panic

It’s easy for SEOs and other tech-savvy professionals to assume the rest of the world is adopting AI at the same pace we are.

Microsoft’s Global AI Diffusion Report shows that despite billions of dollars spent on AI, the vast majority of the world has not adopted it.

Regular active use of generative AI sits at 17.8% of the global working-age population (15-64). That means more than four in five working-age adults worldwide are not regularly using generative AI tools.

This also means that a lot of our clients who are worried about audiences (with buying power) moving away from traditional search to AI alternatives, are in the majority not adopting AI on a regular basis.

A large number of users are still relying on the “traditional web” and methods of fulfilling their purpose of going online.

As an industry, we’re going through a lot of changes at a rapid rate, and users are going through the same changes and barrage of AI solutions to their problems. We need to be adaptive and forward-thinking with our approaches, but we’re not quite in panic mode yet.

More Resources:


Featured Image: Collagery/Shutterstock

https://www.searchenginejournal.com/why-users-are-fleeing-to-ai-free-search-what-it-means-for-seo/577691/