AI Search & Web Analytics Insights

Discover how AI search, RAG retrieval, and headless browsers transform web analytics. Learn to decode direct traffic, evaluate AI retrieval eligibility, and monitor essential geo signals for better insights.

Lydie GOYENETCHE

9/5/202618 min read

## AI Script Mystery: Unmasking the Truth Behind Your Traffic and Retrieval
## AI Script Mystery: Unmasking the Truth Behind Your Traffic and Retrieval

How to Know If Your Website Is Becoming Eligible for AI Retrieval: The First Technical Signals

Since I was forced to switch Cloudflare to DNS-only mode, following changes in Hostinger’s Website Builder and server architecture, something unusual has happened to my analytics.

The pages I had been strengthening through entity SEO, GEO and internal linking began attracting substantially more traffic — but much of it now appears as Direct / None. At the same time, long and previously little-visited URLs are increasingly becoming landing pages for that traffic.

There is some context behind this. My website contains more than 17,000 internal links, and much of my work focuses on entity relationships, semantic architecture, internal linking and GEO — although, to be perfectly honest, I increasingly struggle to see where modern SEO ends and GEO begins. Once search engines and AI systems retrieve documents, entities and passages from the same public web, the boundary becomes surprisingly blurry.

But this evolution raises a much harder question.

Who — or what — is actually visiting these pages?

Are these real people arriving without a referrer because of privacy settings, copied links, applications, stripped referral data or cookie restrictions? Are some visits generated by AI systems, automated browsers or bots that Hostinger's infrastructure does not classify or filter in the same way Cloudflare previously did? And if Plausible or GA4 records a visit because its JavaScript executes, can I safely assume that there was a human behind it?

Not necessarily.

That uncertainty becomes even more interesting when looking at the rest of the data. My website has long been visited from AI interfaces. In the period shown here, for example, Plausible identifies 262 visitors from ChatGPT, 65 from Perplexity and 27 from Claude. Yet since changing my Cloudflare configuration, explicit AI referrals have become much less visible while Direct / None traffic has grown dramatically.

Something in the observable pattern has changed.

And that leads to the question behind this article: am I simply seeing an analytics artefact caused by infrastructure, referral stripping and bot behaviour, or am I beginning to observe the first technical signs that a larger part of my website is becoming eligible for AI retrieval?

Because retrieval eligibility would matter for a very different reason than bot traffic itself. The goal is not to attract more crawlers. The goal is for the information contained in the site to become sufficiently accessible, understandable and retrievable that it can eventually be selected in response to real user queries — and bring real people back to the website.

That is where my idea of cognitive nurturing comes in: not chasing every click, but creating a sufficiently rich and interconnected information environment for a person — or initially an AI system assisting that person — to progressively discover my expertise.

The difficulty is that there is no “AI retrieval eligibility” metric in Google Analytics, Plausible or Search Console.

So we have to look for the first technical signals.

From Crawlability to Retrieval: When Direct Traffic Starts Behaving Strangely

If AI retrieval were visible through a neat analytics channel called “AI Retrieval”, this article would be considerably shorter.

Unfortunately, it is not.

One of the most puzzling changes I observed after modifying my Cloudflare configuration was not simply an increase in Direct / None traffic. It was the behaviour and geographical distribution of that traffic.

Over a single month, visitors classified as Direct / None appeared from an unusually wide range of countries. France remains important, as expected, and part of this traffic is unquestionably mine: I regularly visit the website while maintaining it, updating articles and checking technical changes.

But my own activity represents only part of the pattern.

The remaining traffic is scattered across the United States, Japan, the Netherlands, Spain, Italy, Canada, Colombia, Poland, the United Kingdom, China, Sweden, Germany, Pakistan and many other countries, with isolated visits appearing from places as diverse as Taiwan, Ukraine, Vietnam, Jordan, the Philippines or Egypt.

Some geographical areas associated with major hosting and network infrastructures also appear repeatedly in my analytics. But geography alone proves nothing. An IP address located in the United States, China or the Netherlands does not tell me whether the request originated from a reader, a VPN, a proxy, a cloud infrastructure, an automated browser or an AI-related service.

And this is precisely where the analysis becomes difficult.

Direct Traffic Does Not Mean Human Traffic

Analytics platforms use Direct / None essentially as a residual category: they received a visit but could not identify a usable referrer or campaign source.

A real person can therefore appear as Direct traffic after:

  • typing or pasting a URL;

  • opening a bookmark;

  • moving between applications;

  • using privacy protections;

  • following a link whose referral information has been stripped;

  • navigating through certain messaging or AI interfaces.

But an automated system capable of executing enough browser-side JavaScript may potentially generate analytics events too.

So the fact that Plausible or GA4 records a session does not, by itself, establish human presence.

This distinction becomes particularly important when studying AI retrieval. Traditional crawlers that simply request HTML often remain visible only in server or CDN logs. But the emerging ecosystem is much broader than traditional bots: search agents, browser automation, retrieval systems and user-triggered fetches do not necessarily behave like a classic crawler.

The More Interesting Signal Is What They Visit

Geography is therefore intriguing, but it is not the signal that interests me most.

The URL distribution is far more revealing.

My Direct traffic is no longer concentrated primarily on the homepage or on a handful of obvious commercial landing pages. It increasingly reaches long, deep and highly specific URLs across different languages and topical clusters.

Among them are pages about:

  • internal linking and backlinks;

  • entity SEO and GEO;

  • AI Overviews;

  • declining organic traffic;

  • B2B marketing;

  • international SEO;

  • neurodivergence and leadership;

  • spirituality;

  • luxury marketing;

  • strategic consulting.

Some of these pages receive only one or two visits. Others receive dozens or more than a hundred.

Even more interestingly, several deep articles show measurable engagement rather than a simple page hit: two, three or four minutes on the page, substantial scroll depth, and in some cases several pageviews generated by very few visitors.

That creates two completely different behavioural families in the same Direct dataset.

On one side:

1 visit → 0 seconds → 100% bounce

appearing across a large number of geographically dispersed countries.

On the other:

deep URL → several minutes → meaningful scroll depth → sometimes additional pageviews.

The first pattern could easily contain automation, failed or partial executions, privacy effects or other technical noise.

The second looks considerably more compatible with genuine consumption — although analytics alone still cannot tell us whether the consumer was a person, a browser-mediated agent or some combination of the two.

Et là je placerais juste après une phrase qui devient presque la thèse de ton article :

The first sign of retrieval eligibility may therefore not be “more AI traffic”, but a change in how widely and deeply your information corpus is being explored.

Why Deep URL Exploration Matters More Than Raw Bot Volume

This distinction matters particularly on my website because its internal architecture is unusual. It currently contains more than 17,000 internal links, connecting articles, entities, services and concepts across French, English and Spanish content.

I did not build this architecture simply to increase PageRank circulation. I have increasingly used internal linking as a way of making relationships between documents and entities explicit.

If retrieval systems are attempting to locate useful passages or documents rather than merely rank ten blue links, the relevant question changes.

It is no longer only:

“Can the crawler discover this page?”

It becomes:

“Can a retrieval system navigate the corpus sufficiently well to find the document that best answers a particular information need?”

This is also why I am increasingly uncomfortable drawing a hard boundary between SEO and GEO. Both ultimately depend on machines being able to discover, interpret and retrieve information from a structured corpus. GEO changes the interface and the selection mechanism, but it does not magically remove the technical foundations of SEO.

Important caveat: None of these analytics patterns proves that an AI system retrieved a page in response to a user query. They are observational signals. Server logs, bot identification, HTTP behaviour and referral data are needed before drawing stronger conclusions.

What My Server History Taught Me About AI Crawling

I no longer have access to the same level of server-side visibility I once had, because my current Cloudflare configuration is running in DNS-only mode. But the historical evolution of the site is still useful, because it shows how easy it is to misread automated traffic when you do not yet understand the infrastructure behind it.

When I launched the website in 2025, it was built on a completely new domain, with no previous semantic history and no established authority.

Very quickly, I started seeing traffic almost exclusively from the United States, with a noticeable concentration around places such as Council Bluffs.

At the time, I did not really understand what I was looking at.

Was it legitimate crawling? Cloud infrastructure? Automated traffic? Search-engine activity? Something else entirely?

My first reaction was defensive.

I moved the site behind Cloudflare and started limiting, then sometimes blocking, parts of that traffic.

Looking back, that may have been one of my first important technical mistakes.

I cannot prove that I blocked Google or another major crawler in a way that directly harmed rankings, but I now know that aggressive filtering rules can interfere with legitimate search-engine and AI crawling if they are not configured very carefully.

About a year later, after removing or relaxing some of those restrictions, the site's visibility began to improve.

At the same time, I was also making the site more structurally explicit: improving internal linking, adding JSON-LD, reinforcing entity relationships and creating a much denser semantic architecture.

That is when another pattern became visible.

When AI Crawling Became Hard to Ignore

As the site's structured data and internal semantic connections increased, AI-related crawlers began appearing much more frequently.

At certain moments, request volumes and crawl activity reached levels that would have looked alarming on a small shared server.

Cloudflare, however, absorbed much of that traffic without turning it into an immediate infrastructure problem.

This period taught me another important lesson: high crawl volume does not necessarily mean an attack, but it does require interpretation.

A sudden increase can come from:

  • legitimate search-engine crawling;

  • AI crawler exploration;

  • repeated retries;

  • malformed redirects;

  • bot loops;

  • cache inconsistencies;

  • or genuine malicious traffic.

And without understanding the CDN and proxy layers, it is very easy to confuse one with another.

The CDN Layer I Had Not Fully Understood

At the time, I was configuring Cloudflare without realizing that Hostinger also had its own CDN and caching layer behind the Website Builder environment.

I only understood the importance of that later, after changes in Hostinger's builder and server configuration created new behaviour that was difficult to explain.

One of the clearest symptoms was the appearance of new 404 errors in Google Search Console that did not match what I expected to see on the site.

That forced me to revisit the entire delivery chain.

After a long series of discussions with Hostinger's support system — including Kodee, its AI assistant, and likely human support staff as well — the practical conclusion was that keeping Cloudflare in proxy mode was adding too much uncertainty to the configuration.

I eventually switched Cloudflare to DNS-only mode.

And from that moment, something changed again.

The explicit crawler visibility I had become accustomed to decreased, while Direct / None traffic became much more prominent in analytics.

This is where the central question of this article really begins.

When infrastructure changes alter what you can observe, how do you distinguish between lost visibility, legitimate retrieval activity and simple technical noise?

That is the problem I am trying to solve here.

The Real Lesson: Do Not Confuse Visibility With Behaviour

Looking back at the evolution of the site, the strongest lesson is not that “AI crawlers increased because I added JSON-LD.”

That would be too simplistic.

The real lesson is that crawler behaviour is shaped by several overlapping layers at once:

  • the semantic structure of the site;

  • internal linking;

  • structured data;

  • robots directives;

  • CDN rules;

  • firewall behaviour;

  • caching;

  • redirects;

  • origin-server configuration;

  • and the crawler's own purpose.

The more complex the stack becomes, the harder it is to interpret any single metric in isolation.

And that is exactly why server logs, analytics and referral data should be treated as complementary signals rather than as interchangeable proof.

Cette version garde ton vécu, mais elle évite d’affirmer des causalités trop fortes.

Je mettrais ensuite une transition assez nette vers le cœur de l’article :

“My current challenge is therefore not to prove that AI systems visit the site — that part is obvious — but to understand whether their behaviour is becoming more compatible with retrieval, selection and eventual human referral.”

The Technical Analysis: Why AI Retrieval Can Become Visible in Analytics Without Looking Like “AI Traffic”

The first difficulty is that modern web retrieval no longer fits neatly into the old distinction between a crawler and a human browser.

For years, search engines have already been capable of rendering pages rather than simply downloading raw HTML. Google made this transition explicit in 2019, when Martin Splitt and the Search Central team announced that Googlebot had moved to an evergreen Chromium rendering engine. Since then, Google Search has been able to process JavaScript with a modern Chromium environment rather than relying solely on the initial HTML response.

By 2026, Google describes the process even more clearly: after Googlebot fetches a page, its Web Rendering Service can execute client-side JavaScript, process CSS and XHR requests, and analyse the rendered state of the document. Google also specifies that rendering is performed in a stateless environment, with local storage and session data cleared between requests.

This matters because analytics scripts are JavaScript too.

But there is an important nuance.

Google's own documentation explains that the rendering system may deliberately avoid fetching resources that are not necessary to understand the page, including reporting or error-related requests. In other words, a rendering engine being capable of executing JavaScript does not mean that GA4 or Plausible will necessarily be triggered every time a crawler renders a page.

That distinction is essential when interpreting unexplained Direct traffic.

A Headless Browser Can Trigger Analytics — But Not Every AI Fetch Does

Technically, any automated system controlling a browser environment such as Chromium can execute the same analytics JavaScript as a human browser.

This is easy to demonstrate experimentally. In 2025, Plausible ran a controlled test using Puppeteer, a headless-browser automation framework. Their simulated bot opened pages in a browser context and executed JavaScript. GA4 recorded the automated visits as legitimate traffic, whereas Plausible filtered the tested scenarios.

That experiment is particularly relevant here because it establishes something important:

JavaScript execution is not proof of human presence.

An automated browser can execute GA4. It can create pageviews. It can potentially trigger events normally associated with a browser session.

However, it would be a mistake to conclude that every ChatGPT, Perplexity or Google AI retrieval uses a headless browser and therefore appears in analytics.

The public documentation does not support that claim.

OpenAI distinguishes, for example, between OAI-SearchBot, used to help discover and surface content in ChatGPT Search, and GPTBot, whose purpose is different. OpenAI explicitly says that allowing OAI-SearchBot is necessary for content to be eligible for inclusion in ChatGPT Search.

Perplexity goes even further in documenting two different mechanisms.

PerplexityBot is a crawler used to surface and link websites in search results. But Perplexity-User is different: it can fetch a page in response to a specific user question in order to help answer that question and potentially link to the page.

That distinction gives us a much better model than simply saying “AI bots use headless browsers.”

There are now at least three observable classes of machine access:

background crawling → search indexing/discovery → user-triggered fetching

And they do not necessarily leave the same technical footprint.

From Classical Crawling to Query-Triggered Retrieval

This is where the evolution from classical search to AI-assisted search becomes important.

A traditional crawler generally operates from a crawl queue. It discovers URLs through links, sitemaps and other signals, then revisits them according to its own crawling priorities.

Retrieval systems introduce another possibility: a URL can be fetched because it is useful for answering a particular information need.

Perplexity explicitly documents this distinction. Its Perplexity-User fetcher may access a page when a user asks a question, rather than as part of general crawling.

Google's AI Mode shows the same conceptual shift from another direction. Google now describes AI Mode as using query fan-out: a user's question is split into multiple subtopics and searches are run simultaneously across several data sources.

This is much closer to retrieval behaviour than to the old mental model of “Google crawls the whole site and later ranks one page.”

The system is effectively looking for information that satisfies a particular sub-question.

That makes deep and highly specific URLs much more interesting.

Not because a deep URL automatically proves RAG, but because targeted document selection is exactly the type of behaviour retrieval-based systems are designed to perform.

RAG Changed the Unit of Interest

The conceptual foundation for this predates today's AI search products.

The Retrieval-Augmented Generation architecture introduced by Patrick Lewis and colleagues in 2020 formalised the combination of a generative model with an external non-parametric memory. Instead of relying only on information stored in model parameters, the system retrieves relevant external documents at inference time and uses them to construct an answer.

That distinction is useful for understanding what we now observe on the public web.

The important object is no longer necessarily “the website.”

It may be:

one document
one passage
one entity
one supporting fact
one URL relevant to one sub-question.

That is why a retrieval system can behave very differently from a classical crawler.

A crawler can traverse:

homepage → category → article → related article

A retrieval system may effectively jump straight to:

/en/bing-chatgpt-and-googles-ai-overviews-why-search-results-differ-and-what-it-means-for-indexing

because that document appears relevant to a specific information need.

That is exactly why the pattern I am observing — isolated visits to long and highly specialised URLs across several topical clusters — interests me more than raw traffic volume.

But Direct / None Still Does Not Mean “AI”

This is where analytics attribution becomes dangerous.

In Plausible and GA4, Direct traffic fundamentally means that the analytics platform has not received a usable source attribution.

That can happen for many reasons:

a URL was typed or pasted;
a bookmark was used;
an application opened the page without transmitting a referrer;
privacy protections removed attribution information;
a redirect lost the original referrer;
or an automated browser loaded the page without providing one.

So Direct / None is an attribution state, not a visitor identity.

And Plausible itself warns that suspicious bot traffic can appear as a sudden increase in Direct traffic, especially when sessions show close to 100% bounce and near-zero duration.

Plausible's numbers are useful here. The company says its filtering system currently blocks roughly 4 billion bot visits per month across its network and filters approximately 32,000 known data-centre IP ranges. In one comparison, Plausible reported seeing 18 times more pageviews in raw server logs than in Plausible analytics, illustrating the enormous gap between HTTP requests and traffic considered human enough to appear in analytics.

That is extremely important for my own case.

If Plausible is already rejecting such a large proportion of automated requests, the Direct sessions that remain cannot simply be equated with “bots.”

But neither can they automatically be equated with humans.

They are what remains after filtering.

The Geographic Scatter Is Interesting — But Not Proof

The global distribution of my Direct traffic is therefore intriguing.

France makes sense.

But the same dataset contains sessions from the United States, Japan, the Netherlands, Italy, Canada, Colombia, Poland, China, Sweden, Pakistan, Taiwan, Vietnam and many other countries.

Some of these sessions show:

100% bounce
0 seconds duration
one pageview

Others show:

several minutes of engagement
meaningful scroll depth
multiple pageviews

That difference matters more than geography itself.

Plausible notes that some VPN traffic may originate from IP ranges associated with data centres, and that distinguishing a genuine visitor behind a VPN from automated traffic can be difficult purely from the IP layer.

So an apparent visitor in Amsterdam, Virginia, Singapore or Tokyo could represent many different technical realities:

a person;
a company VPN;
a privacy proxy;
a cloud service;
a browser automation environment;
or an automated fetch.

Therefore I would not say:

“Global Direct traffic proves that AI agents are fetching my pages.”

I would say:

“The combination of globally distributed Direct traffic, deep-URL access and heterogeneous engagement creates a pattern that cannot be explained reliably from analytics alone.”

And that is precisely why logs matter.

B2B Makes Attribution Even Harder

The B2B hypothesis also needs to be framed carefully.

Corporate users increasingly browse behind VPNs, security gateways, managed browsers and privacy filters. Those technologies can alter or suppress normal attribution signals.

But I would avoid writing that ChatGPT Enterprise or Microsoft Copilot “deliberately strips referrers” unless we have explicit product documentation proving it.

What we can safely say is this:

the more intermediary layers exist between a user and a website, the less reliable traditional referrer attribution can become.

This problem predates generative AI. Messaging applications, email clients, security gateways and privacy-enhancing browsers have all contributed for years to what marketers loosely call dark social or unattributed traffic.

AI assistants add another intermediary layer.

A user may ask an AI system a question, receive a cited source, and then open that source from an application or protected corporate environment. Depending on how the link is opened, attribution may survive — or disappear.

OpenAI actually helps with this problem in the opposite direction: ChatGPT Search now adds utm_source=chatgpt.com to outbound search links so that publishers can identify that traffic.

Therefore, when that parameter is present, the attribution is relatively clear.

When it is absent, we should not reverse the logic and conclude:

“there is no referrer, therefore this probably came from AI.”

That would be methodologically weak.

What Has Actually Changed Between 2019 and 2026

Seen over several years, the evolution becomes clearer.

In 2019, the major technical transition was rendering. Googlebot became evergreen Chromium and modern JavaScript became part of normal search processing.

By 2020, RAG research formalised the use of external non-parametric memory alongside generative models.

By 2024–2025, AI search products increasingly exposed retrieval directly to users: instead of merely returning ranked links, systems could retrieve information and synthesise answers.

By 2026, the distinction has become explicit in crawler documentation.

OpenAI separates search discovery from potential training through different user agents.

Perplexity separates background indexing through PerplexityBot from user-triggered access through Perplexity-User.

Google publicly describes AI Mode as executing simultaneous searches across multiple data sources through query fan-out.

The web therefore no longer has one simple machine visitor.

It has an ecosystem of:

  • index crawlers

  • renderers

  • search bots

  • retrieval fetchers

  • browser agents

  • human users arriving through AI interfaces.

That is why classical web analytics has become increasingly ambiguous.

So What Would Actually Count as an Early Retrieval Signal?

Not one metric.

I would look for convergence.

A stronger early signal would be something like this:

**deep URL diversity increases

  • machine-access logs confirm legitimate search/retrieval agents

  • error rates remain low

  • the behaviour is recurrent rather than a one-off spike

  • several semantic clusters are explored

  • some AI referrals eventually appear

  • a portion of those referrals shows genuine human engagement.**

Only then does the hypothesis become interesting.

And even then I would call it: retrieval eligibility rather than: proof of retrieval authority.

That wording is more defensible, and it actually makes the article more sophisticated.

AI retrieval does not have a single analytics signature. What we can observe is a convergence of server access, deep-document exploration, user-triggered fetching, referral attribution and downstream human engagement.

Conclusion — Retrieval Eligibility Is a Pattern, Not a Single Metric

Through all this noise — noise that Plausible is supposed to filter, and that my entity reconciliation work was also meant to reduce — two things stand out.

First, my website is clearly being discovered and reused within its professional ecosystem. Other SEO practitioners cite my work. Natural backlinks appear without outreach. Some Direct / None visitors stay for more than a minute, and occasionally for ten, twenty or even thirty minutes on highly specialised articles. That is difficult to reconcile with the idea that everything I am seeing is meaningless bot noise.

Second, discovery is no longer limited to one language or one obvious section of the site. Articles in French, English and Spanish are being reached across very different semantic clusters. Google continues to test and reassess my entity against a wide variety of queries, to the point where Search Console data can look almost incoherent to a novice: impressions appear across topics, languages and intents that do not fit neatly into a conventional keyword-ranking report.

To someone used to analysing SEO, GEO and marketing together, however, the picture is more intelligible.

The site is being evaluated not only as a set of isolated pages, but as an increasingly dense information corpus built around entities, topics, relationships and expertise. That does not prove that every unexplained session is an AI retrieval event. It does suggest that the domain is being explored far beyond the handful of commercial queries I originally expected to target.

And I have a useful counterexample.

I also manage a WordPress website for a parish. Its domain may well have stronger conventional authority signals than my own professional site, and its content is perfectly indexable. Yet its crawl patterns, Direct traffic and Search Console behaviour tell a very different story. It does not attract anything close to the same intensity of automated exploration or unexplained deep-URL activity.

That comparison matters.

It suggests that the frenetic activity observed on my own domain is not simply the inevitable consequence of putting any website online, nor of having an older or more authoritative domain.

Something about the content itself, the semantic structure of the corpus, the entity relationships, the internal linking and the topics being covered appears to influence how often machines — and perhaps some very discreet humans — come looking for information.

This still leaves uncertainty.

Some of that activity may be crawlers.
Some may be retrieval systems.
Some may be automated browsers.
Some may be privacy-protected users.
Some may simply be people reading quietly without leaving an obvious attribution trail.

But that uncertainty does not make the data useless.

It changes the question.

Instead of asking:

“Can I prove that this visit came from an AI?”

the more useful question becomes:

“Is my information corpus increasingly discoverable, retrievable and useful enough to be repeatedly explored by search engines, AI systems and eventually real users?”

In my case, the answer appears to be increasingly yes.

And perhaps that is the most practical definition of early retrieval eligibility we can currently observe from outside the black box:

not a sudden burst of AI referrals, but the gradual transition from isolated indexing to recurrent, multilingual and deep exploration of a coherent information corpus.

For SEO and GEO, that may be one of the most important signals to learn to read over the next few years.

SEO & GEO Technical FAQ: Deciphering AI Bots, RAG, and Analytics

How does AI Retrieval-Augmented Generation (RAG) technically differ from traditional search engine crawling?

Traditional crawling relies on background indexers operating from a queue (e.g., standard Googlebot discovering URLs via sitemaps and links) to map domains for a centralized index. In contrast, AI search engines use Retrieval-Augmented Generation (RAG). As formalized in the foundational 2020 paper by Patrick Lewis and his team, RAG architecture dynamically retrieves external, non-parametric memory at inference time to construct an answer. From a technical and patent perspective, systems like Google's AI Overviews utilize mechanisms such as "query fan-out," splitting a user's prompt to run simultaneous searches across various data sources. Consequently, modern web retrieval introduces real-time, user-triggered fetching (documented by user agents like Perplexity-User or OAI-SearchBot), which targets deep, highly specific URLs to extract a single entity or passage rather than crawling a site sequentially.

Why do AI search agents often appear as "Direct / None" traffic in platforms like GA4 or Plausible?

Analytics attribution relies heavily on HTTP referrer headers. Modern AI retrieval systems frequently use headless browser environments (like Puppeteer or Playwright) to fetch and render pages accurately. As Martin Splitt from the Google Search Central team highlighted in 2019, search bots have moved to evergreen Chromium rendering engines capable of executing client-side JavaScript to understand dynamic state. If an AI agent fetches a document and its headless browser executes the analytics JavaScript without passing a referrer—or if a human clicks a citation from a privacy-shielded enterprise AI tool that deliberately strips referral data—the visit is classified as "Direct / None". The execution of JavaScript is therefore not absolute proof of human presence, but an artefact of how modern rendering engines and privacy protocols interact with tracking tags.

Does Generative Engine Optimization (GEO) replace the technical foundations of traditional SEO?

No, the two disciplines are inextricably linked and rely on the same machine-readable web. While GEO shifts the focus toward how generative models select and synthesize answers at the interface level, it depends entirely on the structural integrity of the underlying data. As noted by Lydie Goyenetche, consultante SEO & GEO, GEO changes the interface and the selection mechanism, but it does not magically remove the technical foundations of SEO. A strong semantic architecture, dense internal linking, and clear entity relationships are what make a corpus "retrieval eligible." AI systems look for structured passages, schema markup (JSON-LD), and factual entity associations to confidently build their answers, making the technical rigor of SEO the fundamental prerequisite for GEO visibility.

EUSKAL CONSEIL

9 rue Iguzki alde

64310 ST PEE SUR NIVELLE

07 82 50 57 66

euskalconseil@gmail.com

Mentions légales: Métiers du Conseil Hiscox HSXIN320063010

CGV & Mentions légales

Ce site utilise uniquement Plausible Analytics, un outil de mesure d’audience respectueux de la vie privée. Aucune donnée personnelle n’est collectée, aucun cookie n’est utilisé.

communication digitale
communication digitale