Exploring Beyond Semantic Silos

Discover how I moved beyond traditional semantic silos and 17,000 internal links to explore entities, vector spaces, and identity resolution. This empirical investigation delves into modern search systems and Google’s machine logic, providing insights into the future of search.

VEILLE MARKETINGMARKETING

LYDIE GOYENETCHE

8/30/202636 min read

One day, after finally building the semantic silos across the four languages of my website, completing my blog articles and working extensively on the internal linking structure, I was confronted with a rather puzzling observation.

The search engine kept sending visitors to almost always the same articles. My SEO content, however, remained surprisingly discreet. In Google Search Console, SEO-related keywords and topics I had developed extensively barely appeared, with only low impression volumes.

The paradox was all the more striking because my SEO and marketing content was numerically more abundant, often more specialized, focused on niche topics and, in some cases, supported by backlinks from industry professionals.

On paper, everything seemed coherent.

But Google appeared to have built a very different interpretation of my website.

That was when I tried something I had never been taught in my SEO training. I extracted my Google Search Console data and asked Gemini to visually map my entity.

And suddenly, the problem became visible.

My semantic universe appeared fragmented across several disconnected clusters, with a strong imbalance in favor of a small number of already successful articles. That imbalance did not reflect the number of articles I had published, the semantic richness of the website, or even the professional positioning I had deliberately chosen to build.

Because I am not a web developer, and because entity SEO had never been part of what I had learned up to that point, I simply had not been asking the right questions.

I was still thinking in terms of the traditional SEO framework: keywords, semantic silos, internal linking and backlinks.

Yet one question was becoming increasingly difficult to ignore: what if my CMS itself had structural limitations capable of neutralizing part of that work?

At the time, my website contained nearly 17,000 internal links. The internal linking was there. The content was there. The keywords were there. The backlinks were there too.

So why did the search engine keep favoring certain areas of my website while almost ignoring the ones that best reflected the profession I actually wanted to practice?

Why invest so much technical and editorial effort if, in the end, I was allowing the search engine to identify me mainly according to the signals it found easiest to process, rather than according to the profession, expertise, ethical direction and professional positioning I had consciously chosen?

By applying SEO as I had been taught, I was getting some results.

But not the SEO results I had hoped for.

That was when my way of thinking started to change.

Perhaps the problem was no longer simply about optimizing pages or linking them together more effectively.

Perhaps I first needed to make explicit the identity all those pages were supposed to build.

In other words, I had built semantic silos.

But perhaps I had not yet built an entity.

 Instagram, User Signals, and the Invisible Architecture of a Website

When an Instagram campaign began to change the way I understood SEO

Somewhat exhausted and increasingly confused by the results of my website, I eventually tried something that, at first glance, had very little to do with the semantic silos and roughly 17,000 internal links I had patiently built.

I used Instagram.

I launched an SEO-focused advertising campaign targeting New York, with posts sending users directly to my website.

The visitors coming from Instagram did not stay particularly long. They did not become a loyal audience either: as far as I could tell, only two returned afterwards.

And yet, something else caught my attention.

I started receiving more visitors from the United States, especially from New York. Then, after the campaign ended, I noticed that this American presence did not disappear entirely and that some of my website’s rankings in the US were improving.

I obviously could not conclude:

Instagram made my website rank higher on Google.

That would be confusing correlation with causation.

But the experience forced me to consider a much more interesting hypothesis: perhaps I had been giving too much importance to the signals that were easiest for me, as an SEO consultant, to see — keywords, backlinks, internal linking — and not enough to the much wider set of signals that the search engine itself could potentially process.

What we now know about Google’s use of user signals

For a long time, the role of clicks in Google rankings remained particularly opaque.

US antitrust proceedings against Google have since made much more precise information publicly available.

In testimony made public in January 2025, Pandu Nayak, Google Search Vice President and engineer, described Navboost as a traditional signal that uses, among other things, how often users click on a document for a given query. The document also states that this data can be segmented by location and device type, and that Navboost uses the most recent 13 months of data. (justice.gov)

This absolutely does not mean that a visitor arriving from Instagram directly becomes a Navboost signal: Navboost is related in particular to interactions with search results.

But it confirms something far more important for my reasoning: Google does not build rankings solely from page text and backlinks. Aggregated behavior from its own users is also part of its systems.

And this is not a recent idea.

A Google patent with a priority date going back to June 28, 2004, titled Deriving and using interaction profiles, describes systems capable of using different interaction profiles, including click duration, repeated clicks, query reformulation and other behavioral information. The inventors named are Alexis Jane Battle, David Ariel Cohn, Carrie Elizabeth Grimes and John Ogden Lamping. (patents.google.com)

The patent distinguishes, for example, between shorter and longer clicks and considers the possibility of using such information to estimate satisfaction with a result. It also explicitly warns that a single click is not a sufficiently reliable indicator and that multiple metrics may need to be combined. (patents.google.com)

That distinction matters.

SEO therefore cannot be reduced to a simplistic formula such as:

more time on page = better rankings.

The reality is much more complex.

The famous “dwell time”: useful for thinking, dangerous as an SEO certainty

The term dwell time is often used too loosely in SEO.

Yet several Google-assigned patents do describe mechanisms that distinguish between short and long interactions.

Patent US9223868B2, published on December 29, 2015, describes, in some implementations, interactions of under 80 seconds as short, interactions between 80 and 200 seconds as medium, and interactions above 200 seconds as long. (patents.google.com)

Another Google patent, US8326826B1, also discusses the use of query and click logs and gives, in one example, a threshold of 30 seconds to distinguish a short click from a long one. (patents.google.com)

These numbers should absolutely not be interpreted as SEO thresholds to optimize for.

They simply show that Google has long studied interaction sequences, not merely the existence of a click.

In October 2023, Pandu Nayak also explained under oath that Google has access to a very large volume of session logs and that there is a constant trade-off between the additional value of those data and the computational cost of processing them. (justice.gov)

This was precisely what interested me.

I was no longer trying to determine whether a handful of visitors from New York had individually “sent an SEO signal.”

I was beginning to understand that the search engine was observing an informational environment far richer than the one I was manipulating every time I moved a few links inside a semantic silo.

Then I discovered that my architecture was not the architecture I thought I had built

It took me roughly six more months to understand another important element.

At the time, my website was running on Hostinger Website Builder.

In my mind, I had built an architecture.

I had pillar pages.

Child pages.

Semantic silos.

SEO topic areas.

Marketing sections.

Specialized content.

Carefully designed internal bridges.

But the conceptual architecture I had imagined was not necessarily materialized in the same way inside the technical structure of my CMS.

Hostinger itself explains that the homepage normally lives directly at:

domain.tld

while other pages typically take forms such as:

domain.tld/contact

domain.tld/about

and so on. (support.hostinger.com)

The builder allows pages to be organized in navigation and supports links and dropdown menus. (support.hostinger.com)

But what I mentally imagined as:

/seo/
↳ /seo/local-seo/
↳ /seo/entity-seo/
↳ /seo/technical-seo/

could in reality look much more like:

/seo
/local-seo
/entity-seo
/technical-seo

This obviously does not mean that Google considers all of those URLs equivalent.

Google has other ways to reconstruct hierarchy: internal links, navigation, anchors, content, sitemaps, semantic context and external signals.

Hostinger also automatically generates a sitemap.xml, specifically designed to help search engines discover pages and understand the general structure of the website. (support.hostinger.com)

But I could no longer assume that the hierarchy I had designed mentally would automatically become the hierarchy understood by the search engine.

And that was when the real problem appeared.

I had built an editorial hierarchy. Google still had to reconstruct it.

I thought I had told Google:

this is my core expertise; these are my pillar pages; these are the secondary contents supporting them.

In reality, what I had mainly created was a large number of documents and a large number of relationships between them.

Google still had to determine which ones mattered most.

It had to reconstruct the structure.

Evaluate the pages.

Observe the links.

Compare the topics.

Interpret behavior.

Understand which queries each document answered.

And ultimately decide which version of the website deserved to appear in search results.

Some parts of my website already had advantages.

They received more traffic.

Some had gained visibility faster.

Users had found them useful.

And one geographic area — the United States — was beginning to generate more signals than I had originally expected.

That geographic dimension becomes particularly interesting when compared with documents made public in 2025: Google describes Navboost as being able to use click data segmented by location and device type. (justice.gov)

Again, this does not prove that my Instagram campaign created my US rankings.

But it makes the question I was beginning to ask far less absurd:

could the real geography of interactions around a young website contribute, alongside many other signals, to the space in which Google begins testing its documents?

The irony: my social accounts were saying “United States” too

There was another detail that was almost comical.

My Instagram and Facebook accounts had, without my really intending it, been configured around the United States. I had never managed to fully correct that setting.

On one side, I was carefully building a website intended to represent a very specific professional activity.

On the other, some peripheral signals were potentially telling a different story.

Advertising targeted at New York.

American visitors.

Social accounts configured around the US.

Early improvements observed in American rankings.

Taken individually, each of these signals was weak.

Taken together, they raised a question that my traditional training in keywords and internal linking had never really taught me to ask:

who constructs the semantic and geographic identity of a website when its owner does not formalize it clearly enough?

The search engine is under no obligation to choose the profession I chose

This was probably the most important change in perspective during that period.

I had spent an enormous amount of time saying:

  • this page is about SEO.

  • this one is about GEO.

  • this one supports my pillar page.

But I had not worked enough on a much more fundamental question:

what allows a machine to understand that all these pages represent the expertise of the same person and the same organization?

Because a search engine is under no obligation to consider central what the author of the website personally considers central.

If it encounters several semantic universes, a technically weak or ambiguous hierarchy, and very different performance levels across content, it has to make its own decisions.

And Google has a very large number of signals available for doing so.

A public document from the US antitrust proceedings stated in 2025 that Google uses more than 100 raw signals, which are then aggregated into various higher-level signals. Systems cited include PageRank, Navboost, Q*, RankEmbed, RankBrain and DeepRank. (justice.gov)

The search engine I was trying to influence with internal linking was therefore far more complex than the mental model with which I had begun learning SEO.

That realization led to a second, much more technical question:

if my CMS does not make the relationships defining my business explicit enough, can I express them in another way?

That was the point at which I gradually began moving away from a model focused only on pages and started working instead on entities, identifiers and relationships.

I was in a difficult position.

I wanted to offer SEO and B2B marketing services designed to generate qualified traffic and reduce customer acquisition costs for companies. Yet, judging by my own website, I was struggling to achieve precisely what I intended to sell.

There was another problem too.

I did enjoy SEO — but mostly through its editorial, strategic and creative dimensions. The more technical side often felt strangely artificial to me: a succession of optimizations, rules and adjustments that I could apply without always feeling that I truly understood the deeper logic behind them.

So I was faced with a fairly simple choice.

Either I stopped.

Or I kept going.

I chose to keep going.

But from that point on, I stopped relying only on what I had been taught. I began learning empirically: forming hypotheses, running tests, observing what changed, comparing signals, making mistakes, correcting them, and trying to understand why the search engine behaved the way it did.

That was when SEO became genuinely fascinating to me.

Because I was no longer simply applying a method.

I was trying to understand a system.

When Google Started Looking More Like Geometry Than a Crawler

I began to suspect that I was looking at the wrong model

As I continued observing the behavior of my website, something increasingly bothered me.

The traffic patterns I was seeing did not really resemble the mental model I had been taught.

I had imagined Google mainly as a system crawling pages, following links, identifying semantic relationships and progressively understanding the architecture I had carefully built.

In that model, my semantic silos mattered enormously.

My internal linking mattered.

My pillar pages mattered.

And logically, if I strengthened the connections between those elements, the search engine should progressively understand which subjects were central to my professional positioning.

Except that what I was observing did not behave quite like that.

Some pages seemed to gain visibility despite not occupying the strongest positions inside my carefully designed architecture.

External signals appeared capable of shifting the geography and thematic distribution of my traffic.

Certain documents seemed to be tested in spaces that I had not deliberately created through internal linking.

And the resulting movements did not feel linear.

They felt spatial.

I remember thinking that Google seemed to be performing some kind of three-dimensional calculation.

I did not have the mathematical vocabulary for it.

I had never seriously studied vector mathematics.

But intuitively, what I was observing looked less like a robot walking through a network of links and more like a system calculating distances, proximity, directions and relationships between pieces of information.

Years later, this intuition no longer seems as strange as it initially did.

Modern information retrieval systems increasingly represent language and information mathematically.

Google's own researchers played a major role in that evolution.

In 2018, Google researchers Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova published BERT — Bidirectional Encoder Representations from Transformers. The model represented language contextually rather than treating words merely as isolated strings.

One year later, in October 2019, Google Search Vice President Pandu Nayak publicly announced the deployment of BERT in Google Search to improve the engine's understanding of queries and the relationships between words within them.

What mattered to me was not the acronym BERT itself.

It was the underlying conceptual shift.

Search was no longer adequately described as:

word → page containing that word.

Meaning could be represented from context.

Relationships could be learned.

And similarity could increasingly be calculated rather than merely declared through exact lexical matches.

That distinction would eventually change almost everything about the way I approached SEO.

I started asking AI a different kind of question

At that particular moment, I was working on the Spanish version of my website.

More specifically, I was linking my Spanish pillar pages with the corresponding blog articles, trying once again to strengthen the semantic coherence of the structure.

But while doing it, I had begun questioning the model itself.

Instead of asking AI:

How should I improve my internal linking?

I started asking questions closer to:

How does Google mathematically determine that two documents, concepts or entities are close to one another?

Or:

If two pages are not directly connected by my website architecture, can Google nevertheless determine that they belong to the same semantic space?

And eventually:

Does Google represent meaning mathematically?

Those questions led me somewhere I had not expected.

Through many conversations, searches, mistakes and reformulations, I started encountering terms such as embeddings, vector spaces, entity reconciliation, Knowledge Graph identifiers and Machine IDs.

And then I discovered the KGMID.

I almost fell off my chair.

The KGMID changed the scale of the problem

A Google Knowledge Graph Machine ID — historically appearing in forms such as /m/... or /g/... — identifies an entity inside a knowledge graph rather than identifying a web document.

That distinction seems small until one realizes what it changes.

A URL answers:

Where is this document?

An entity identifier attempts to answer:

What thing does this information refer to?

This idea had already been developing inside Google's research and patent ecosystem for years.

Google patent US9336211B1, Associating an entity with a search query, describes mechanisms for identifying entities and associating search queries with those entities rather than treating every search solely as an isolated collection of strings.

This is precisely the distinction that Bill Slawski, one of the SEO industry's most respected analysts of Google patents, spent years emphasizing.

His work repeatedly drew attention to Google's transition from systems centered largely on strings and links toward systems capable of identifying things, attributes and relationships between things. His contribution was important precisely because he treated patents not as secret ranking-factor lists, but as evidence of the problems Google's engineers were attempting to solve.

And this distinction radically changed the way I looked at my website.

I had created relationships between pages.

But that did not automatically mean that Google had resolved all those pages as evidence about one clearly identified person, one organization, one body of expertise, and one coherent set of relationships between them.

My internal linking could tell Google:

Article A → Pillar Page B

But my real professional identity required something considerably richer:

Lydie Goyenetche → founder / consultant → Euskal Conseil

Euskal Conseil → offers → SEO / GEO / B2B marketing

Lydie Goyenetche → has expertise in → entity SEO

Article A → authored by → Lydie Goyenetche

Article A → about → entity SEO

Euskal Conseil → serves → French / Spanish / international markets

These are not simply hyperlinks.

They are typed relationships between entities.

And that was the moment when my 17,000 internal links suddenly looked rather different to me.

They were not useless.

Far from it.

But they were solving only one class of relationship.

Google had been studying entity associations for years

The more I investigated, the more I realized that the problem I had intuitively stumbled upon was not marginal.

Google patents had been exploring entity-query relationships, entity detection and semantic associations for years.

The patent family around Associating an Entity with a Search Query, for example, describes identifying a particular entity and determining queries associated with that entity.

The significance for SEO is subtle.

A search engine does not necessarily have to understand every page independently.

It can attempt to determine:

  • what entities appear in the information;

  • how those entities relate to queries;

  • what attributes belong to those entities;

  • what other entities they are associated with;

  • and whether different mentions are actually references to the same underlying thing.

This moved the problem far beyond keyword optimization.

It became a problem of identity resolution.

And an identity-resolution problem cannot be solved simply by repeating the same keyword on fifty interconnected pages.

Then came vectors

The second shock was discovering that modern information retrieval does not necessarily depend on exact words being connected in an explicit chain.

Information can also be represented numerically.

An embedding is a numerical representation in which information is encoded as coordinates in a multidimensional space.

The mathematical details are substantially more complicated than the simplified diagrams normally used to explain embeddings, but the basic principle was enough to transform my understanding:

semantic similarity can become mathematical proximity.

Two concepts can be represented as relatively close even if they are not expressed using exactly the same words.

This is precisely why the term vector space kept appearing in my conversations with AI.

What I had intuitively imagined as three dimensions was, mathematically, potentially hundreds or thousands of dimensions.

The human brain cannot visualize such a space.

A machine does not need to.

It can calculate within it.

This is one reason why Google's adoption of neural language models was so significant.

When Pandu Nayak explained Google's deployment of BERT in Search in 2019, the central point was that the system could interpret words in relation to all the other words around them rather than processing meaning through a simplistic sequence of isolated terms.

And the picture became even clearer through information released during the US antitrust proceedings.

SEO analyst AJ Kohn, commenting on Pandu Nayak's 2023 testimony, noted the importance of systems including RankEmbed-BERT, DeepRank and RankBrain, and summarized RankEmbed in particularly useful terms:

“Embed is clearly about word embeddings and vectors.”

Kohn's formulation is simple, but its implication is enormous: part of modern ranking involves transforming language into mathematical representations that machines can compare.

This does not mean:

Google converts my entire business into one vector and ranks that vector.

That would be an unjustified simplification.

But it does demonstrate that my original mental model — a crawler simply walking through my internal-link structure — was dramatically incomplete.

The geometry I had imagined was not entirely imaginary

This was the part that fascinated me most.

When I had looked at my traffic and imagined some strange three-dimensional calculation, I obviously had not independently discovered vector mathematics.

But I had noticed a behavior that a purely hierarchical model did not explain very well.

A tree works like this:

Homepage

SEO pillar

Entity SEO article

A vector space works conceptually differently.

An SEO article may simultaneously be close to:

SEO

entity resolution

Knowledge Graph

digital marketing

Google Search

structured data

B2B acquisition

France

Spain

and perhaps dozens or thousands of other dimensions.

There is no requirement that those relationships follow the menu structure I designed.

That was a profound realization.

My website architecture was hierarchical because humans and CMSs like hierarchies.

Machine representations did not necessarily have to be.

RankEmbed, DeepRank and the disappearance of my simple SEO diagram

The antitrust testimony was particularly striking because it exposed something SEO practitioners rarely get to see directly: several different stages of Google's ranking pipeline.

Pandu Nayak described traditional signals, machine-learning systems and deep-learning systems operating at different stages of retrieval and ranking.

AJ Kohn's detailed analysis of that testimony highlights systems such as Navboost, RankBrain, RankEmbed-BERT and DeepRank, while also stressing that different systems operate on different candidate sets and at different moments in the ranking process.

This matters enormously.

It means that there may not be one single calculation called “the Google algorithm”.

There are successive transformations.

Candidate retrieval.

Filtering.

Scoring.

Semantic interpretation.

Machine-learning adjustments.

Behavioral signals.

Reranking.

And potentially other systems about which the public knows very little.

My semantic silo therefore existed inside a much larger computational environment.

It was not wrong.

It was simply not the whole machine.

My semantic silo was a map. Google was building representations.

This is where my original SEO model began to crack.

A semantic silo is something I construct.

I decide that Article A supports Pillar Page B.

I create links.

I choose anchors.

I define navigation.

I decide what is parent and what is child.

But Google does not have to accept my map as reality.

It can build its own representations from other evidence.

Content.

Links.

Queries.

User interactions.

Language.

Location.

External mentions.

Structured data.

Known entities.

Relationships between entities.

And learned numerical representations of language.

My architecture was therefore not the system.

It was one source of evidence presented to the system.

That sentence seems obvious to me today.

At the time, it was not obvious at all.

And it echoes something Bill Slawski's patent analysis repeatedly demonstrated: the search engine's task is not simply to reproduce the architecture supplied by a webmaster, but to infer what entities, meanings and relationships the available evidence actually supports.

And that exposed another weakness in my website

The more I understood entity resolution, the more uncomfortable the situation became.

My website contained extensive content about SEO.

It contained marketing content.

It contained professional information about me.

It described Euskal Conseil.

It linked articles to pillar pages.

But the existence of those elements did not guarantee that Google would reconcile them into the exact entity structure I had in mind.

Nor could I simply decide:

I want a KGMID.

An identifier is the result of a machine identifying or reconciling something sufficiently clearly inside its own knowledge system.

I could supply evidence.

I could clarify relationships.

I could reduce ambiguity.

But I could not simply create Google's internal certainty by declaration.

This distinction also explains why structured data fascinated me later.

Schema markup can tell a machine what relationship I claim exists.

But the machine can still compare that declaration with the rest of the web.

The declaration:

Person → worksFor → Organization

is not equivalent to independent evidence supporting that relationship.

The distinction between declared semantics and corroborated semantics would eventually become central to the way I approached entity SEO.

What respected SEO specialists had already understood

At this point I discovered that some people in the SEO industry had been moving toward this systems-based interpretation for years.

Bill Slawski had spent more than a decade analyzing search patents relating to entities, query interpretation and semantic search.

AJ Kohn's reading of the Nayak testimony similarly focused not on isolated “ranking factors”, but on the architecture of information retrieval and the succession of systems through which documents are selected and reranked. He argues that understanding those mechanics makes him a better SEO because it changes how he interprets volatility, relevance and experimentation.

More recently, Mike King, founder of iPullRank, has described this broader approach as relevance engineering — moving SEO away from mechanical optimization toward understanding the systems that determine whether information is retrievable, interpretable and relevant.

That expression resonated strongly with what I was discovering.

I had started SEO by modifying pages.

I was slowly becoming interested in the machinery that decided what those pages meant.

And that was very far from what I had initially learned

This contrast was particularly striking because my initial SEO training had not been poor.

I had followed professional SEO training with Paul Grillet, a recognized French SEO practitioner.

I had learned useful and legitimate techniques.

Keywords mattered.

Internal linking mattered.

Backlinks mattered.

Technical optimization mattered.

None of that had suddenly become false.

My mistake had been more subtle.

I had confused techniques that influence a search engine with a model explaining how the search engine works.

Those are not the same thing.

Internal linking can provide structure without being the entire representation.

Keywords can provide lexical evidence without defining meaning by themselves.

Backlinks can transmit authority without resolving every entity.

Structured data can state relationships without forcing Google to believe them.

And embeddings can help represent semantic similarity without replacing links, entities or behavioral signals.

Search was beginning to look less like a checklist and more like a multi-layer information-retrieval system.

And once I understood that distinction, I could no longer approach SEO in quite the same way.

I had begun with a cocoon.

I had discovered a graph.

And behind the graph, I was beginning to see vectors.


 I Could Not See Google’s Vector Space, I Could Still Change the Signals Entering It

I could not discover Google’s dimensions — but I could observe what seemed to move the system

Once I had started thinking in terms of embeddings, vectors and entities, another problem immediately appeared.

If Google was representing information inside some kind of multidimensional space, what were the axes of that space?

At first, I imagined something almost naïve:

one dimension for SEO,

one for geography,

one for expertise,

one for authority,

one for user behavior,

one for language,

and perhaps different weighting factors connecting all of them.

Mathematically, that is probably not how modern embeddings work.

The dimensions learned by neural models are generally not human-readable variables called “SEO”, “authority” or “France”. They are latent mathematical dimensions learned from data.

So I could not realistically identify Google's vector axes.

Google was obviously not going to publish them either.

What I could do, however, was something much more practical:

observe which changes in the information environment around my website appeared to alter the way machines interacted with it.

From that point on, experimentation became more important than the SEO rules I had originally learned.

I started imagining different confidence levels between signals

Another intuition followed naturally.

If Google uses many different classes of information, it seemed unlikely that every signal carried the same weight or the same degree of confidence.

An internal link.

A geographic signal.

A Business Profile.

An external mention.

An explicit entity identifier.

A structured relationship.

A query interaction.

A known organization.

A known person.

Surely those pieces of evidence could not all be treated identically.

I therefore started imagining the Knowledge Graph not simply as a set of nodes and edges, but as a system in which some relationships would be more strongly supported than others.

Interestingly, Google's own Enterprise Knowledge Graph documentation now describes entity reconciliation using concepts such as clusters, similarity or distance, and confidence scores between 0 and 1. Google explains that an entity located further from other members of its assigned cluster receives a lower confidence score.

This is not documentation of Google Search ranking.

That distinction matters.

But it demonstrates that within Google's own knowledge-graph technology, entity resolution can indeed be treated as a problem involving distance, clustering and confidence, rather than merely binary connections.

That was remarkably close to what I had begun to imagine empirically.

I started by connecting the three strongest pieces of my digital identity

At that stage, three digital properties represented my professional identity particularly clearly:

my website,

my Instagram account,

and my Google Business Profile.

My Google Business Profile was particularly interesting because it already existed inside Google's ecosystem and possessed its own persistent identifier.

I therefore began experimenting with custom JSON-LD.

My first objective was simple:

make explicit that these separate digital objects referred to the same professional identity.

A normal HTML link essentially tells a machine:

Document A → URL B

JSON-LD allowed me to express something more structured:

Person → worksFor → Organization

Organization → url → WebSite

Person → sameAs → external profile

Article → author → Person

Article → about → Concept

The distinction may look small.

It is not.

A hyperlink creates a connection between resources.

A semantic graph can also specify what type of relationship exists between those resources.

I was beginning to move from linking documents to describing entities.

Google had been working on entity normalization long before I discovered it

When I began researching the technical background, I realized that the problem I was encountering was not new at all.

One particularly revealing Google patent is “Entity normalization via name normalization” — US20070198600A1, later granted as US8700568B2.

Its priority date goes back to February 17, 2006.

The patent describes methods for examining objects and facts, identifying candidate objects that may refer to the same thing, comparing them, and normalizing or merging representations when they appear to represent the same underlying entity.

This immediately reframed what I was trying to achieve.

My problem was no longer merely:

How do I tell Google what this page is about?

It was also:

How can a machine determine that different mentions, profiles, documents and attributes refer to the same underlying person or organization?

That is an entity-resolution problem.

And entity resolution is much more fundamental than keyword repetition.

Then I discovered that Google also associates entities with queries

Another Google patent made the picture even clearer.

“Associating an entity with a search query” — US9336211B1, with a priority date of March 13, 2013, describes systems for identifying queries associated with a particular entity and, conversely, identifying an entity associated with a query.

The patent family also includes later versions such as US9870423B1, US10789309B1 and US11294970B1, all dealing with relationships between search queries and entities.

This mattered enormously for my reasoning.

Because the problem was no longer simply:

keyword → document

It could also involve relationships such as:

query → entity

entity → associated queries

query → documents associated with entity

That is a completely different model from the simple crawler I had imagined at the beginning of my SEO journey.

My internal linking could influence document relationships.

But the search engine could simultaneously be working on another layer entirely:

what entity does this information refer to, and for which queries is that entity relevant?

I began coding thematic authority nodes, not only internal links

At the same time, I still had a large amount of traditional SEO work to complete.

I was linking my Spanish pillar pages to their corresponding blog articles.

I was strengthening thematic connections.

But now, instead of thinking only in terms of anchor text and PageRank circulation, I began experimenting with a second layer.

I started coding structured thematic relationships inside my pages.

A blog article could be represented as an Article.

It could have an author.

That author could reference the same persistent Person node.

The Person could be connected to an Organization.

The article could reference concepts through about or mentions.

The organization could reference the website.

Instead of hundreds of isolated declarations, I gradually started thinking about reusable nodes.

Conceptually:

Article

↓ author

Person

↓ worksFor

Organization

↓ url

WebSite

and simultaneously:

Article

↓ about

Concept

This was beginning to resemble a graph rather than a collection of pages.

My first implementations failed because my website builder was not a server-side environment

Of course, understanding the idea did not mean I immediately knew how to implement it correctly.

My first codes failed.

I was using Hostinger Website Builder, and I initially treated custom code almost as though I were working inside a Cloudflare Worker.

That was technically wrong.

A Cloudflare Worker operates at the HTTP request/response layer and can manipulate what is delivered to the user or crawler.

A Website Builder gives me a much more constrained injection point inside the final HTML.

The layers are not equivalent.

For JSON-LD to be correctly interpreted as structured data inside an HTML document, I needed to wrap the JSON inside:

<script type="application/ld+json">

and then close the script properly.

That sounds elementary.

For a developer, it probably is.

For me, it was another important realization:

semantic modeling only matters if the machine actually receives syntactically valid data.

The graph in my head was irrelevant if the browser or crawler only saw broken code.

The first signal that caught my attention was not ranking — it was crawling

Something unexpected happened while I was doing this work.

It was not an immediate increase in human traffic that caught my attention.

It was machine traffic.

As I progressively updated pages with structured data and entity relationships, I started seeing much more crawler activity on my server.

Search-engine bots appeared repeatedly.

AI-related crawlers also became much more visible.

And what particularly intrigued me was that the activity often seemed concentrated around pages I had recently modified.

At first, I was tempted to think:

the entity reconciliation code is causing the crawling.

Today I would never state that as a fact.

That would confuse correlation with causation.

A crawler can revisit a page because its content changed.

Internal links may have changed.

A sitemap may have been refreshed.

The URL may have received new signals.

Historical crawl patterns may matter.

Different bots also operate according to completely different policies.

So I cannot demonstrate that my JSON-LD itself caused the increased crawling.

What I can say is much narrower: during a period of intensive structural modification, I observed a strong increase in machine activity, particularly around recently modified pages.

That observation encouraged me to continue experimenting.

At first, I thought one global identity block would be enough

My first assumption was quite logical.

Why repeat the same information everywhere?

I thought I could place one major identity graph in the global <head> of the website and explain:

  • who I was,

  • what Euskal Conseil was,

  • which profiles corresponded to us,

  • and how the organization and person were related.

But gradually, I understood that I had confused two different problems.

The first problem was:

Who is the entity?

The second was:

How is this specific document related to the entity?

A global declaration such as:

Lydie Goyenetche → worksFor → Euskal Conseil

does not automatically express: Article 143 → author → Lydie Goyenetche

or: Article 143 → about → Entity SEO

or: Article 143 → publisher → Euskal Conseil

or: Article 143 → isPartOf → WebSite

That relationship belongs at document level.

This was why my builder made the problem more painful.

I could not simply solve everything once globally.

I needed persistent nodes that could be referenced consistently throughout the website.

This is where @id and @graph became much more important to me

The breakthrough was understanding that I did not necessarily need to redefine a new person or organization independently on every page.

I needed stable identifiers.

If my organization had a persistent @id, for example:

https://example.com/#organization

then the same node could be referenced from multiple pages.

Likewise:

https://example.com/#person

https://example.com/#website

Then each article could connect itself to those already identified nodes.

Conceptually, this reduced a major source of ambiguity.

Without persistent IDs, a machine might encounter hundreds of separate objects all called “Lydie Goyenetche”.

With stable IDs, I was at least explicitly declaring: these references are intended to describe the same entity.

Again, declaration is not proof.

Google remains free to validate, ignore or reinterpret structured data.

But I was reducing the amount of identity reconstruction I was asking the machine to perform from scratch.

Wikidata and DBpedia became external semantic anchors

The next step was to extend the logic beyond my own website.

For important concepts, I began experimenting with Wikidata and DBpedia references.

Instead of simply writing:

  • "Knowledge Graph"

  • inside structured data, I could point toward a persistent URI representing the concept.

That distinction mattered to me because strings are ambiguous.

An identifier is much less ambiguous.

My logic gradually became:

Article → about → Concept URI

rather than simply:

Article → contains words related to concept

I began calling this, somewhat informally, AI-ready content.

I do not mean that adding Wikidata or DBpedia URLs magically optimizes a page for an LLM.

There is no evidence for such a simplistic claim.

The technical value is narrower and more defensible:

persistent external identifiers can provide additional disambiguation clues about what entity or concept a structured statement intends to reference.

This matters in systems where identity resolution precedes interpretation.

Google patents show that semantic matching itself can use embeddings

At this point, the entity graph and my earlier intuition about vectors started to connect.

Google patent “Semantic matching and retrieval of standardized entities” — US20210303638A1, granted as US11481448B2, has a priority date of March 31, 2020.

The patent explicitly describes using embeddings, clusters and hierarchical structures to semantically match input information against standardized entities.

This is especially important because it bridges two concepts that SEO discussions often treat separately: entity resolution and vector similarity.

An entity can have an explicit identifier.

But a system may still need to determine whether an incoming string, document or description corresponds to that standardized entity.

Embedding-based semantic matching provides one possible technical mechanism for doing so.

This does not prove that Google Search uses this exact patent implementation for my website.

Patents describe possible systems, not necessarily current production architecture.

But it demonstrates that within Google's technical research, entities and vector representations are not separate worlds.

They can participate in the same reconciliation problem.

My intuition about “weights” was becoming more technically plausible

At this point, my original mental picture began to make more sense.

Not because I had discovered Google's secret ranking coefficients.

I had not.

But because different Google systems were clearly dealing with concepts such as:

  • entity identity,

  • semantic similarity,

  • query association,

  • clustering,

  • confidence,

  • distance,

  • and scoring.

For example:

“Entity normalization via name normalization” — US8700568B2 deals with resolving representations of entities.

“Associating an entity with a search query” — US9336211B1 deals with relationships between entities and search queries.

“Semantic matching and retrieval of standardized entities” — US11481448B2 explicitly involves embeddings and semantic matching of entities.

And Google's current Enterprise Knowledge Graph documentation describes confidence scores based partly on similarity and distance inside entity clusters.

Again, these are not four pieces of evidence proving one hidden Google Search equation.

They are evidence of something broader.

Google's information systems repeatedly encounter the same mathematical problems: identity, similarity, confidence, association and relevance.

That was exactly the family of problems I had started seeing behind my traffic.

Then I wondered: if entity reconciliation matters so much, surely Google must offer a tool for it

At some point, I asked myself a very simple question.

If Google spends so much effort resolving entities, surely it must offer some kind of technology for doing this.

It does.

I discovered Google's Enterprise Knowledge Graph Entity Reconciliation system.

My first reaction to the interface was that it looked slightly like a NASA control panel.

But conceptually, it was fascinating.

Google describes reconciliation as a clustering process in which different records can be grouped because they appear to represent the same underlying entity.

Each resulting cluster can receive a machine identifier, and Google's documentation describes a confidence score measuring how strongly a record belongs to that cluster.

That was extremely close to the kind of mechanism I had been trying to imagine from outside.

But Google's Enterprise reconciliation service and my website experiments were not the same thing

This distinction is crucial.

Using Enterprise Knowledge Graph does not mean asking Google Search:

Please reconcile my company and improve my ranking.

There is no public documentation supporting that conclusion.

Google's Enterprise Knowledge Graph product is designed to reconcile datasets and build knowledge systems.

What I was doing on my website was different.

I was modifying the source itself.

Instead of submitting a dataset to a reconciliation API, I was trying to make each web document more explicit about its semantic relationships.

My pages could contain information such as:

  • who authored the document,

  • which organization published it,

  • what entity the document was about,

  • which concepts it mentioned,

  • how the person and organization were related,

  • and which persistent external identifiers could help disambiguate some of those concepts.

The objective was not to force Google to accept my graph.

It was to give the machine a cleaner graph to evaluate.

Under the hood, I was finally beginning to separate four different problems

This work helped me distinguish processes that I had previously treated as though they were the same thing.

Description

What does the page claim?

Identification

Which real-world object does a name or identifier refer to?

Reconciliation

Do several records, pages or profiles describe the same underlying entity?

Ranking

For a particular query, which documents or entities should appear, and in what order?

These are radically different computational problems.

A JSON-LD statement can help describe something.

An @id can help maintain identity within my own graph.

A Wikidata URI can help clarify the intended external concept.

An entity-resolution system may then determine whether multiple representations correspond to the same entity.

A ranking system still has to decide whether that entity or document is relevant to a query.

This changed my SEO model completely.

I was no longer simply trying to create a website that Google could crawl.

I was trying to create a website in which the most important relationships required less implicit reconstruction.

And strangely enough, while I was learning to make those relationships more explicit, the machines kept coming back.

My Entity Graph Was Becoming Clearer, but My CMS Was Still Flattening the Site

I had rebuilt the taxonomy in code — but not in the actual URL structure

As my understanding of entity relationships improved, I became increasingly aware of another limitation.

My website builder imposed a fundamentally flat architecture.

In the structure I would ideally have wanted, my pages might have looked like this:

/seo/

/seo/entity-seo/

/seo/knowledge-graph/

/seo/technical-seo/

That kind of structure expresses hierarchy directly in the URLs.

The pillar page is clearly above the child pages.

The child pages are clearly grouped inside the same thematic section.

But my builder did not really allow me to construct the site that way.

Instead, the structure looked much closer to:

/seo

/entity-seo

/knowledge-graph

/technical-seo

Technically, all those pages remained almost direct children of the root domain.

And that distinction mattered more than I had initially realized.

I tried to compensate for the CMS limitation with structured data

Because I could not fully express the hierarchy through the native architecture of the site, I began rebuilding it in code.

I used JSON-LD to make parent-child relationships more explicit.

I defined pillar pages.

I linked supporting articles to those pillars.

I created thematic relationships between different parts of the site.

Conceptually, I wanted to tell the machine:

SEO Pillar Page

↓ hasPart

Entity SEO Article

↓ hasPart

Knowledge Graph Article

And I tried to reproduce equivalent structures across French, Spanish and English.

From a semantic point of view, the graph became much clearer.

But from a document-topology point of view, the CMS remained flat.

That was the important limitation.

JSON-LD could describe the hierarchy, but it could not create a real hierarchical document graph

This was another lesson I learned gradually.

Structured data can express relationships.

It can tell a machine that one page belongs to another thematic structure.

It can define:

  • isPartOf

  • hasPart

  • about

  • mentions

  • author

  • publisher

  • and other relationships.

But it does not physically transform: domain.com/entity-seo into:

domain.com/seo/entity-seo/

The actual URL remains where the CMS places it.

The site's crawlable architecture remains the architecture that is really delivered.

That means Google can receive two forms of information at the same time.

One layer says:

this article belongs to this SEO pillar.

Another layer says:

this article sits directly beneath the root domain, just like dozens of other pages.

Those two signals are not necessarily contradictory.

But they are not identical either.

The homepage naturally became the strongest structural node

This helped me understand another behavior I had been observing.

My homepage was capturing a surprisingly large proportion of my core professional queries.

SEO queries.

GEO queries.

Local professional queries.

And, unexpectedly, queries coming from several languages.

At first, I wondered whether Google had simply failed to understand my pillar pages.

But another explanation seemed increasingly plausible.

The homepage was structurally dominant.

It sat directly at:

domain.com

Almost every major section of the website connected back to it.

It represented the organization globally.

It connected my services.

It connected my personal identity.

It connected my multilingual content.

It had the strongest historical presence.

And because so many other pages were also positioned directly beneath the same root, my pillar pages did not benefit from the same explicit hierarchical reinforcement they might have received inside a structure such as:

/seo/entity-seo/

From a graph perspective, the homepage had become an extremely dense node.

Google may have understood the topics while still preferring the wrong URL

This distinction became very important to me.

At first, I interpreted the homepage ranking for so many professional queries as evidence that Google still did not understand the architecture.

But perhaps that was not the right conclusion.

Google could understand that:

this article belongs to entity SEO,

this page is the SEO pillar,

this person is associated with Euskal Conseil,

this organization offers SEO and GEO services,

and these pages belong to the same multilingual professional ecosystem.

Yet when the ranking system had to select one URL for a broad query, the homepage could still remain the strongest candidate.

That meant I had to distinguish between two problems: semantic understanding and URL selection.

They are not the same thing.

A machine can understand what a document is about and still decide that another document is a stronger result for the query.

My entity SEO was becoming stronger than my URL-level SEO

This was probably one of the most useful conclusions from the experiment.

The reconciliation work seemed to help Google understand the broader identity.

The person.

The organization.

The services.

The topics.

The relationships.

The multilingual consistency.

But the distribution of queries across individual URLs remained imperfect.

In other words:

the entity appeared to be becoming clearer faster than the document hierarchy.

This was not necessarily a failure of the structured data.

It was a reminder that structured data and information architecture solve different problems.

JSON-LD can clarify meaning.

It can reduce ambiguity.

It can connect entities.

It can express relationships that are difficult to infer from raw text alone.

But it cannot entirely compensate for a CMS that exposes most documents at approximately the same structural level beneath the root domain.

The root domain was still telling Google a very strong story

I had therefore reached a strange situation.

Semantically, I was telling Google:

Article → Pillar Page → Topic

But structurally, the site still looked closer to:

Root → Article

Root → Pillar Page

Root → Service Page

Root → Another Article

Root → Another Pillar Page

The root remained common to almost everything.

And the more I connected the person, organization, services and thematic nodes together, the more the homepage itself could become the central meeting point for those signals.

This may explain why it continued to capture broad professional queries across several languages.

The very work that improved entity coherence may also have reinforced the prominence of the page sitting at the center of the whole system.

Again, this is an interpretation of the behavior I observed, not a disclosed Google ranking rule.

But it gave me a much better model of what I was seeing.

A semantic hierarchy is not the same as a crawl hierarchy

This became one of the clearest technical lessons of the entire project.

I could describe a semantic hierarchy in JSON-LD.

I could create relationships between entities.

I could identify parent topics and supporting content.

I could make the graph more explicit.

But a semantic hierarchy does not automatically become a crawl hierarchy.

A crawler still sees:

  • URLs,

  • links,

  • navigation,

  • sitemaps,

  • canonicals,

  • HTML,

  • response codes,

  • and the real topology of the site.

Structured data adds another interpretation layer.

It does not erase the layers beneath it.

That distinction matters enormously for anyone trying to use entity SEO to compensate for CMS limitations.

My builder was not preventing Google from understanding the site — it was limiting how precisely I could distribute authority

This is probably how I would describe the problem today.

The builder was not making the site invisible.

It was not preventing indexing.

It was not preventing Google from understanding entities.

But it was limiting my ability to express the document hierarchy as precisely as I wanted.

And because the architecture remained flat, authority could concentrate more strongly around the root domain and the homepage.

My structured graph could clarify:

what belongs together.

But the technical architecture still influenced:

where authority accumulates.

That difference became impossible for me to ignore.

The multilingual behavior made the effect even more visible

The phenomenon became especially interesting because my pillar pages were progressively being tested across French, Spanish and English.

Google appeared to understand that equivalent thematic structures existed in several languages.

The tests showed a degree of symmetry.

That was encouraging.

It suggested that the entity and thematic relationships were being understood beyond a single linguistic environment.

But the homepage still captured many broad professional queries across those languages.

So once again, I was seeing two things happen at the same time:

multilingual semantic understanding was improving

while

URL-level distribution remained imperfect.

That paradox told me more about the limitations of my CMS than almost any technical audit could have done.

I had improved the graph without fully changing the topology

That is probably the simplest way to summarize what happened.

I had improved the graph.

I had clarified entities.

I had strengthened thematic relationships.

I had connected pillar pages and supporting content.

I had made the multilingual structure more explicit.

But I had not fully changed the topology imposed by the CMS.

And that distinction changed the way I understood structured data.

JSON-LD was not a replacement for architecture.

It was a semantic layer added on top of architecture.

Powerful, useful and sometimes extremely clarifying.

But still a layer.

The root domain remained the root.

The homepage remained the strongest hub.

And the site continued to remind me that a machine can understand a structure conceptually without reproducing exactly the hierarchy I intended it to rank.

Conclusion — I Have Not Finished the Experiment

I could probably write much more about all of this.

Far more, in fact.

But at this point, I am already exceeding what most people are realistically willing to concentrate on while reading an article on a screen.

And perhaps that is a good place to stop.

Because my work is not finished anyway.

What I can say today is that the technical reconciliation work I carried out has coincided with several changes that are difficult for me to ignore.

My Instagram account began capturing local professional queries much more clearly.

My impressions around SEO and GEO increased.

My traffic became more balanced across the different thematic areas of the website.

And Google appeared to understand my professional identity in a much more structured way than before.

But there was also a price to pay.

The site seemed to enter a new evaluation phase

What I observed looked very much like a partial reset.

Not a disappearance from Google.

Not a penalty.

More like a situation in which the search engine appeared to be reassessing the website almost as though it were discovering it again.

The difference was that this time, it did not seem to test isolated pages in the same way.

I began seeing what looked like groups of thematically connected pages being tested together.

What I would describe, in my own working vocabulary, as clusters of authority nodes.

I must be careful here.

I cannot prove that Google officially “reset” my site.

Nor can I prove that every ranking movement I observed corresponded to a specific RAG reevaluation cycle.

Those are interpretations based on repeated observations.

But the pattern was sufficiently consistent to interest me.

When one thematic cluster moved, other related pages often seemed to move with it.

And each period of testing appeared to delay the moment when traffic became denser and more stable.

I would like to call it mass traffic — but that would be dishonest

I sometimes catch myself wanting to say that I am waiting for “mass traffic”.

That would sound impressive.

It would also be slightly ridiculous.

The profession I chose is already a niche.

SEO is specialized.

GEO is even more specialized.

Entity SEO is more specialized again.

And the particular intersection I am exploring — entity reconciliation, structured graphs, retrieval systems, semantic representations and generative search — is narrower still.

My articles are not written for millions of people.

They are written for the relatively small number of people who are curious enough to go this far into the mechanics of Search.

So the objective is not necessarily massive traffic.

It is relevant traffic from the right semantic neighborhood.

That distinction matters much more to me now than raw visitor volume.

The Knowledge Panel was probably the strangest validation

One of the most surprising outcomes was seeing a Knowledge Panel associated with my identity.

Even more interestingly, the panel included an AI-generated summary.

I had never written that summary.

I had never used those exact words to describe myself.

The text was Google's own synthesis of the information it had gathered.

For me, that was fascinating.

Because it provided something I had been trying to understand throughout this entire experiment:

a visible representation of how the machine had reconciled my identity.

It was no longer only:

what do I say that I am?

It became:

what does Google currently think I am?

Those are two very different questions.

And for someone studying entity SEO empirically, the second one is considerably more interesting.

My company name produces yet another representation

Something else happens when I search for my company rather than my personal name.

Google presents the information differently.

That again suggests something important.

A person entity and an organization entity may be connected, but they are not identical objects.

The relationships surrounding them are different.

The attributes are different.

The queries associated with them can be different.

The pages Google chooses to surface can therefore also be different.

This may sound obvious from a Knowledge Graph perspective.

But seeing it emerge progressively from my own website made the concept much more concrete.

I was no longer dealing with a single keyword universe.

I was watching different entity representations appear depending on the object being queried.

The multilingual tests have been even more revealing

Another particularly interesting observation concerns my pillar pages.

They have progressively been tested across my different languages.

And what surprised me was the relative symmetry of those tests.

French.

Spanish.

English.

Equivalent thematic structures appeared to receive visibility in each linguistic environment.

That does not mean every page has equivalent rankings.

Nor does it prove that Google has built some perfectly unified multilingual representation of my entire site.

But it strongly suggests that the thematic relationships I have been constructing are no longer being interpreted solely inside one language.

For me, this is one of the most encouraging results.

Because one of the hardest questions behind a multilingual website is not:

Can Google translate this?

It is:

Can Google understand that these different linguistic expressions belong to the same underlying professional and thematic identity?

I am beginning to see evidence that it can.

The website still has a lot to teach me

There is still an enormous amount of work left.

Some nodes need strengthening.

Some relationships probably remain ambiguous.

Some articles still need structured reconciliation.

Some multilingual connections need clarification.

Some experiments will undoubtedly prove useless.

Others may contradict conclusions I currently believe to be correct.

And that is precisely why I find the project so interesting.

When I began learning SEO, I wanted to know which techniques worked.

Today, the question that interests me much more is:

what is the machine trying to understand?

Once that question changed, SEO stopped being a checklist.

It became an investigation.

I am still conducting it.



EUSKAL CONSEIL

9 rue Iguzki alde

64310 ST PEE SUR NIVELLE

07 82 50 57 66

euskalconseil@gmail.com

Mentions légales: Métiers du Conseil Hiscox HSXIN320063010

CGV & Mentions légales

Ce site utilise uniquement Plausible Analytics, un outil de mesure d’audience respectueux de la vie privée. Aucune donnée personnelle n’est collectée, aucun cookie n’est utilisé.

communication digitale
communication digitale