Topical authority: how pillar-and-cluster content architecture actually works
Pillar-and-cluster architecture is the most useful content planning model in SEO and the most frequently misapplied. This is how to choose clusters, map queries to URLs, wire the internal links, and tell whether any of it is working.
Authority is earned at topic level, not page level
The observation that started all of this is genuine and easy to verify for yourself. Sites that cover a subject properly — thoroughly, from several angles, over time — begin to rank for questions inside that subject they never explicitly targeted. Sites that publish one page per keyword across forty unrelated subjects do not, even when the individual pages are competently written.
It is worth being precise about what that observation does and does not prove, because the industry has built a great deal on top of it. There is no confirmed site-wide “topical authority score” at Google. The only system Google has publicly described under that name is framed around news, where it considers things like how notable a source is for a topic or location and how a publisher’s original reporting gets cited by others. Anyone quoting you a topical authority percentage is quoting a vendor’s model rather than a search engine.
What the model does explain is more mundane and more actionable. Comprehensive coverage produces more pages matching more queries. Related pages linked to each other give the crawler context about what a page covers and which page is the important one. And depth is the practical precondition for saying anything non-obvious, which is what Google’s guidance keeps returning to when it distinguishes a unique expert take from commodity content that could have come from anyone.
So treat topical authority as a planning heuristic rather than a metric. It is a good answer to “what should we publish next”, and a bad answer to “what should we report”.
What a pillar page is, and what it is not
A pillar page is the canonical page for the broadest question in a topic. Its job is to orient someone who has not yet decided what they want, then route them to the page that answers their actual question.
This is where the model is most often misapplied. The common interpretation — that the pillar should be an exhaustive ultimate guide covering everything the cluster covers — produces a page that competes with its own children, that nobody reads to the end, and that is impossible to keep current. If the pillar answers everything, the cluster pages have nothing to add, and you have built one long page with a navigation problem.
There is a genuine disagreement here. One camp argues the pillar must be exhaustive because length and comprehensiveness are what rank for head terms; the other argues the pillar is a hub whose value is structural. Our position, and you may reasonably disagree: the pillar should be complete at its own level of abstraction and no deeper. A reader who only reads the pillar should leave genuinely informed, and anyone wanting the next level of detail should be sent somewhere specific.
A pillar that works usually contains:
- A direct answer to the head question in the first two paragraphs. Not throat-clearing. If the page is “what is X”, define X immediately and precisely.
- The decision the reader is actually trying to make, named explicitly, because the head query is almost always a decision in disguise.
- A section per sub-topic, each summarising the sub-topic honestly and linking to the page that treats it properly, with an anchor that describes what the reader will get.
- At least one thing that only you can provide — a comparison table, a decision rule, a piece of first-hand experience. Otherwise the pillar is a table of contents with adjectives.
- A structure that survives additions. You will add cluster pages later, and the pillar has to accommodate them without a rewrite.
Length is an output of that, not an input to it. A pillar that needs four thousand words should have four thousand; one that needs fifteen hundred is finished at fifteen hundred.
Why three to seven deep clusters beat twenty shallow ones
Most content plans fail on arithmetic rather than strategy. Twenty clusters means twenty pillars plus a hundred-odd cluster pages, all needing research, editing, updating and link maintenance. Nobody has that capacity, so the work degrades: pages get shorter, research gets thinner, and eventually the plan is met by producing pages that restate what is already on the first page of results.
That end state has a name in Google’s spam policies. Scaled content abuse describes many pages generated primarily to manipulate rankings rather than to help users, and explicitly includes using generative tools to produce many pages without adding value. A twenty-cluster plan run by a two-person team converges on that pattern whether or not anyone intended it.
The selection rule we would use, in order:
- Can you say something non-obvious here? If your only contribution to a topic would be a competent summary, that topic is not yours to own. Pick the ones where your experience produces information the existing results do not have.
- Does it connect to something you sell? Not every page needs commercial intent, but every cluster should have a plausible route to a commercial page. A cluster with no such route is a hobby.
- Is the topic big enough to contain real sub-questions? A cluster needs genuine internal structure. If you cannot list eight distinct decisions inside it, it is a page, not a cluster.
- Can you maintain it? Anything factual decays. A cluster about a platform that ships quarterly is a standing commitment, and taking on five of those is how sites end up full of confidently wrong pages.
Three excellent clusters beat twenty thin ones for the same reason three excellent screenshots beat ten mediocre ones: the audience judges you by the depth of the best thing they find, not by the count.
Query mapping: one intent per URL
Before writing anything, build the map. It is a spreadsheet, it takes a day, and it prevents most of the problems in this article.
| Column | What goes in it | Why it matters |
|---|---|---|
| Query | The actual phrasing people use | Anchors the page to real language, not internal vocabulary |
| Intent | Informational, comparative, transactional, navigational | Two intents on one URL means one of them is served badly |
| Stage | Problem-aware, solution-aware, vendor-aware | Determines tone, depth and what the page asks for next |
| URL | The single page that owns this query | The point of the exercise |
| Role | Pillar or cluster | Sets the linking direction |
| Must answer | The one question the page has to resolve | Becomes the brief |
| Information gain | What this page has that page one does not | If blank, do not write the page |
The last column is the one people skip and the one that decides whether the cluster is worth building.
How cannibalisation actually starts
The mechanism is less dramatic than the folklore. Google chooses one URL to represent a query, and when several of your pages serve the same intent it may choose differently over time. The consequence is not a penalty — it is dilution. Your external links, internal links and accumulated click history split across pages that would have been stronger as one, and your reported position oscillates in a way that makes the content look unstable to whoever reads the report.
Two practical notes. First, sharing a keyword is not the problem; sharing an intent is. A comparison page and an implementation guide can both discuss the same platform without competing, because they answer different questions. Second, when you do need to consolidate, the strength of your options is documented: a permanent redirect is the strongest signal, rel="canonical" is a strong signal, and listing a URL in a sitemap is a weak one. Choose accordingly, and merge the content rather than simply redirecting the weaker page into oblivion — the redirect moves signals, but the reader still needs the material that was on the page you removed.
The internal link contract between pillar and cluster
Internal links are the part of this model that does the actual work, and the part most often left to whoever happens to remember.
Google’s guidance is short and worth following literally. A link is only crawlable if it is an <a> element with an href — not an onclick handler, not a framework router directive with no href. Anchor text should be descriptive and specific, and Google names “click here” and “read more” as the pattern to avoid. Every page you care about should have a link from at least one other page on your site. There is no ideal number of links, and stuffing them is a spam pattern.
Nielsen Norman Group’s work on information scent explains why the anchor text advice matters beyond crawling: readers decide whether to click by estimating, from the label and its surroundings, how likely the destination is to answer their question. A vague label produces a weak estimate, and weak estimates produce no click. The sentence that helps a search engine understand a destination is the same one that helps a human decide to go there.
The contract we would hold a cluster to:
- The pillar links to every cluster page, in context, with an anchor that describes the destination’s specific value.
- Every cluster page links back to the pillar, once, early, where the reference is genuine.
- Cluster pages link laterally only where the adjacency is real. Manufactured cross-links between unrelated siblings help nobody and read as filler.
- Anchor text varies. The same destination linked from four pages should be described four ways, because those pages are making four different points about it.
- No page in the cluster is orphaned. If a page has no inbound internal link, it is not part of the cluster no matter what the plan says.
- URLs are grouped in directories. Google notes that grouping similar topics in directories helps it learn how often URLs in each directory change, and it makes the structure legible to humans reading the address bar.
Do this once at the end of each cluster as a deliberate pass, not opportunistically while writing. Retrofitting links across a cluster takes an afternoon; noticing three years later that half your pages are orphaned takes a full audit.
Cadence and the compounding curve
Publishing cadence matters less than people think and consistency matters more, for an unglamorous reason: a cluster produces very little until it is nearly complete, and half-finished clusters are the most common way content programmes fail. Six clusters at forty per cent completion perform worse than two clusters finished.
So the rule is sequence, not speed: finish one cluster before starting the next, and prefer a slower rate you can actually sustain to a burst followed by nine quiet months.
Expectations should be set against what Google actually says rather than against a vendor’s chart. Improvements can take effect in a few days or over several months, and there is no guarantee that changes produce a noticeable impact at all. Google also advises waiting at least a full week after a core update finishes rolling out before analysing your site — which is a useful discipline generally, because the alternative is attributing normal variance to whatever you shipped most recently.
Finding the coverage gaps you have not answered
The best gap-finding methods cost nothing and use data you already own.
Search Console, read sideways. Sort your queries by impressions with a low average position. Those are questions Google already associates with your site but does not think you answer well. That list is a content plan written by the search engine. Note the caveat when reading position: Search Console reports the topmost position your property occupied, averaged across impressions, so a page competing with a sibling shows you the better of the two rather than the situation.
Query fan-out as a planning exercise. Google describes its AI features decomposing a question into sub-questions and synthesising across the results. Do that manually: take each head question in your cluster, write down the six sub-questions a thorough answer would require, and check whether you answer each one somewhere. The unanswered ones are your next briefs.
The questions your own people are asked. Sales calls, support tickets, onboarding sessions. These are the highest-quality source of real queries in the business and almost nobody mines them systematically. A monthly half-hour with whoever answers the phone will out-perform most keyword tools.
Where the topic is discussed without you. Forums, communities, review sites, conference Q&A. You are looking for the phrasing people use when nobody is trying to rank for anything.
Whichever method you use, apply the same filter at the end: what will this page contain that the current first page of results does not?
Refresh, rewrite or consolidate: a decision rule
Every established site reaches the point where editing beats publishing. The decision is easier with a rule than an argument.
| Situation | Action | Why |
|---|---|---|
| Ranks well, facts have aged | Refresh in place | Keep the URL and its accumulated signals; update the specifics and the date |
| Ranks moderately, thin against current results | Rewrite at the same URL | The intent is right, the execution is not; there is nothing to gain from a new URL |
| Two or more pages serving one intent | Consolidate into the strongest URL, redirect the rest | Stops the dilution; merge the material rather than discarding it |
| No impressions, no links, no purpose | Remove and redirect to the nearest relevant page | Reduces maintenance surface; do not expect a ranking gain from the deletion itself |
| Good page, wrong place in the structure | Leave the URL, fix the internal links | Most “underperforming page” problems are linking problems |
One honest caveat about pruning, because it is oversold. Removing weak pages reduces the amount of material you have to maintain and can improve how coherent a section looks to a reader. Treating deletion as a growth tactic in itself is a different claim, and it is not one anyone can support from Google’s documentation. Delete pages because they are not worth keeping, and expect a tidier site rather than a traffic increase.
Measuring topical authority without a vanity metric
“Number of keywords ranking” is the metric of choice here and it is close to useless, because it counts position-eighty rankings for terms nobody searches alongside the ones that matter.
Four measures that are harder to game and easier to act on:
- Coverage of the cluster’s query set. Define the set of queries the cluster is supposed to serve, once, and track the share where you appear at all. This rises as the cluster completes and it is the closest honest proxy for the thing the model claims to build.
- Average position within that fixed set. Fixed is the operative word — a query set that changes each month measures nothing.
- Distinct queries per URL. A page earning impressions across many related phrasings is a page the engine understands well. A page earning impressions on one exact phrase is a page that got lucky.
- Entry pages for conversions, not last click. Cluster pages rarely convert directly; they get read, remembered, and returned to. If your reporting only credits the last page before the enquiry, the entire programme will look like it does nothing. Fixing this is an analytics problem before it is a content one.
The leading indicator worth watching is the appearance of queries you never targeted. When a cluster starts earning impressions on questions that were not in the plan, the coverage has become general rather than page-specific, which is the actual phenomenon underneath the phrase “topical authority”.
A build plan for a single cluster
Deliberately ordered so the cheap, high-yield work happens before the expensive work.
- Choose the cluster against the four selection criteria above. Write down, in one sentence, why this topic is yours.
- Build the query map. Every row gets an intent, a URL and an information-gain note. Rows with a blank information-gain column get deleted, not written.
- Audit what you already have. Most sites have partial coverage. Decide per existing page: refresh, rewrite, consolidate or leave.
- Write the pillar first, so the structure exists before the children. Publish it even though the cluster pages it will link to do not yet exist.
- Publish cluster pages in order of commercial proximity. The ones nearest a buying decision first — they take the same effort and return sooner.
- Run the linking pass once the cluster is substantially complete: pillar to every child, every child back, lateral links only where real, varied anchors, nothing orphaned.
- Set the baseline. Record the fixed query set and current coverage before you expect any movement, because retrofitting a baseline is impossible.
- Leave it alone long enough to mean something, then review coverage, distinct queries per URL and the unplanned queries appearing. Then start the next cluster.
How this site is structured, as a worked example
It is fair to ask whether we do this ourselves, so here is the architecture in plain terms.
Services sit in a hub with genuine parent-child nesting where the child is a real specialism rather than a keyword variant — search engine optimisation has technical, local, ecommerce and generative-engine children, and mobile app development has platform children. Each service page is the commercial destination for a topic, and it carries its own sources, because a page making claims should show its working.
Articles live in a separate section, and each one declares which service it belongs to, which sibling articles it relates to, and which sources support it. That declaration is not decorative: the template builds internal links from it, which is a structural guarantee that no article is a dead end and every cluster routes somewhere commercial. Every factual claim has to be attached to a source that resolves, because the alternative is a site full of numbers nobody can check.
The body links in this article follow the contract described above — into content marketing where the adjacency is production capacity, into technical SEO where the problem is crawlability and canonicals, and into the wider search practice where the question is strategy. The adjacent reading on getting cited by AI search is here because query fan-out changed how coverage pays off, and protecting rankings through a redesign because the fastest way to destroy a cluster is to migrate it carelessly.
None of that requires an agency to reproduce. It requires a decision about which three topics you intend to own, the discipline to finish one before starting the next, and someone willing to delete a planned article when it turns out to have nothing to add.