[ Case study ]
Faceted navigation had generated hundreds of thousands of near-empty URLs — filtered lists with zero listings. Crawl budget drowned in them, quality pages aged out of the index, and Search Console coverage was a wall of 'Crawled – currently not indexed'.
CLIENT a regional services directory — FOCUS Classify every URL pattern
Representative examplesEvery case study in this library is an illustrative composite of the kind of engagement we deliver — written to show our method and standards, not to name clients.
Regional services directories live and die by their listings being findable, and this one had spent years adding facets users genuinely value — filter by suburb, service, opening hours — without anyone counting the URLs those filters created. Empty filtered views outnumbered real listings by tens to one. The platform is a bespoke build with a templating layer that can express rules per template and not much beyond that; a small team runs the product, and its developer already splits time between feature work and maintenance.
Faceted navigation had generated hundreds of thousands of near-empty URLs — filtered lists with zero listings. Crawl budget drowned in them, quality pages aged out of the index, and Search Console coverage was a wall of 'Crawled – currently not indexed'.
We proposed classifying before cutting: a full crawl would separate URL patterns by whether the facet combination ever returns listings, and each pattern would get one rule — index, canonical, or noindex — applied template by template with crawl statistics watched between waves so a mistake stayed contained. Genuinely empty archives would be retired with 410s rather than left in noindex limbo. Users keep their filters; only the crawler-facing side changes. History mattered too: previously indexed URLs got individual treatment where equity existed.
Just as important is what we ruled out, and why:
The crawl separated listing patterns by value — which facet combinations ever returned results — and assigned each a rule: index, canonical, or noindex.
Rules rolled out per pattern with crawl stats watched between waves, so a mistake was contained to one template.
Genuinely empty archive pages were retired with 410s — honest removals instead of limbo noindexes.
Delivered by the growth pod — technical SEO specialist over 7 weeks, with working increments reviewed with the client every week.
Obstacle
Classification hit seasonal facets: combinations that return results every summer were judged empty during the winter crawl, because 'ever returns results' was measured over one snapshot.
Handled: We reclassified against a rolling twelve-month sample of internal facet usage and listing counts by month, giving seasonal patterns index rules with standing review dates.
Obstacle
The templating engine cached aggressively, and two of the fourteen templates kept serving pre-rule markup for days after their noindex shipped.
Handled: The between-waves crawl check caught both; the developer added a cache purge to the rollout step, and the wave re-ran with a clean verification.
Obstacle
Retiring empty archive pages with 410s risked discarding a small set of junk-pattern URLs that had collected real links from an old press feature.
Handled: We cross-checked retired URLs against the backlink index; the handful with links redirected to their parent listing category, and the rest were removed without ceremony.
The headline: indexed urls, with 'crawled – not indexed' reports falling accordingly, over the 10 weeks after full rollout — 340k → 12k, read from Search Console coverage. A second check: impressions on quality listing pages, same period at +18%.
New listings surface in days now, and the editorial team stopped resubmitting pages that had never actually disappeared from anything but Google's patience. Crawl reports fit on one screen, which changed their meetings: the product owner asks what URL pattern a new feature will generate before it ships. The developer keeps the rules file open in a tab — the discipline survived the engagement, which is the part we care about most.
The result was read from Search Console coverage against the pre-engagement baseline over the stated window, with a guardrail check on impressions on quality listing pages, same period. Where platform-reported numbers and business outcomes differ, this record says which layer it is quoting.
What we would do differently
We would have started the log analysis on day one — the crawl-budget picture was clearer in the logs than in any crawler export.
[ Related service ]
[ Related builds ]
4 27Weekly prompt-set answers citing or summarizing the brand's content, weeks 1 to 10
14 31Qualified inquiries per month attributed to service pages, over the 8 weeks after rollout versus the 8 before
[ Next step ]
Next case study