# How to stop Googlebot wasting its crawl on PrestaShop filters

> On catalogues with untreated facets, the share of crawling devoted to filters often exceeds half the requests. How to measure it in the server logs, why noindex and robots are not enough, and the limits to respect.

- Page: <https://www.datafirefly.com/en/2026/09/25/googlebot-crawl-waste-prestashop-filters/>
- Language: en
- Published: 2026-09-25
- Last updated: 2026-09-25
- Other languages: [fr](https://www.datafirefly.com/2026/09/25/crawl-filtres-prestashop-budget-exploration/index.md), [es](https://www.datafirefly.com/es/2026/09/25/googlebot-rastreo-filtros-prestashop/index.md), [de](https://www.datafirefly.com/de/2026/09/25/googlebot-crawl-budget-prestashop-filter/index.md), [it](https://www.datafirefly.com/it/2026/09/25/googlebot-crawl-filtri-prestashop/index.md), [pl](https://www.datafirefly.com/pl/2026/09/25/googlebot-crawl-filtry-prestashop/index.md), [nl](https://www.datafirefly.com/nl/2026/09/25/crawl-filters-prestashop-crawlbudget/index.md)
- Index: <https://www.datafirefly.com/en/2026/llms.txt>

On a large catalogue, a considerable share of bot requests targets filter combinations that were never meant to be indexed. That time is taken from the time that would have served to discover your new products.

This article covers measurement and implementation. The criteria for choosing which facets to open have been detailed separately.

## Measuring the waste before acting

No decision should be taken without having looked at the server logs. They are the only source that shows what bots actually request, as opposed to what you think they request.

The method fits in three steps.

**Extract the bot requests** over four weeks, filtering on the user agent and verifying authenticity by reverse DNS lookup of the address. A share of visits claiming to be bots is not.

**Classify the requested addresses** into four families: product pages, categories without filters, categories with filters, and the rest. This classification is done by URL pattern and takes a few minutes once the file is loaded into a spreadsheet or a database.

**Calculate the distribution.** On catalogues with untreated facets, the share devoted to filters frequently exceeds half the requests, and it can reach 80% on the most open configurations.

Two complementary figures are worth noting. The **visit frequency on your product pages**: if a product is only visited once a month, your crawl budget is saturated elsewhere. And the **discovery delay** of your latest new products, between going live and the first visit.

## What the logs reveal

Three findings come back systematically on catalogues that have never been treated.

**Deep combinations.** Addresses carrying three, four or five simultaneous filters, which have no chance of being indexed and yet represent a significant volume.

**Sort parameters.** The same product list requested in six different orders. None of these states brings anything to the index.

**Facets with no results.** Combinations that return no product, explored all the same because the link exists on the page.

This last point is the most revealing: your interface offers links to empty selections, which is already a problem for your visitors before being one for crawling.

## Why the classic directives are not enough

The reflex is to put a noindex directive on these pages. It settles indexing, not crawling: the page must be visited for the directive to be read.

Blocking in the robots file prevents crawling, but it has two drawbacks. Pages already indexed can stay there, without their content being re-read to understand they should be removed. And the block is hard to work around cleanly: an address that is blocked but massively linked from your pages remains known.

These two tools remain useful. They simply intervene after the bot has discovered the address.

## Not giving the link

The complementary approach consists of not expressing these links as links.

The principle: the navigation behaviour is produced by a script attached to an element that is not a link in the protocol sense. The visitor clicks and navigates normally, the bot finds no address to follow.

Three conditions for this to be clean.

**The element must remain accessible.** A clickable element that is not a link must carry the attributes that make it usable with a keyboard and understandable by a screen reader. Without that, you fix a crawl problem by creating an accessibility problem.

**Navigation must remain functional.** The address must change, going back must work, the link must remain shareable. Obfuscation concerns the way the link is expressed in the code, not the disappearance of the address.

**The scope must be limited.** Apply this treatment only to the links you do not want crawled. Your main navigation, your categories and your open facets must remain normal links.

## What not to do with it

Two excesses to rule out clearly.

**Hiding links to content you want indexed**, in the hope of controlling value distribution. It is an old practice, ineffective, and it deprives your pages of incoming links.

**Presenting the bot with a page different from the visitor's.** Link obfuscation does not change the content served, it changes the way an interface element is written. The difference is sharp and must be held to: serving distinct content depending on the agent is a sanctioned practice.

## The other levers

Obfuscation is not the only tool, and it is not always the first.

**Reduce the number of filters offered.** Many catalogues display twelve facets of which four are used. Removing the useless ones mechanically reduces the combinatorics and simplifies the interface.

**Hide filters with no results.** A value that returns nothing should not be offered. That removes useless links and improves the experience.

**Load the filters on demand.** On mobile, the filter panel is often collapsed: if its content is only present after opening, the links are not in the served page.

**Take care of the sitemap.** It must contain your useful pages and nothing else. A sitemap that lists filter combinations explicitly invites their crawling.

## Checking after implementation

Four indicators, measured before, then at one month and at three months.

**The distribution of bot requests** by URL family, using the same method as at the start. It is the direct measure of the effect.

**The total number of pages crawled per day**, in the crawl statistics. Careful with the interpretation: a decrease is expected and desirable here, contrary to the usual belief.

**The visit frequency of product pages**, which should increase since the freed time redeploys.

**The discovery delay of new products**, which is the final indicator: it is the reason the whole project is run.

A point of method: allow three to four weeks before concluding. The redistribution of crawl budget is not visible immediately.

## What can go wrong

Three situations to watch after implementation.

**Useful pages become unreachable.** If a facet you wanted open ends up in the obfuscated scope, it loses its incoming links. Check that your open pages remain linked from a normal place.

**Navigation broken for the keyboard.** The test takes thirty seconds: walk through the filters with the tab key and check that they are reachable and activatable.

**A traffic drop on facet pages.** If you had some receiving traffic and they were caught in the scope, the drop appears in three to six weeks. Hence the value of recording positions beforehand.

The  implements this device on PrestaShop 8 and 9: selective treatment of filter and sort links with exclusion of open facets, preservation of the navigation behaviour and accessibility attributes, and a scope configurable per category.
