PS PrestaShop Intermediate

AI Competitor — Competitor Price Monitoring

Install and configure competitor price monitoring on PrestaShop 8 and 9: competitor profiles, catalogue crawl, AI extraction, alerts and analytics.

Updated Module version 1.1.1

Overview

AI Competitor tracks your competitors’ prices from inside your PrestaShop back office. You declare a competitor once, with its name and its selectors; the module can then pull the full list of its products from its sitemap, match that list against your catalogue, read prices at regular intervals, detect significant moves, alert you by email and propose an adjusted price based on your strategy.

Compatible with PrestaShop 8.0 to 9.x, PHP 7.4 to 8.3, native multistore. No Composer dependency, no external JavaScript library.

Installation

  1. In your back office, open Modules → Module Manager → Upload a module.
  2. Upload the dfaicompetitor.zip file.
  3. The module creates its tables and the Catalog → AI Competitor menu, which holds six pages: dashboard, competitors, monitored URLs, competitor catalogue, reports and analytics, import/export.

A unique cron token is generated on install. You will find it on the module dashboard.

Upgrading from a 1.0.x version

Just upload the new ZIP over the existing installation. The upgrade script creates the new tables, adds the missing columns, and above all generates a competitor profile for every competitor name already in use before attaching the matching URLs to it. Your existing data is preserved and there is nothing to re-enter.

Reinstalling a package carrying the same version number as the one already in place does not replay the upgrade script, and the PHP opcode cache may keep serving the old files. Always check the version number shown on the dashboard after an upload.

Initial configuration

Open Catalog → AI Competitor. The dashboard holds the statistics and five configuration blocks.

AI provider

The AI fallback is optional but recommended: it takes over when a competitor site publishes neither structured data nor a usable selector. Three providers are supported:

  • Mistral AI (default), the cheapest, suggested model: mistral-small-latest
  • Anthropic Claude, the most accurate on complex pages, suggested model: claude-haiku-4-5-20251001
  • OpenAI, a good middle ground, suggested model: gpt-4o-mini

Enter the API key of the chosen provider. It is masked on display; leave the field blank on a later save to keep the stored key. You pay the provider directly, with no DataFirefly markup.

Alerts

  • Notification email: recipient of the alert digests and the weekly report.
  • Change threshold as a percentage, 3 % by default. Below it, no alert is created.
  • Weekly report day, Monday by default.

Scraping

  • URLs per cron batch: how many URLs are read at each run, 20 by default.
  • Default interval in hours, applied to new URLs, 24 h by default.
  • HTTP timeout and User-Agent. The shipped User-Agent identifies the bot by name (DataFireflyBot); some sites allow a named bot they would otherwise block.
  • History retention in days, 180 by default.

Catalogue crawl

  • Pages per slice: how many competitor pages are opened in one pass, 15 by default.
  • Seconds per slice: the time budget of each pass, 20 by default. Keep it below your PHP max_execution_time.
  • Match threshold: minimum title similarity to accept an automatic match, 82 by default.
  • Monitor matches automatically: creates a monitored URL as soon as a crawled page is matched to one of your products. Off by default.
  • Advance crawls from cron: on by default.

Adjustment strategy

Three strategies determine the suggested price, always computed from the cheapest competitor:

  • Match: the same price as the lowest competitor.
  • Undercut by X %: X % below, 1 % by default.
  • Premium at X %: X % above, for a deliberate upmarket position.

Declaring a competitor

Since version 1.1.0 a competitor is an entity of its own, not a label retyped on every URL. Open Catalog → AI Competitor → Competitors → Add. The form has three tabs.

Identity

The name shown everywhere, the domain without a scheme (example.com), the currency, the reading interval, the forced AI extraction switch and a free notes field. The domain is what lets the module attach a new URL to the right competitor automatically.

Selectors

CSS selectors for the price, the product name, the SKU, the image and the availability. Leave them empty when the site publishes JSON-LD, which most modern shops do. Supported syntax: tag, .class, #id, [attribute=value], descendant and direct child. Example: div.product-price > span.amount

This is the main gain of the competitor record: the day a site rebuilds its theme and breaks a selector, you fix it here once and the fix propagates to every URL of that competitor that has no selector of its own.

Catalogue crawl

The sitemap URL (leave empty to read the Sitemap: directives of the domain’s robots.txt), the URL of a starting category page, the selector for product links and the one for the next-page link, the URL patterns to keep or to drop, the delay between requests, the product cap and robots.txt compliance.

Patterns accept a plain substring or a slash-delimited regular expression. On a typical shop, /product/ as a keep pattern is often enough to isolate product pages in a sitemap that also holds the blog and the CMS pages.

Crawling a competitor catalogue

Open Catalog → AI Competitor → Competitor catalogue, pick the competitor, the discovery source and the product cap, then start the crawl.

Two sources are available:

  • Sitemap, recommended. The module reads the configured sitemap or the one advertised by robots.txt, follows sitemap index files recursively and keeps the addresses matching your patterns.
  • Category pages, when no sitemap is usable. The module starts from the listing URL, collects product links through the competitor’s selector and follows the next-page link.

Discovery is followed by an enrichment pass that opens each page, extracts title, SKU, EAN, brand, price, availability and image, then attempts the match against your catalogue.

The crawl advances in time-bounded slices driven by the open page with a progress bar, and resumes at every cron pass. A sitemap with tens of thousands of URLs is processed without ever reaching PHP’s max_execution_time. You can close the page: the work continues on the cron side.

Matching against your catalogue

The module tries three routes, in decreasing order of reliability:

  1. EAN13, looked up on your products and their combinations. Certain match, score 100.
  2. Reference, supplier reference or MPN. Score 95.
  3. Title similarity, with a weighting that gives model numbers and brands more weight than common words, and a penalty when the two titles differ a lot in length. The score must exceed the configured threshold, 82 by default.

Every catalogue row displays the method and score of its match. A doubtful match is corrected by hand: the matching field offers autocomplete on your catalogue, by name or by reference. The Re-match buttons run the operation again on unmatched rows, or on all of them, which is useful after fixing references on your side.

A matched row becomes a monitored URL through its Monitor button, or in bulk through the group action on a selection. The created URL inherits the competitor’s name, selectors, currency and interval.

Adding a URL by hand

The crawl does not replace one-off entry. Open Catalog → AI Competitor → Monitored URLs → Add:

  • Competitor: the record the URL will inherit from. Leave empty to fill everything by hand.
  • Product: picked from your catalogue.
  • URL: the full address of the competitor product page.
  • Competitor name, CSS selector, currency, interval: filled from the competitor when you leave them blank.

If you pick no competitor, the module looks for the one whose domain matches the address you entered and attaches the URL to it. The Scrape button on each row triggers an immediate reading, handy to validate a URL as soon as you create it.

How extraction works

For each URL the module tries four methods in cascade and stops at the first that succeeds:

  1. JSON-LD: Product/Offer structured data, published by the vast majority of e-commerce sites. Price, currency, availability, and for the crawl the title, SKU, GTIN and brand, are read directly with no configuration.
  2. OpenGraph: the product price amount meta tags.
  3. CSS selector: the competitor’s one, or the URL’s own if it has one.
  4. AI: a cleaned HTML excerpt goes to your provider, which returns the price, the currency and the stock state. International price formats are handled (1 299,90 as well as 1.234,56 or $49.99).

Always start without a CSS selector: JSON-LD is enough in most cases. Only add a selector or forced AI if the Status column shows no_price.

Import and export

The Import / Export page is there to carry a configuration from one shop to another.

Competitors as JSON

The export produces a file holding the names, domains, selectors and crawl settings. Tick the competitors to include, or tick none to export them all. On import, matching is done first by name then by domain; a checkbox decides whether an existing competitor is updated or skipped, and a preview mode reports the outcome without writing anything.

Monitored URLs as CSV

Expected columns: product_id, product_reference, competitor_name, competitor_domain, url, price_selector, currency_iso, use_ai_extraction, scrape_interval_hours, active.

Only competitor_name and url are mandatory, plus product_id or product_reference to know which product the URL belongs to. An unknown competitor name creates its record automatically, so a single file is enough to bootstrap an entire setup. The delimiter is detected automatically, and the import starts in preview mode, reporting line by line what would be created, updated or skipped.

Competitor catalogue as CSV

An export of everything the crawl discovered, with the matched product, the method and the score. A checkbox restricts the export to matched rows only. Useful for an offline audit or to share with your purchasing team.

Setting up the cron

Periodic operation relies on a token-secured endpoint shown on the dashboard. Schedule a call every 30 minutes:

*/30 * * * * curl -s "https://your-shop.com/index.php?fc=module&module=dfaicompetitor&controller=cron&token=YOUR_TOKEN" > /dev/null

At each run the cron reads the batch of URLs whose interval has elapsed, sends the pending alert emails, ships the weekly report if it is the right day, advances a running catalogue crawl by one slice, and purges history beyond the configured retention. The response is a JSON summary.

The token can be regenerated at any time from the dashboard, in which case existing cron jobs must be updated. A manual trigger is also available through the Run cron now button.

Alerts

Five alert types are generated, each with a severity:

  • price_drop / price_rise: a move beyond the threshold. A move of 10 % or more becomes critical.
  • undercut: a competitor goes below your tax-included selling price. Always critical.
  • out_of_stock / back_in_stock: availability transitions detected through structured data or the AI.
  • scrape_error: only after 3 consecutive failures, to remove the noise of transient outages.

Alerts are grouped into a single digest email per cron run.

Reports and analytics

The Reports and analytics screen fits in four tabs, behind a permanently displayed KPI strip: products compared, share of products where you are the cheapest, average gap against the market floor, open alerts, scraping health. A period selector covers 7, 30 or 90 days.

Overview

Your market position as a bar (cheapest, in the middle, most expensive), the products where you are losing on price ranked by decreasing gap, those where you have room to raise, a per-competitor table showing who undercuts you most often, and the AI-written executive summary. Opportunities export to CSV.

Alerts

The list filterable by type, severity, state and free text, with a chart of the daily volume by severity. Each row is acknowledged or reopened without reloading the page, and the open-alert counter updates in the strip.

Product analysis

Pick a monitored product: the module plots your price against each competitor over the period, shows the adjustment suggestion with the current price and the proposed price side by side, and lists the sources with an immediate reading button on each.

Since version 1.1.0 every reading also records your own price at that moment. That is what makes the gap history reconstructible. On a shop upgraded from a 1.0.x, the curve of your price only becomes exact from the readings taken after the upgrade; before that date the module falls back to your current price.

Health

The failing URLs, with their status, their consecutive failure count, their last error and a retry button, plus a success rate per competitor. This is the first place to look when a figure seems wrong: a competitor whose scraping has broken skews the averages.

Price suggestion

In the product analysis tab the module shows your current price, the competitor minimum, average and maximum, the sources, and the suggested price from your strategy. The Apply button writes the price onto the PrestaShop product.

Application is never automatic. That is deliberate, to avoid mirrored races to the bottom between competitors running the same kind of tool. The tax-included to tax-excluded conversion is handled automatically from the product’s tax rate.

Troubleshooting

The crawl discovers no page

First check that the competitor’s domain is entered without a scheme and without www.. Then open https://domain/robots.txt in a browser: if it holds no Sitemap: directive, enter the sitemap URL by hand on the competitor record. If the sitemap exists but nothing is kept, your filter pattern is probably too strict: empty it to see what comes back, then tighten it.

The crawl discovers pages that are not products

Set the keep pattern to the segment common to the site’s product pages, for example /product/ or /p/, and the drop pattern to whatever pollutes the result, for example /blog/. Then clear the competitor’s catalogue and start again.

Many pages stay unmatched

First look at whether the competitor pages expose a SKU or an EAN: without a shared identifier, only title similarity can work. Set the SKU selector on the competitor record, run the crawl again, then use Re-match. If your titles differ a lot from the competitor’s, lower the match threshold to 75 and check the resulting matches before leaving it there.

Status “no_price”

None of the methods found a price. Check the URL in a browser, add a price CSS selector on the competitor record, or turn on forced AI extraction.

Status “error” or “blocked”

error means the page did not respond (HTTP 400 and above, or a timeout). blocked means the site’s robots.txt disallows that address to our bot. In the first case, customise the User-Agent, raise the timeout or widen the interval. In the second, it is a decision of the target site that is yours to weigh.

AI extraction is not working

Check that the API key is valid and that the model exists at your provider. AI call errors are logged in Advanced Parameters → Logs with the [dfaicompetitor] prefix.

Emails do not arrive

Test the PrestaShop email configuration in Advanced Parameters → E-mail. The module uses the native mail system with its own EN and FR templates.

Good practice and compliance

  • Leave robots.txt compliance on. It applies to discovery as well as to reading product pages, honours a published Crawl-delay, and applies the standard’s precedence rule: between an Allow and a Disallow that both match, the longer rule wins.
  • Keep a delay between requests of at least one second and a sensible product cap on the first crawl of a site you do not know.
  • Widen the reading intervals: 6 to 24 h per URL is plenty for price monitoring.
  • Keep an identifiable User-Agent. It is the fair practice expected in competitive monitoring, and some sites tolerate a named bot they would otherwise block.
  • Reading public prices is generally lawful in Europe, but you remain responsible for how you use the collected data and for judging the terms of service of the sites you target.
Was this page helpful?

Still stuck? Contact support