Public web crawler

How Modou reads approved public websites

ModouDiscoveryBot checks selected public pages to find timely business opportunities for Modou customers. It is a small, polite crawler rather than a general internet search engine.

Reviewed 24 July 2026

Purpose and identity

The crawler reads approved public registry pages and newly discovered public detail pages. Customer watch pages and customer search-result pages are not added to the shared crawler in this first collection stage. Requests identify as ModouDiscoveryBot/1.0 and link back to this page.

What it can access

The crawler uses ordinary HTTP first. A separate, tightly limited browser may render JavaScript only for a source that Modou has explicitly approved for browser access. It does not sign in, solve CAPTCHAs, bypass paywalls, use proxies, download files through the browser, or bypass website controls. It checks robots.txt before fetching a page.

Rate limits

The HTTP worker makes no more than two requests at once and the optional browser handles one page at a time. Modou never requests one host in parallel and waits at least five seconds between requests to the same host.

What is stored

Modou stores the page title, up to 20,000 characters of cleaned public text, source links, language, publication and fetch dates, and HTTP validation data. It does not store raw HTML and removes email addresses and phone numbers from indexed text.

Removal and correction

You can block the crawler through robots.txt. To ask Modou to remove an already indexed page or correct the crawler's behaviour on your site, email hi@modou.io with the affected URL. We will review the request and remove eligible indexed content.

Crawler questions and removal requests

Email hi@modou.io and include the website, affected page URLs, and a short explanation. We do not require you to create a Modou account.