APEX FLOWWEB SCRAPER

APEX FLOW WEB SCRAPER / REFERENCE LIBRARY

Robots rules, rate limits and retries

A reliable collector reads access instructions, paces requests and reports refusals. Repeatedly retrying a blocked page is not a useful success strategy.

By Apex Flow Labs · Reviewed 2026-10-08

Robots handling

Apex retrieves robots.txt per origin and applies its parser before page collection. Missing robots files (404 or 410) are handled differently from refusals or server errors. Rules cached for one host must not accidentally apply to another host.

Use bounded retries

For rate limiting or temporary server failures, use a limited number of retries and honor a bounded Retry-After delay. A request that repeatedly fails should end with an explicit failure, not an indefinite spinner.

Know the release behavior

The collector retries selected transient responses, including 429, 502, 503 and 504, with at most three attempts. A robots crawl delay outside the supported bound stops collection for that path rather than silently ignoring the rule.

Separate policy and mechanics

Robots directives are a technical input, not a complete determination of permission or legal rights. Choose sources appropriate to your use and do not treat a successful response as permission to republish everything in it.

← All reference articles · Current capabilities and limits →