APEX FLOWWEB SCRAPER

APEX FLOW WEB SCRAPER / REFERENCE LIBRARY

Web crawling vs. web scraping

Crawling chooses the next URLs to visit. Scraping extracts useful information from the pages that were visited. A reliable system needs separate limits and evidence for both.

By Apex Flow Labs · Reviewed 2026-10-08

Design the frontier

Use a queue of normalized URLs with an explicit page budget and depth limit. A ten-page job should admit at most ten pages, even when several workers discover links at once. Apex uses an atomic database operation to enforce the frontier budget.

Avoid accidental expansion

A category page can link to login pages, filters, language variants and endless combinations of query parameters. Restrict the starting scope. Apex follows same-origin links, uses depth limits and removes fragments for URL identity. Review query handling before collecting signed or parameter-sensitive URLs.

Measure coverage separately

A completed job means the admitted work finished. It does not mean every page on the site was found. Record pages discovered, admitted, completed and failed separately. Pagination, JavaScript navigation and orphan pages can leave gaps.

A practical starting point

Try one seed, ten pages, one level of links and HTTP-only mode. Inspect the results before expanding. If only the seed matters, set max_pages=1 and avoid a crawl altogether.

← All reference articles · Current capabilities and limits →