APEX FLOW WEB SCRAPER / REFERENCE LIBRARY
How to budget a web crawl
A crawl budget should cap pages, elapsed time, bytes and concurrency. A page count alone cannot bound the cost of large files or expensive browser pages.
Set explicit limits
The current Apex job API accepts 1–1,000 pages, depth 0–10, 5–3,600 seconds and a byte allowance up to 500 MB. HTTP concurrency is capped at four; browser rendering is serialized. These are release limits, not a promise that every maximum-sized run will succeed.
Understand what bytes mean
The HTTP fetcher charges received body bytes to the job budget and bounds decoded response size. Browser subrequests are also mediated by the guarded fetcher. Cache use and retries influence the work done; inspect events and counters when comparing runs.
Leave room for overhead
Robots requests, redirects, retries, parsing and browser startup add work beyond the number of useful records. Saved output also consumes disk space. This release does not provide a global retained-result disk quota; operators must monitor and manage storage.
A sensible pilot
Start with ten pages, one depth level and a short deadline. Inspect average response size and failure reasons. Increase one limit at a time. When a limit is reached, keep the completed records and report the limited state explicitly.
← All reference articles · Current capabilities and limits →