Owned execution
Run the core on your computer or server, without a required hosted scraping API account.
APEX FLOW WEB SCRAPER / REFERENCE LIBRARY
A dated record of what Apex Flow Web Scraper implements, what has been tested and what remains outside this release.
Run the core on your computer or server, without a required hosted scraping API account.
Keep selectors, match counts and source values beside extracted fields.
Inspect progress, cancel work and recover jobs after a worker interruption.
Bound pages, depth, time, bytes and concurrency before collection.
Expose missing required fields instead of quietly treating incomplete records as complete.
Review changed fields and text between saved runs. Missing observations do not prove deletion.
Store explicit extraction rules and use them across collection runs.
Keep structured JSON and export spreadsheet-oriented CSV; extracted Markdown is also available.
Search collected results without sending them to an external model by default.
Nine MCP tools expose the same job workflow to compatible clients. Each client requires configuration.
33 local tests passed: 31 core tests and two private web boundary/export tests. Coverage includes public URL policy, robots behavior, cache revalidation, byte accounting, cancellation, recovery, rendering, recipe evidence and CSV formula escaping. An owned-server smoke run returned example.com with HTTP 200 in approximately 1.55 seconds. One page is not a throughput benchmark.
The initial Geode homepage sample extraction passed at widths of 375, 768 and 1440 pixels with no horizontal overflow. The automated accessibility scan reported no violations in the tested rule set. Automated checks do not establish full accessibility conformance or guarantee every device behaves identically.
No residential proxy network, CAPTCHA solving, authenticated account collection, OCR, global search index, visual workflow recorder, immutable snapshot archive, automatic recipe repair or contracted uptime SLA. Browser resource restrictions can prevent some sites from rendering. Retained results need operator storage management.
Apex owns its orchestration, job, evidence and interface code. The engine uses open-source foundations including Python, aiohttp, Beautiful Soup, lxml, Playwright, pypdf and the MCP SDK. Their licenses remain applicable. The public engine license does not grant rights to the commercial Geode website design.