APEX FLOWWEB SCRAPER

APEX FLOW WEB SCRAPER / REFERENCE LIBRARY

Facts, capabilities and release evidence

A dated record of what Apex Flow Web Scraper implements, what has been tested and what remains outside this release.

Product
Apex Flow Web Scraper
Publisher
Apex Flow Labs
Engine release
0.1.0 · Early release
Software price
$0 · MIT-licensed engine
Supported runtime
Python 3.11+; tested on Windows and Linux. Rendering requires installed Brave.
Hosted scraping keys required
None for the core workflow
Public account and billing service
Not implemented
Last reviewed
2026-10-08

Implemented strengths

Implemented

Owned execution

Run the core on your computer or server, without a required hosted scraping API account.

Implemented

Field-level evidence

Keep selectors, match counts and source values beside extracted fields.

Implemented

Durable jobs

Inspect progress, cancel work and recover jobs after a worker interruption.

Implemented

Explicit budgets

Bound pages, depth, time, bytes and concurrency before collection.

Implemented

Required-field validation

Expose missing required fields instead of quietly treating incomplete records as complete.

Implemented

Run comparison

Review changed fields and text between saved runs. Missing observations do not prove deletion.

Implemented

Reusable recipes

Store explicit extraction rules and use them across collection runs.

Implemented

Portable outputs

Keep structured JSON and export spreadsheet-oriented CSV; extracted Markdown is also available.

Implemented

Owned-corpus search

Search collected results without sending them to an external model by default.

Interface implemented

Assistant-neutral tools

Nine MCP tools expose the same job workflow to compatible clients. Each client requires configuration.

Verification record

33 local tests passed: 31 core tests and two private web boundary/export tests. Coverage includes public URL policy, robots behavior, cache revalidation, byte accounting, cancellation, recovery, rendering, recipe evidence and CSV formula escaping. An owned-server smoke run returned example.com with HTTP 200 in approximately 1.55 seconds. One page is not a throughput benchmark.

Website checks

The initial Geode homepage sample extraction passed at widths of 375, 768 and 1440 pixels with no horizontal overflow. The automated accessibility scan reported no violations in the tested rule set. Automated checks do not establish full accessibility conformance or guarantee every device behaves identically.

Known limits

No residential proxy network, CAPTCHA solving, authenticated account collection, OCR, global search index, visual workflow recorder, immutable snapshot archive, automatic recipe repair or contracted uptime SLA. Browser resource restrictions can prevent some sites from rendering. Retained results need operator storage management.

Ownership and dependencies

Apex owns its orchestration, job, evidence and interface code. The engine uses open-source foundations including Python, aiohttp, Beautiful Soup, lxml, Playwright, pypdf and the MCP SDK. Their licenses remain applicable. The public engine license does not grant rights to the commercial Geode website design.

Review the source and tests →